Google’s new Gemini 1.5 AI can dive deep into oceans of video and audio

Just last week, Google unveiled its new AI chatbot lineup, featuring Gemini Advanced—its best bot, based on its most powerful large language model, Gemini 1.0 Ultra. But Gemini 1.0 Ultra’s reign as the company’s flagship LLM could turn out to be brief.

Today the company is announcing Gemini 1.5 Pro, an update to its middle-tier LLM. It says the improvements result in an LLM in the same zip code, power-wise, as Gemini 1.0 Ultra. And in a briefing for reporters on Wednesday, Google DeepMind principal scientist Oriol Vinyals showed off videos of Gemini 1.5 Pro performing some pretty spectacular feats of AI.

According to Google, Gemini 1.5 Pro punches above its weight in part because it’s engineered for efficiency, both when it’s being trained and when it’s generating content. It can also handle more tokens—the data points an LLM divides a piece of content into to process it. Gemini 1.0 could deal with 32,000 tokens at a time. By default, Gemini 1.5 has a capacity of 128,000 tokens, the same as OpenAI’s GPT-4 Turbo model. But Google will let some customers try a version with a capacity of 1 million tokens, and says it’s tested the LLM with 10 million tokens.

Those of us who aren’t AI scientists may have trouble getting our heads around those numbers. For Gemini 1.5 Ultra, they translate into an hour of video, 11 hours of audio, more than 700,000 words of text, or 30,000 lines of programming code—all of which help Gemini 1.5 deal with inputs that are way more complex than a typical typed-in prompt or photo of your cat.

During its press briefing, for instance, Google showed a video in which it fed more than 400 pages of transcribed air-to-ground audio from the Apollo 11 moon landing to Gemini 1.5 Pro, which divvied it into 326,678 tokens. That allowed the LLM to ace the request “Find 3 funny moments. Make a list with only the quotes.” When Google gave the LLM a scrawled drawing of an astronaut’s boot taking a step, Gemini 1.5 Pro figured out that it referenced Neil Armstrong’s iconic declaration.

In another demo, Gemini 1.5 Pro turned Buster Keaton’s 45-minute silent comedy Sherlock Jr. into 696,161 tokens. It was then able to summarize the film’s plot, answer a question about the writing on a slip of paper that appears partway through it, and pinpoint the moment represented by another hasty sketch. In a third demo, the LLM ingested a grammar guide for Kalamang—a language spoken by fewer than 200 people—and was then able to translate between it and English with human-like proficiency, according to Google.

Why didn’t the company focus on readying Gemini 1.5 Pro for deployment rather than immediately applying its new advances to its top-of-the-line Ultra version, which would theoretically result in an even more, well, Ultra LLM? The bigger an LLM’s training set, the trickier it is to make it perform satisfactorily, which gave the midrange Pro version an advantage as a test bed for Google’s latest work.

“Very naturally, the first set of models that we trained to completion is the Pro series, which is on the smaller side compared to Ultra,” Vinyals told me during the briefing. “That’s the reason why, in general, this might become available earlier.”

For now, the Gemini 1.5 Pro LLM is in private testing with a select group of customers of Google’s Vertex AI cloud service and AI Studio software development platform. Google isn’t saying when its power might be available to more developers or—via its Gemini chatbots—mere mortals. Nor did Vinyals share anything about what Gemini 1.5 Ultra might be able to accomplish or when it could appear.

But with Google’s AI rivals also making progress at a furious clip—on Tuesday, The Information’s Aaron Holmes reported that OpenAI is developing a search engine—the company has every incentive to make its best LLM available far and wide as soon as it can.

https://www.fastcompany.com/91029527/google-gemini-1-5-ai-llm?partner=rss&utm_source=rss&utm_medium=feed&utm_campaign=rss+fastcompany&utm_content=rss

Created 10mo | Feb 15, 2024, 4:30:04 PM

Other posts in this group

TikTok is full of bogus, potentially dangerous medical advice

TikTok is the new doctor’s office, quickly becoming a go-to platform for medical advice. Unfortunately, much of that advice is pretty sketchy.

A new report by the healthcare software fi

Dec 25, 2024, 12:30:03 AM | Fast company - tech

45 years ago, the Walkman changed how we listen to music

Back in 1979, Sony cofounder Masaru Ibuka was looking for a way to listen to classical music on long-haul flights. In response, his company’s engineers dreamed up the Walkman, ordering 30,000 unit

Dec 24, 2024, 3:10:04 PM | Fast company - tech

The greatest keyboard never sold

Even as the latest phones and wearables tout speech recognition with unprecedented accuracy and spatial computing products flirt with replacing tablets and laptops, physical keyboards remain belov

Dec 24, 2024, 12:50:02 PM | Fast company - tech

The 25 best new apps of 2024

One of the most pleasant surprises about this year’s best new apps have nothing to do with AI.

While AI tools are a frothy area for big tech companies and venture capitalists, ther

Dec 24, 2024, 12:50:02 PM | Fast company - tech

The future belongs to systems of action

The world of enterprise tech is built on sturdy foundations. For decades, systems of record—the databases, customer relationship management (CRM), and enterprise resource planning (ERP) platforms

Dec 23, 2024, 10:50:06 PM | Fast company - tech

Bluesky users report AI bots, disinformation, and copycat accounts

Bluesky has seen its user base soar since the U.S. presidential election,

Dec 23, 2024, 10:50:05 PM | Fast company - tech

Banning Chinese-made drones could hurt some Americans

Russell Hedrick, a North Carolina farmer, flies drones to spray fertilizers on his corn, soybean and wheat fields at a fraction of what it

Dec 23, 2024, 8:40:03 PM | Fast company - tech

Tomas_r2