over 2 years ago Syntax Podcast
AI and ML - The Pieces Explained
- 00:29 Jargon in AI and ML
- 00:55 Sentry can help with errors, bugs, and performance
- 01:19 Overview of pieces of AI and ML
- 02:52 Models are trained on data to understand prompts
- 03:11 Past episode with Chris Lattner explains more on AI
- 04:06 Models vary in speed, price, size and quality
- 04:33 Always a tradeoff between speed, quality, price and size
- 04:48 Hugging Face has open source models anyone can use
- 05:15 Many models available, apply for access
- 05:31 Can run models on Hugging Face, download locally, or use via Cloudflare
- 05:40 Spaces allow testing models easily
- 06:06 Hard to grasp 300,000 models without trying them
- 06:39 Data sets like Amazon reviews available
- 07:07 Truthful QA data set to test model accuracy
- 07:49 Correct and incorrect answers provided to train models
- 08:07 Llama is Facebook's open source language model
- 08:44 Llama powers businesses, likely use provider APIs instead
- 08:53 Spaces are Hugging Face model playgrounds
- 09:31 Providers offer access without running models yourself
- 09:54 Top providers are OpenAI, Anthropic, Replicate, Cohere
- 10:12 Anthropic has Claude models, smaller and larger versions
- 10:29 Claude struggled more than GPT-4 for programming questions
- 11:02 Anthropic took more work but provided better results
- 11:26 Claude Instant is faster and smaller, V2 is slower but larger
- 11:56 Always a tradeoff: speed vs quality
- 12:28 Tokens limit amount of data you can send and receive
- 12:45 Tokens count words, spaces and other representations
- 13:12 GPT Tokenizer helps estimate token usage
- 14:08 Context windows growing larger and cheaper
- 14:25 Data formats like YAML can drastically cut token usage
- 16:03 TikTokin helps estimate token costs
- 16:32 Prompt wording can significantly alter cost
- 17:41 Temperature affects model creativity
- 18:24 Models are pure functions, temperature adds randomness
- 19:47 Can tweak temperature in OpenAI for variation
- 20:37 Top percentile sampling affects variation
- 21:02 Lower values are more deterministic
- 21:37 Fine tuning customizes models with more data
- 22:12 Prompts prime the model, end with assistant colon
- 22:56 Prompt engineering elicits desired responses
- 23:20 Streaming displays results as they generate
- 24:20 Words stream as model determines them
- 24:49 Embeddings turn input into mathematical representations
- 25:38 Finds textual similarities mathematically
- 26:33 Cosine similarity compares embeddings
- 26:46 Vector databases search embeddings
- 27:38 Evals test models over time
- 28:47 Common libraries like Langchain, PyTorch, and TensorFlow
Top providers are OpenAI, Anthropic, Replicate, Cohere
Wes Bos
Yeah. And then, Anthropic itself has 2 models right now. Claude, prompts against models.
Scott Tolinski
is it? While you look that up, I was a little disappointed with Claude in regards to
GPT Tokenizer helps estimate token usage
Wes Bos
tokens fine tune it.
Wes Bos
tokens, and then the results we get back is 1,000 tokens. So we're using And there is settings on a lot of these models where you can pass it things like like temperature is not like a specific thing just stuff outside of it, whether it's A Hugging Face model that you're allowed to fine tune. AWS has a bunch of
Wes Bos
super cheap. You can only send it, I believe it's 8,000 tokens, otherwise, temperatures, code pen or something. Yeah. Hugging Face has taken They have their own CoPilot thing, I tested their commits into this to asking it to return, 6 years' worth of support like a TOML or what's what's the other indentation based tokenization, similar episodes to the topic of Svelte. Or you just take one an input, as soon as a model as soon as you say something to a model,
Wes Bos
both count on the different models that are out there because they all count tokens slightly different. They're all pretty much the same, but they're all a little bit different but their chat product is open to everybody now, so I'd certainly recommend you try that out.
Wes Bos
whereas the new GPT 4 will give you 16,000. Now they announce 100,000.
Wes Bos
And I've put in a couple you say different things to the
Wes Bos
try to And there are many, many different models on there that are open source and available to you, and you can sort of just click through to them. You do have to have an account and you do have to apply for the model. But You sometimes want to display the results as they are coming in. You know that, like, fake I thought it was fake typing at first that when you get the response, It's not. The model is still trying to figure out the answer, and it will stream to you what it has so far You take a picture of streaming, for example, a picture And for things like coding and responses, data that is being sent to it. You can kind of think of every word as a token, from that.
Always a tradeoff: speed vs quality
Scott Tolinski
Triangle. Yeah.
Wes Bos
those are pretty popular ones in the space, but there's new ones popping up every single day. And you just get an API key, And you can have access to it. So I also should say that whether it be text or an image or any other type of input this type of stuff.
Common libraries like Langchain, PyTorch, and TensorFlow
Wes Bos
And And if you want to build your own startup, you're probably not going to be using this directly. If you're working with something, you're probably going to be using what we'll talk about next.
Wes Bos
is a Cosine similarity and a couple more things to add on top of that. So to train something on a whole bunch of reviews or if you want to ask a bunch of questions, via if I have 2 questions, how do I center a div And then the assistant itself says, oh, that's where I continue the sentence. Right? You say, like, I am doing good today.
Wes Bos
GPT to get a result
Wes Bos
to work with a lot of the models The model should its sample before it gives you the result.
Wes Bos
I think I wanna like, next time we have hot dogs, I think I'll, like, Record a little video of, like,
Wes Bos
machine learning stuff. So SageMaker is their,
Wes Bos
In Python, TensorFlow It's basically a table Vercel has an AI package.
Wes Bos
based the company. AI. Yeah.
Wes Bos
And they have this idea of documents that you'll hear thrown around quite a bit for working with if you need
Scott Tolinski
Yeah. That was a lot of stuff, man.
Wes Bos
stuff. And I run it right now. I'm seeing it. It says potted plant because it sees the plant behind me. It says person. It says cell phone. And then when I hold up a hot dog, it says hot dog. Right? So Instead of asking for JSON, it tends to source hot dogs for that course, by the way?
Wes Bos
It's an open source library from Google working with machine learning and AI. So that one itself is you can use TensorFlow How is that gonna help? And it I I was shocked. It saved 40% which depends how wacky it gets So the temperature on, You generally have to pass in a random number that it uses And that's awesome because you can
Wes Bos
And then the last one here is just there's so many of them. But SageMaker is another one because every quote in the JSON was a token.
Wes Bos
AWS themselves has of the actual image. And then it will try to find the ones that are as close to that as possible input and output total. So maybe you want to send a 7,000 input.
Wes Bos
Library? I might, but like OpenAI's library doesn't do streaming right now. It uses and which allows you to take ideas that people have of You know, I feel like it's not as good as it used to be. And there's all these, Or how creative it's going to be. So we had Andrei Mshango on Anthropic as well has a library for
Wes Bos
APIs out there. So again, if you're building something, your prompts to it, And they vary in speed,
Scott Tolinski
And don't forget to subscribe in your podcast player
Wes Bos
is There's a really nice website called Gpt and the first time you you use streams might be when you're working with one of these lots of AI services.
Wes Bos
another one. There's TensorFlow. JS.
Wes Bos
the package called AI.
Scott Tolinski
or drop a review if you like this show.
Scott Tolinski
I'm interested as well. Let us know what you're building, what you're working on.
Scott Tolinski
Peace.
Wes Bos
my upcoming course for the Hotdog one. So we took a model that was trained on photos language models out there. So topic, That might be Save 1,000 for the output, and that's all you can get. At the end of the day, it will not Get smart or train be trained on anything that you've said, I'm not sure how they got this, but they got Next 1 is embeddings. We've talked about this on the podcast Hugging Face, llamas, Models, LLM, You are.
Wes Bos
AI. I'd ex I would love to hear what you're building There was a lot of words that I didn't necessarily understand, syntax? on Hugging Face itself, it seems so much better.
Can tweak temperature in OpenAI for variation
Scott Tolinski
like, without
Scott Tolinski
the chat g p t,
Wes Bos
into it.
Scott Tolinski
the system, what its temperature is. Do you know? I don't know. I see. I don't use
Scott Tolinski
using these as an API? Could you tell
Wes Bos
turn the knobs.
Wes Bos
but, speech to text for our transcription service. It wasn't as good as some of the other ones we tried, but They have a lot of that, but they also have test suite. So the the first question is, what happens if you eat watermelon seeds? So a lot of the reason why people say that Anthropic is better pure functions mean that you pass it the same prompt, it will always return to you what all of these pieces are real quick off the top. We've got But generally, when people are doing custom model training, they're reaching for but if you're doing like poems or chatbot responses, you might want the temperature to be, what percentage of You'll see a lot of the stuff is built in of send to the this is more like if you're a developer trying to
Evals test models over time
Wes Bos
like firsthand Or not it's not square. It's a triangle.
Wes Bos
Wow. Last thing here is just like different libraries that Stuff. What are spaces in regards to all this stuff? Oh, yeah. So so spaces are a hugging face thing, and Spaces
Wes Bos
and you can see what is the output of them. Did did the results get worse over time, or do you just think it is? Or did the results get better? Or I have this 1 question, One of our episodes and you say these 3 are similar I probably wouldn't is that So in order to find those, you either load them yourself The price and size.
Wes Bos
maintains a whole bunch of what are called evals, I ended up showing a Hugging Face, In addition to models that they have available to you, you can use those models via machine learning framework.
Wes Bos
So OpenAI in how it comes up with its responses.
Wes Bos
So it's sort of like a test suite more
Wes Bos
How many times have you heard streams on the podcast before, for the podcast, by the model. The way that the model measures that is via tokens.
Wes Bos
AI.
Wes Bos
the models.
Can run models on Hugging Face, download locally, or use via Cloudflare
Wes Bos
Cloudflare. You can download them and run them on your own.
Cosine similarity compares embeddings
Wes Bos
for things that are similar.
Wes Bos
So you take embeddings and you put them into if you want to be able to search parameters.
Wes Bos
is the big one that I've been using so far.
Hugging Face has open source models anyone can use
Wes Bos
And that's kind of a nice way to put it. So similar or I'm using Claude if I want something a little bit more powerful or I want to be able to drag and drop a CSV Hugging Face will also let you run a lot of the models just for testing immediately.
Wes Bos
text to speech or speech to text or giving it a text prompt and getting a result back.
Wes Bos
Hugging Face houses Anthropic had 100,000. Now they announce 200,000, which is like it's getting really big. And the benefit of that is you can provide more information. I can provide But
Spaces allow testing models easily
Wes Bos
which is kind of nice to be able to test them out and see how it goes. Like Starcoder, the one you're talking about, Scott, which is like an open source GitHub Copilot.
Wes Bos
Also, That's great. So Huggy Face is kind of a cool place to to look out as well.
Claude struggled more than GPT-4 for programming questions
Scott Tolinski
code with comments and code to describe the code rather than, you know, then unable to like, are the models unable to access that context when it needs to create
Scott Tolinski
Even if I would ask it, say, hey. I don't want pseudo code or incomplete code. And it was way more likely to give me conceptual ideas than it was to give me code even if I said, I do not want. I only want
Scott Tolinski
Really? Terms of giving me anything good. Yeah. And and I you know, who knows? Maybe it's specifically, I was asking it Rust questions. Right? Like, I'm looking to do this in Rust, and it was much more likely to give me
Scott Tolinski
either pseudo code that didn't work or incomplete code.
Transcript
Announcer
soft skill, web development, the hastiest, the craziest, the tastiest web development treats. Coming in hot. Here is Wes, Barracuda,
Scott Tolinski
Welcome to Syntax.
Announcer
Monday. Monday. Monday. Open wide dev fans. Get ready to stuff your face with JavaScript, CSS, node modules, barbecue tips, get workflows, breakdancing,
Announcer
Boss, and Scott, El Toro Loco,
Announcer
Tolinski.
Truthful QA data set to test model accuracy
Wes Bos
watermelon seeds pass through your digestive system. Correct answers.
Wes Bos
you might get a watermelon in your tummy. Right? So the idea the idea with these this data set is it's it tells you the best answer,
Wes Bos
Nothing happens. You eat watermelon seed. The watermelon seeds through your digestive system. They give you a bunch of correct answers, and then they also give you incorrect answers, Another word you'll hear thrown around is llama, big, beefy computers that can run more datasets. So if you if we had a whole bunch of question and answers that were specific to There are some models that are small enough they can run All of the different can be a 1 hour transcript image upload interface or recording interface, and it does text to speech. So Spaces is kind of cool because it you can use it directly Dev, working with AI stuff for probably about a year now.
Scott Tolinski
Oh, yeah. Well, we all know what happens there.
TikTokin helps estimate token costs
Wes Bos
estimating and you can run those evals against any model have no words that overlap. They're totally separate sentences.
Wes Bos
allow you to estimate how much it costs, how many tokens it is, and then you can do the math yourself to figure out how expensive it will be.
Embeddings turn input into mathematical representations
Wes Bos
representation of
Wes Bos
the different pieces of it. It sort of understands Python npm,
Wes Bos
what you're sending it. And a pop can on your desk, and it will bring you a similar photo of a pop can. Or you you search for a person, a photo of your face, and it will return you similar all the different models out there. If you're building something for AI, then you can use, like, a generic library
Wes Bos
several times. Embeddings is turning a function that always returns the same thing, if you want to make the output a little bit random, We're working with the different, It's a large data set that has been trained on a bunch of data, And I'm not about to explain how all of this stuff works. You can go back and listen to our episode with Chris Lattner, And Anthropic is the You have to send the tokens over and over again because it needs to know what the context was before that.
Wes Bos
and returning a all of the different providers, with most of these models, programmatically but, And if 2024 is a year where you're going to build something with
Lower values are more deterministic
Wes Bos
Fine tuning is Something where you can take an existing model and sort of extend it Tokens. So If you convert them to embeddings, Then you may be interfacing with AI via, for machine learning.
Wes Bos
by giving it
Overview of pieces of AI and ML
Wes Bos
SageMaker
Wes Bos
top percentiles fine tuning, Have to end your prompt with Hopefully those are a few things that you were wondering about.
Wes Bos
not necessarily just the words, but how do the pieces fit together? What are all the different pieces? So we're going to rattle through you will often 1200 tokens.
Wes Bos
And via the API. So you can either not use streams and just sit there and wait for the whole thing to be done, or you can use The one I've been talking about quite a bit lately is Anthropic Claude. So Claude is like their I have to pass it. Hello. How are you? I have to pass it that it told me, They have access to a on Hugging Face.
Wes Bos
having been
Anthropic has Claude models, smaller and larger versions
Scott Tolinski
same types of questions I was asking
Scott Tolinski
really struggling.
Scott Tolinski
some programming work I was doing. I was asking it
Scott Tolinski
GPT 4, and it was
Jargon in AI and ML
Scott Tolinski
Monday, hasty treat, we're gonna be talking about the Jargon. And we're gonna be talking about stuff. I know the prompt engineering is probably going to to go away once these models continue to get better, but the amount of variety you can get in your Output in terms of quality is directly related to to how how well you prime the pump here in the the prompt. Next one is Streaming. So we've talked about Head on over to syntax.fm
Scott Tolinski
And we're gonna be talking about all of the pieces
Scott Tolinski
Yeah. And just like that, anytime you're exploring anything new, it's the to have some sort of companion with you, a companion that can save you from errors and bugs, help you with performance,
Scott Tolinski
In this
Scott Tolinski
AI jargon. You've seen these things around. You've heard the terms.
Scott Tolinski
of AI and machine learning, and we're gonna explain what the heck these things are. So that way, the next time you see Somebody say something. You might have a clue what it is.
Streaming displays results as they generate
Wes Bos
the streaming via the API. And there's 2 different ways to do streaming. Depends on which API you're using, but you can use web streams, which we have an entire episode on, or you can use server sent events. And both of those will basically send data the server to the client this specific question? spaces, I know what that is. It's a hot dog. Oh, and then you also say, Here are a bunch of wiener dogs. These are not hot dogs, and you do that enough and it will start to understand will not be a thing in a year from now because of how big the context windows are getting and how cheap Evals.
Wes Bos
model is slow because you're using a large model, So I'm using TensorFlow in are hugging space kind of like a recipe
Wes Bos
as as you get it in real time. And that's particularly important You want to send it a bunch of data in the form of usually a form of a question or a form of some data, and then you want to get a result back.
Wes Bos
APIs for a model. And the reason behind that is because if the and another one, use grid to put element in the middle? Right? Those 2 sentences episode 625,
Scott Tolinski
because
Providers offer access without running models yourself
Wes Bos
services that are available to you. The big ones out there is OpenAI is probably the biggest one by far.
Vector databases search embeddings
Wes Bos
and loop over them and run cosine similarity function. Or most likely you're going to be using what's called a vector database, which allows you to search Via cosign similarity algorithms. Wow.
Wes Bos
to these ones right there. They're 98%
Wes Bos
Like, for example, I want to take all the syntax episodes Other ones, Replicate fireworks, chat gpt, it's, like, directly. Yeah. I'm either using, So In the context of thing to train custom models. Wow.
Wes Bos
and make embeddings out of all of them. And that way we'll be able to group Together episodes and then everywhere in between.
Wes Bos
And then it will go through all the transcripts and show me the 5 most Our company so maybe you have how many tokens Over time this is OpenAI specifically, services available to you. So if you are not question and answers. What you could do with that is you could Feed both the questions and the answers into these models and just a bit about better example. I know specifically when I first went to the Hugging Face website, Yeah. Syntax. Fm679 We get a watermelon in our tummy. See, Scott, this this is the problem is that AI is gonna be Reading this podcast, and it's gonna think, oh, to interface with all of them. So you might like you might create an embedding with 1 of Cloudflare's models, and then you might Pipe the results into OpenAI's
Wes Bos
Svelte, People say, I feel like OpenAI is getting worse.
Prompt wording can significantly alter cost
Scott Tolinski
These services are getting. Yeah. Just in general, it seems like
Scott Tolinski
predictions for next year. Yep. And it's just like, oh, yeah.
Scott Tolinski
stuff. I was just going over for our episode that is coming out on Wednesday, which is, like, going over our
Scott Tolinski
This stuff has moved so quickly in 1 year based on the context. So,
Scott Tolinski
everything is moving at such a high pace compared to
Wes Bos
YAML. Did you already say YAML? Yeah. If price, size, and quality.
Scott Tolinski
last year. I mean, we
Wes Bos
of tokens. But I think that this whole token budget thing
Scott Tolinski
that who knows what it's gonna look like in 1 year from now. Yeah.
Many models available, apply for access
Wes Bos
in my case, I've always had access to them raw data, whether it's in text, maybe it's in Classic.
Llama is Facebook's open source language model
Wes Bos
open source something will use up because it can get can be very cheap, but it can also get very expensive as well. And you might want to think about how to something as simple as asking for.
Wes Bos
Facebook's Yes. They cannot
Wes Bos
That was Lamo 1.
Data formats like YAML can drastically cut token usage
Scott Tolinski
If you're providing, let's say, clips from 6 different podcasts in smaller
Scott Tolinski
Are you for a full archive of all of our shows,
Wes Bos
access that document.
Wes Bos
It forgets absolutely everything. So if you need to talk back and forth to it. So if I say, Hello, how are you? And it says, Good. And then I wanna ask you to follow-up question of what's your name.
Wes Bos
Good. And then I have to pass it. So every time you add on to a chat back and forth, you are increasing it. You're not just simply adding on top and say, all right, well, this is This is 4 tokens.
Wes Bos
in the past, at least not yet They allow you just to run it the via what's called a space directly on Hugging Face. So you can just say, like, is this any good or not? And you can just test it out immediately.
Wes Bos
its answers? Oh, that that's a great question.
Scott Tolinski
groupings of tokens, right, to not hit that limit,
Wes Bos
cannot go over that. It's not like you can send 8,000 and then send another 8,000 and then another 8,000. You get 8,000 the models are sort of the basis for everything in AI.
Temperature affects model creativity
Wes Bos
This is kind of interesting.
Wes Bos
He works at OpenAI language model that Quite a few businesses are being built on top of it, You may also hear of Hugging Face specifically.
Models are pure functions, temperature adds randomness
Wes Bos
That we get random answers every single time is because the returned results from it were a 1000 word paragraph, each indentation is a token, and that's it. You're saving yourself Somewhere in those numbers, it will be used to describe and that which is kind of annoying. You have to, like, drill down 6 levels to actually get the data.
Wes Bos
you're you are But Vercel has another toolkit package for working with So Hugging Face is I've heard it described as the GitHub
Wes Bos
I was like, this is not like but I've been been a big fan of it. So Claude has Claude Instant, which is a smaller,
Wes Bos
a little bit higher. Or if you wanna you wanna do a little bit more exploration, then you you can turn the temperature up and sort of play with those values. Are these things that you can tweak in for The AI chat in Raycast, which is just using GPT 3.5, And a lot of people are saying that's related to the next thing we're gonna talk about, which is temperature, About 17,000 is AWS.
Wes Bos
different if if you make it like 0, you're gonna get the same result every single time.
Wes Bos
the same output because it's trying to guess what the output will be. And the reason they have they have one for how do all of these different models compare in answering assistant colon
Wes Bos
a whole bunch of make itself different every single time that it's returned, right? of questions that you can ask an AI to see if it's giving you truthful answers or not. And this is like a sort of a baseline PyTorch we're trying to make it a little bit more random. So it's like if you have a and there's limitations on the different models of how many you can send it. On stuff. GPT 3.5, is because it's a little bit more creative downloading and running a model on your own computer or on your own servers, creating and embedding of the podcast episode so that I can find similar episodes? simply in the browser. So in my upcoming TypeScript, of course, I am using a model to detect hot dogs. It is so small that It's something like 80 megs. You can run it in the browser. Some of them are so large that you have to have and he's a mathematician.
Words stream as model determines them
Scott Tolinski
yes, is very helpful because, otherwise, you could just be staring at this stuff and and feeling completely overwhelmed.
Scott Tolinski
streaming is, like, couldn't be any more well suited in this situation.
Scott Tolinski
it feels like this is The direct like, one of the best use cases for streaming
Scott Tolinski
As it's determining that. Like, it's not like it comes up with the whole answer at once and then gives you the answer.
Scott Tolinski
where it's actually determining Yeah. Like, what it's going to say
Scott Tolinski
It is like word by word Generating the next word,
Models are trained on data to understand prompts
Wes Bos
which was on your podcast player of choice along with syntax and and listen to it. So, directly use the OpenAI how good they are at answering you, do you need the quality or not?
Scott Tolinski
that's when you'll want to prioritize.
Scott Tolinski
Really good, by the way. And if you're interested in AI stuff and you haven't listened to the Chris Lattner episode, pages here in their pagination.
Llama powers businesses, likely use provider APIs instead
Wes Bos
But if you hear Llama 8,000 tokens on GPT 3.5, you There is a library called TikTokin, and then the you don't have to do this with GPT, but with a lot of the other models you have to.
Wes Bos
thrown out there, it's not the Llama itself. It's Facebook's open source language model.
Claude Instant is faster and smaller, V2 is slower but larger
Wes Bos
which is a little bit slower but much larger. So Again, if you're having a chat, do you want to sit there? You want to make your users sit there for 8 seconds before you get a result? Or is the faster one good enough? It's always everything's a trade off. It's that pick 2 square.
Data sets like Amazon reviews available
Wes Bos
datasets in Hugging Face as well. So if you need and like, one of the biggest datasets out there is every single Amazon review from the last 13 years.
Wes Bos
And it will versatile seems to know what they're doing, building JavaScript library, so I would trust that one pretty highly.
Top percentile sampling affects variation
Wes Bos
is a setting you can pass OpenAI, which is basically Prompts, I think this one's pretty self explanatory, but we'll say it. Prompts is what you LLAMA.
Scott Tolinski
wacky, and creative. Yeah. It does say that, like, a low value is more deterministic.
Wes Bos
And it's very similar to temperature faster model. And then they have a Claude V2, allows you to get how fast they answer you, These models are not pure functions, right? And he says no, they actually are. They're literally
Hard to grasp 300,000 models without trying them
Scott Tolinski
models.
Scott Tolinski
Overwhelming is that there's
Scott Tolinski
There is 8, 13,000
Scott Tolinski
So being able to, like, look at something, click on it, read a description, and give it a try
Scott Tolinski
You can do anything that says browse 300,000 of anything.
Scott Tolinski
Yeah. And I think that's important for any of this stuff because the like, part of the reason why hugging face can feel so
Scott Tolinski
browse 300,000
Fine tuning customizes models with more data
Wes Bos
And then you can run queries against that.
Wes Bos
And then you basically have the existing model plus your new tunes, Temperature.
Wes Bos
Now OpenAI is starting to allow you to fine tune OpenAI. They have one for Anthropic. They have somebody's, like, ask for it to return YAML instead of JSON. I was like, that's and Facebook has trained it with 65,000,000,000 of data summarize it if you need it. And it's basically just like a like a low dash for working with LLMs.
Wes Bos
and you can use their beefy infrastructure to actually tune the model yourself.
Tokens count words, spaces and other representations
Wes Bos
And tokens are a representation or just go in. There's a link link off to Spotify directly for that episode. If you want, or I can or what? How do these work together? So That's kind of what we hope to do here. So we'll start off with the sort of the basic one, which is models or LLM. LLM stands for large language model, The temperature is often very low, SSE, web streams, embedding, Vector, VectorDB, a $400,000 without having to download or really do anything. Yeah.
Wes Bos
How much data you can send it and receive back is limited
Wes Bos
But then every time you have a space, that's also a token. So if you have parameters. So This is a pretty, pretty large one. Oh, no, sorry. No, there's 70,000,000,000 and also give you a kind of an idea of how much it might cost if you want to send that much data. So blonde hair, blue eyes, toolkit
Sentry can help with errors, bugs, and performance
Scott Tolinski
Use the coupon code tasty treat, all lowercase, all one word, to get 2 months for free.
Scott Tolinski
So let's get into it, Wes. Yes.
Scott Tolinski
help you with all kinds of things, maybe even get user feedback. I'm talking about a tool like Century at century.i0.
Wes Bos
So this is basically 25.
Finds textual similarities mathematically
Wes Bos
thin face, model that knows what things are. You basically show it a 100 pictures of hot dogs, and then you show it another picture. You go, what's this? Right? And it says, And if you're just using indentations, I was, like, clicking on stuff. I'm like, but what is this. Like, what like like, what is it? You know? Yeah. How do I use this? Yeah. The 1st time I got there, I wanted to get StarCoder working, and I was just like, alright. Yeah. How do I get Starcoder working? What do I have to do? You're telling me how big they are, but you're telling me I can run them here or Evals, Langchain, PyTorch, TensorFlow,
Wes Bos
like all of these different values chat gpt see
Wes Bos
you would be able to mathematically Syntax. Fm6 And once you convert,
Wes Bos
how much those questions overlap. Are they similar questions. Are they close to each other? if it's really big, then you go for something like AWS as SageMaker And I did find that it took me a little bit more work to get
Wes Bos
photos of that person. It's because it understands
Wes Bos
these mathematical equations.
Prompts prime the model, end with assistant colon
Scott Tolinski
the AI? So that it responds in the in the ways that you want it to. Because it it is funny because
Wes Bos
model if you are sending text and Those are some things, but that is because I I could just say which will That's neat. Next we have is just a bunch of, And then it drives the fill in the blank for you via your prompt. And a lot of people are talking about prompt engineering, which is essentially like, How do So I'm sure you can buy your your specific prompt, not Fetch, but Axios under the hood. So you get the big Axios response, to There's some interesting summarizer And the AI understands tons same thing with images. That's how Google Lens works. Right? You search for to all models. But most of these models will allow you to pass in some sort of value, especially the ones that you are using via an API,
Models vary in speed, price, size and quality
Wes Bos
And it's all again, it's a trade off between speed, or that they're telling us.
Anthropic took more work but provided better results
Wes Bos
it to do what I want it to do. But On the flip side, once I did figure it out, the given the types of content that was talked about in this episode.
Wes Bos
questions or whatever. Yeah. Actually, when I switched over the syntax just a photo of a hot dog on my phone,
Wes Bos
to Anthropic Claude, ton of datasets, but one of the most popular datasets is called Truthful QA.
Context windows growing larger and cheaper
Wes Bos
6 little clips from from a couple podcasts. Yeah.
Wes Bos
transcripts for 2 podcasts instead of
Correct and incorrect answers provided to train models
Wes Bos
which is watermelons grow in your stomach.
Past episode with Chris Lattner explains more on AI
Wes Bos
something that understands Tokenizer of this podcast is about 16,000
Wes Bos
the models have been trained on a bunch of data. The very basic example is, hundreds and hundreds of models out there that are trained on doing things like image creation or a captions file, and it's able to process that, This is in server equipment to even possibly run it special CPUs, things like that, If you have it lower, then it will generate things that are very similar every single return. If you have it higher, it's going to be a lot more hopefully you leave this with So Langchain is a
Wes Bos
What these things are? Obviously, it's a lot more complicated than that, but at the basis, a model what it is. So At a very high level example is Or I have all these podcast episodes. Which one is the best at And I was like,
Wes Bos
You can just search 679 and if we missed anything on this list. Yeah.
Spaces are Hugging Face model playgrounds
Wes Bos
in Hugging Face Unpredictable, it's a little bit annoying. I'm sure OpenAI will update it at some point, but