AI-powered captions are becoming increasingly common, but Rylo CEO and co-founder Tomer Aharoni believes accessible communication needs to go much further than simply converting speech into text.
Shari Eberts speaks with Aharoni about the evolution of Nagish into Rylo and the company’s expansion from FCC-certified captioned phone calls into a broader AI communication platform for people who are deaf and hard of hearing. Aharoni shares how the company began with a simple experiment while he was a computer science student at Columbia and why Rylo now views accessibility as the starting point rather than the end goal. The conversation explores how AI could provide communication context that traditional captions may miss, including emotion, tone, accents and environmental sounds.
Aharoni also discusses emerging tools such as sentiment analysis, potential fraud detection features, and Rylo Sign, an effort to use AI to translate between signed and spoken language when an interpreter may not be available.
Looking ahead, Aharoni explains why he expects basic captioning to become increasingly ubiquitous and how AI could help reduce the additional planning and communication barriers that people who are Deaf and hard of hearing often encounter in everyday life.
Full Episode Transcript
Welcome to This Week in Hearing. I’m Shari Eberts. Many of you know Nagish, the FCC-certified app providing free AI-powered real-time captioning of phone calls with no human operator on the line. But recently, the company rebranded to Rylo and announced $85 million in new funding to expand expand from captioned phone calls into a broader AI communication platform for people who are deaf and hard of hearing. My guest today is Tomer Aharoni, Rylo’s CEO and co-founder, and he started his journey in 2019 when he was a Columbia computer science student and he hacked his laptop to caption a phone call in a Google Doc. So Tomer, welcome to the show.
Hi, Shari. Thank you for having me. Absolutely. So take us back to that computer science class in 2019. Your phone rang, you couldn’t answer it. How did you turn that into an FCC-certified company? Yeah. So back in 2019, I was sitting in class. I was a computer science student at Columbia. I got a call. I couldn’t pick it up. And the call was followed by a text message saying it was urgent, and my brain started racing. What is it? Do I have to go out of class and answer this immediately? And then I thought, what if I had a way to read my phone calls instead of having to hear them?
So I’d be able to answer a call here during class in a situation where I cannot speak or hear. On the spot, I like— just like you said, I hacked my audio card of the computer So instead of the sound going out of the speakers, it went back into the computer. And then I just opened the Google Doc, started dictating, and got a text of everything that was said in real time. It wasn’t really enough to have a phone call, but it gave me a sense of the situation, which, by the way, wasn’t that urgent. Of course not, right? Yeah. But when I got out of class, I mentioned this to Alon.
Alon is my co-founder. He’s our CTO, and his first reaction was, “What are you talking about? Like, how do deaf people communicate? I’m sure they have a solution.” And that’s how we learned that the deaf community didn’t have a solution. They were relying on humans, namely interpreters, stenographers, captioners that would be an intermediary between them and another hearing person. And to us, being two computer science students, it just felt absolutely insane that the year was 2019 and people can’t have private calls. So we built a proof of concept, had no intention whatsoever to make it a company, but that proof of concept just gained more and more and more traction all the way to the point where we almost had no choice but to make it a company.
That’s a pretty cool origin story. I like that very much. Did you get in trouble from your professor for hacking things in class? Class, or— I mean, it was a computer science class, so I think they, they mostly appreciate it. It was expected. Okay, so that’s good, that’s good. So Nagish means accessible in Hebrew, is that correct? Yes, that’s right. Yeah, so that’s a name that has, you know, a lot of important meaning. Can you talk about why you’re changing the company’s name and why you chose Rilo? Yeah, when we started, we used the first word that came to mind. We’re both regionally Israeli. We spent more than a decade now in the States, but it’s just like we wanted to make calls accessible.
The first word that came to mind was accessible, and in Hebrew it means “nagesh.” But two things happened. One, as we kept building, realized that accessibility is really not the ceiling we’re aiming for. It’s just the floor. We want to do way more than just accessible. We want to do better. We want to change the way people communicate. So the word accessible suddenly felt just small. And the other thing is the word, the quiche wasn’t so accessible because many people didn’t know how to pronounce it. And eventually we’re only operating in the US. We do hope to be able to expand beyond the US, but right now that’s our main market.
We needed a word that people can rely on and actually pronounce. And Rylo is that word. It’s, it’s a play on Relay, which is the type of services that we provide, and rely on the ability to rely on these Relay services. Okay, there we go. And it is easy to say, so that’s good. And so you’ve just said, so accessibility is the floor, not the ceiling. I like that idea. And historically, I guess Relay services were, were maybe thought of as something that people had to use, right? Not something that they necessarily wanted to use. So can you talk about, you know, how that philosophy really impacted your product design and the mission overall to really elevate it from, you know, not just being a floor, but taking it to the, to the ceiling.
Yeah. So if you think about legacy relay services all had a good mission, but eventually they weren’t the type of experience their users wanted to have. They actually were the thing that stood between the end user and the person they wanted to talk to. So many people think of them as a bridge, where in fact they were at the same time a bridge and a friction point. And we said, what if we can create a relay service that is only a bridge, that is really just there to remove any friction, to make life truly accessible— or not even accessible, I’m really trying to avoid that word— because just equivalent, just equivalent to what hearing people get on a call.
When I call someone as a hearing person, I don’t think of whether they’re going to hang up on me, on whether they’re going to accept my call. I don’t have this level of anxiety that the deaf person have when they call someone, and that’s what we wanted to remove. That makes a lot of sense. So let’s just talk about the options that are out there, right? So there are a number of options for the hard of hearing community now in terms of captioning apps, and I know you’ve gone all in on AI. So can you talk about why AI, and you know, what situations where is that better than a human captioner?
And then where, you know, does the AI still have work to do? Yeah, so I’ll start from the end. I think it’s only better in situations where the end user finds it as better. Okay. We always try to avoid making any claims on our AI is more accurate than a human, because what does accurate mean, right? If I look at word error rate, like how many errors does a sentence have when I caption it, then yes, on paper it’s more accurate. But if the sentence is, “I’m not allergic to penicillin,” and the AI is saying, “I am allergic to penicillin,” it doesn’t matter that it only missed one word.
That word was crucial. That said, we rarely make mistakes like that. But the bottom line is we said we’re going to develop the best technology out there and if it’s good enough, people are going to use it. We don’t need to worry about people using it or whether it’s accurate enough. If people are willing to use it and if people prefer it over relying on another human, we did our thing. But again, that’s just the floor. That’s accessibility. Captioning is just the floor. We want to do more than just captioning. We want to make sure that your audio quality is best in class, that is streamed to your hearing aids in real time.
We want to make sure that we caption or give you the information on the call that is not necessarily verbal. I have an accent. If I talk to a hearing person, they know that I have an accent. If I talk to a deaf person, they don’t know that I have an accent. They deserve to know. Where am I sitting right now? Do I have background noise around me? Is it noisy? It’s just giving people exactly the same thing that everyone else gets on a call, That’s what we see as the level of service that we want to provide. Interesting. Okay, so that kind of flows right into this newly announced toolkit that you announced, which includes the sentiment analysis and some real-time fraud detection as well.
They sounded very interesting to me because a lot of people are getting more fraudulent phone calls, unfortunately. So maybe you can give us a real-world example where one of those tools would change the outcome of a conversation. Conversation in a way that captions could not do alone. Absolutely. So I’ll give you an example. We have a lot of users who tend to be older, 80+, and unfortunately, we see it all the time, these users are more prone to fraud. So for instance, they can get a call from someone claiming to be, say, American Express, and that person on the line can tell them, hey, this is ‘X from American Express.
I need to verify your credit card number. Can you please read it out to me?’ And they would just go and read it out to them, and here a scammer just gained access to their credit card. That happens to everyone regardless of whether they use Relay or not. But we as an AI-powered telephony provider actually have the capability to privately detect that a speaker says I’m from American Express, and at the same time to detect that that phone number is not from American Express and just show the user a message, this caller is not who they claim to be. It’s, it’s not for us to take any decisions on behalf of the user, but we can at least flag it.
It’s not something that we offer today, but it’s just one idea of fraud prevention that we could easily integrate into our services. Okay, but so this is sort of something down the line. This is not something that exists today. Yeah, the bottom line is that we want to give our users the utilities that they deserve, make sure that they get to communicate in the most humane way, and that they don’t need to worry about how they communicate, where they communicate, and whether they communicate to the person they’re thinking they’re communicating with. Okay, great. So is it a similar scenario with the sentiment analysis? Because I know sometimes it is very hard for people with hearing loss, myself included, to, judge the emotional content behind what’s being said?
So maybe we don’t get the sarcasm or what have you. How would that sentiment analysis work? So just like you imagine, you’d be talking to someone. And I mentioned that we want to do more than just captioning, right? So captioning only gives you part of the context that takes place on the call. But we want you to know, is the person on the other end angry right now? Are they happy? Are they laughing? What do they mean? And this is where sentiment analysis comes in. So every single call on Rylo has this set of AI layers that actually provides more context than just captions, whether it’s sentiment, whether it’s accents, whether it’s the sound, like the gender of the tone of the person who’s speaking.
Is it a feminine voice or a masculine voice? So we just want to give our users as much context as possible. OK, and so that’s something that’s part of the product today that people would experience. It’s also— it’s rolling out right now. So all these updates, you’re actually hearing them first. They’re rolling out to a small group of users. We’re testing with the community, and over time we expand and eventually launch for everyone. That’s terrific. I know Google is doing some work on that as well. They have these expressive captions. So is that information going to be in that way? So like the, you know, the— letters would look different, or would there be sort of like in parentheses, oh, this person is shouting, or how would that work?
So it really depends. Some users don’t want their captions to be modified, and some users— it’s actually a really tough product decision. For instance, if a person is sneezing, do I want to know? Like, do I really care? So the question is, what am I putting on the user interface? How? Is it part of the captions? Is it a small indication that something happened? That’s why we’re also not releasing it overnight, because we want to see how people react and what’s their preferred form. Yeah, no, interesting. Okay, well, good. Lots of options is always a good thing for a user, so I love that. Yes.
So the toolkit, I guess, also includes Rylo Sign, which provides AI translation between signed and spoken language, and there’s definitely a shortage of certified ASL interpreters in the US. So that is, you know, a big need for a service like this. So how is your Sign product going to close that gap? And how do you think about it? Is it a replacement for interpreters or a supplement or, you know, what exactly is it? So just like with captioning, it’s not a replacement. It really depends on what the user wants to use. Sign language is not just another language that you caption. It’s a very rich, expressive, and yet at the same time abstract language.
And it’s also very unique. It develops differently in different communities. People in different states have different signs for the same word, not to mention all the name signs that are just impossible to grab. What we’re trying to do is we’re trying to develop models that would make sign language slightly more accessible. So if I’m at a situation where it’s just impossible for me to get an interpreter, at least I have a fallback. That fallback can be super accurate. But we’re not making any claims on this is going to be 100% accurate, or this is going to replace sign language interpreters. We’re not there. And we’re not trying to get there.
We’re trying to just give another tool to that community. So they just have a bigger set of tools to choose from. So maybe it sounds like it’s kind of the early days of captioning when it was, if you couldn’t get a human captioner there and there was just nothing you could do about it, you’d say, well, we’ll try the AI and, you know, no guarantees. And obviously that’s improved and improved, you know, over time. It’s not perfect. Nothing’s perfect, but it’s improved quite a bit. And so maybe this is just sort of that individ— like that beginning piece with the sign language. Is that a right way to think about it?
Exactly. So it’s really about progression over perfection from day one to begin with. And then the other aspect here, unfortunately, or also fortunately, captioning is at the interest of almost every single company. There’s so many applications to captions that the biggest giants in the world, the Googles, the OpenAIs, the Apples of the world, they all work on captioning engines. And what happens is that we as a community benefit from these developments. We get the latest tech and the best tech. But with sign language, no one out there is really working on that. I mean, there are a few research teams, but it’s on a very low scale.
And we just decided we never wanted to be a captioning company. We always wanted to be an accessibility company. We were lucky enough that captioning advanced so much that we didn’t have to develop our own models over time. But with sign language, we just couldn’t wait for Google or Apple to decide that they want to develop their own models. So we almost had no choice but to do it ourselves. That makes sense. I mean, and speaking of these big companies, Apple and Google, like you said, a lot of them are sort of integrating this captioning now directly into their operating systems, into FaceTime or what have you.
I mean, do you see a time down the road where a separate app might not be as necessary as it was in the past? Potentially, I think captioning is fully commoditized eventually. It’s very easy to get captions today, and that’s exactly why we’re doing more than captions. Eventually, when you think about these companies, about Apple and Google, they sometimes do things like captioning out of goodwill, but for them, it’s really just a checkmark. It’s, we build captions, we’re done. For us, it’s all we do. It’s our product. It’s, again, not about captions. It’s about speech-to-text and text-to-speech and speech-to-speech and sign language and making sure that everyone is heard and everyone is understood in the best possible way for their needs.
So even if that becomes the case, I think it will be a net benefit for everyone, including us as a company providing accessibility services. Well, that makes sense. All right, so let’s zoom out for a minute, right? Because AI is, you know, pretty much everywhere now in healthcare and hearing aids and earbuds, captioning, or cast, all of this. So what’s sort of your best-case scenario, right? If you look out 5 years from now, what communication barriers are going to be gone and how do you guys fit into that changing and new landscape? We want to live in a world where there is no partitioning, where a deaf person gets to do exactly the things that a hearing person gets to do.
They get the same access to creativity, to education, to the workplace. They get hired at the same rate. They get into the same Ivy League schools that hearing people get into. That’s what we want to achieve. A deaf person should be able to walk into a yoga class and just understand what, what the instructor is saying. They don’t have to live their life with another layer of planning and anxiety. Is this party going to be accessible? Is this event going to have live captions? Can I go into this job interview knowing that the interpreter will be there? That’s all we’re trying to do. And there are 3 ways to do it, or 3 layers to think about it.
It’s in person, over the phone, and at the workplace. And that’s what we’re doing. Perfect. All right. So how can people who are listening to us today learn more about this and how can they give the new toolkit a try? So go on Rylo.com. That’s R-Y-L-O.com. Check out our products. We don’t charge for anything. Accessibility should be free. So it’s 100% free. Our phone call solution is certified by the FCC, which means that customers don’t, as long as they have hearing loss, don’t have to pay for it. So just go on Rylo.com. Excellent. All right. Well, thank you so much for spending the time talking about this today.
It’s very exciting, really, to see so much innovation in accessibility for people who are deaf and also hard of hearing. So thanks so much for being here. Appreciate it. Of course. Thank you for having me. It was a pleasure.
.
Be sure to subscribe to the TWIH YouTube channel for the latest episodes each week, and follow This Week in Hearing on LinkedIn, Instagram and X.
Prefer to listen on the go? Tune into the TWIH Podcast on your favorite podcast streaming service, including Apple, Spotify, Google and more.
About the Panel
Tomer Aharoni is the co-founder and CEO of Rylo, formerly Nagish, an AI-powered communication platform for people who are deaf and hard of hearing. He co-founded the company after developing an early phone-captioning prototype while studying computer science at Columbia University, eventually growing the concept into an FCC-certified service and broader accessibility platform.
Shari Eberts is a passionate hearing health advocate and internationally recognized author and speaker on hearing loss issues. She is the founder of Living with Hearing Loss, a popular blog and online community for people with hearing loss, and an executive producer of We Hear You, an award-winning documentary about the hearing loss experience. Her book, Hear & Beyond: Live Skillfully with Hearing Loss, (co-authored with Gael Hannan) is the ultimate survival guide to living well with hearing loss. Shari has an adult-onset genetic hearing loss and hopes that by sharing her story, she will help others to live more peacefully with their own hearing issues. Connect with Shari: Blog, Facebook, LinkedIn, Twitter.







