Subtitles to transcript
You have the subtitle file of a meeting or a video and you need the text, not the numbers. Load your .srt or .vtt and out comes prose: times gone, sentences rebuilt, and above all the repetitions of scrolling captions removed. It takes a file you already have: nothing is downloaded from any site here.
The real problem is not the times
Anyone can strip numbers and timecodes. The trouble with automatic subtitles is another one: they scroll, meaning every box repeats the tail of the one before to push it up, and the text that comes out says every sentence two or three times. A file like that, handed to someone or something that has to summarise it, is three times longer than it needs to be and full of echoes.
I only have the audio, not the subtitles
It happens to half the people who land here, and the honest answer is that the subtitles almost always exist somewhere already. If a video-call program recorded the meeting, there is almost always a transcript file to download next to the recording. If the video is yours and sits on a video platform, the automatic captions can be taken from the uploader’s panel. If all you really have is an audio file, the most common writing programs have a transcription feature, and the phone keyboard has dictation. There is no speech recognition in here, and that is not an oversight: doing it in the browser means downloading tens of megabytes on every visit, and the models that would fit get about one word in four wrong on spontaneous speech, with mistakes that are real words and therefore invisible to whoever proofreads. Once you have the subtitles, this page does the rest; if the times are off go through Sync subtitles, and to find out whether they can actually be read there is Check subtitles.
How it decides whether to cut
First it measures the whole file: it looks at how many consecutive pairs of boxes overlap in time and genuinely repeat the same words. If they are the majority, those are scrolling captions and it merges them. If they are two or three, those are real repetitions by the speaker, and it touches nothing: so a «no, no, no» stays «no, no, no». The comparison is exact, word by word, never by similarity.
What it did is written down
Under the result you find how many captions it read, how many repeated tails it merged and how many words it removed. The «See the raw text» button shows the captions one per line, with no merging: it is there so you can check, in ten seconds, that nothing went missing.
Speaker names
The names are only there if the platform wrote them into the file: that happens with Teams, Zoom and Meet, and with .vtt files that use <v Name>. YouTube's automatic subtitles do not have them, and in that case it tells you: working out who is speaking from the audio alone cannot be done, and inventing it would be worse than saying nothing. If the names are there, you can also replace them with Speaker 1, 2, 3, which is a way of anonymising them, not of recognising them.
Punctuation and paragraphs
Paragraphs close on a full stop, on a change of voice or on a pause (you set it). If the subtitles have no punctuation, and automatic transcripts often do not, only the pauses are left: it tells you, because punctuation cannot be invented. And words the automatic transcript misheard stay wrong: what is fixed here is the shape, not the content.