Extracting comprehensible speech from audio soup

Hi. I’m hoping someone can help me with an audio query.

Although I’ve been dealing with web-related content for many years, I freely confess that I’m a total novice when it comes to treating sound files because it has never been part of my work process.

However, I now need to optimise a sound file and I’m struggling with various filters in Audacity and getting absolutely nowhere because I simply don’t know enough to know what I don’t know!

Background: My son attended a consultation with a child psychologist and I need to make a transcript of that meeting. I was given a cassette tape of it by the psychologist but either the tape was poor, or the recording equipment was, or the microphone was… or maybe everything was poor. The end result is that the voices sound like they are underwater, with tons of background noise and would could be reverb or distortion. It’s very difficult to make out what people are saying. And it runs about 35 minutes.

Anyway, I digitised the recording and tried to clean it with various filters but got nowhere. I then sampled two 15 second clips, uploaded them on various (free) AI sites and tried to see if online AI could clean them, but they are removing so much of the background noise and reverb (?) that they are clipping the speech and making half of it indecipherable. And the processing is automatic, with no options for me to customise it, even if I knew what changes to make… which I don’t.

Can anyone recommend any software that does a great job of extracting speech from audio soup? Or the best filters to use with Audacity? Or an AI site that could help?

Thanks in advance for any suggestions.

Regards

I’m assuming you digitised the tape but you don’t say how you connected things up to do so. Also, I’m unclear how you don’t know if the tape is bad. Didn’t you listen to it before digitising it?

You’re quite right, there are lots of AI sites that promise to clean up audio. The few I’ve tried didn’t live up to their claims TBH. You could try using Ocenaudio rather than Audacity. It has some speech enhancement AI plugins and it’s free too. Also, the preview of effects (like EQ, for example) is nicer to work with than the Audacity offering.
If the content wasn’t of such a deeply personal nature I’d offer to have a go myself but I realise it’s not something to share.
I hope you find an effective solution.

Hi MrB

Thanks for coming back to me on this. In answer to the above, I first listened to the cassette tape through a couple of tape players and was shocked by the awful quality. Equalisers, etc had little positive effect. When I played it on an old Walkman I could make out a little more of the dialogue.

In an attempt to try and treat the recording myself, I digitised it twice (via LINE OUT and LINE IN connections), once through a PC soundcard and then via a cheap USB Audio capture device (DIGITNOW). The two digitised versions sound indistinguishable to me, and both are faithful copies of the original cassette tape (the awful audio is the same on all).

Thanks very much for the suggestion of OceanAudio. I’ll go over and try it this afternoon.

Obviously, the content of the full interview is personal but if someone were able to point me at the right or optimal filter setup I’d be happy to make a couple of short vanilla clip samples that contain no personal data.

I’ll have a go with OceanAudio. Thanks again!

In my experience, if you can’t decipher what is being said neither can AI. Superhuman AI is not here, yet.

Treble Boost could help intelligibility … Bass and Treble - Audacity Manual

Thanks, Trebor! I’ll try the Treble Boost to see if I can make any improvement.

Point taken about AI not being superhuman. The results I got from a couple of free tests were clearly aimed in the right direction but went too far. In cutting out the reverb (? - I can’t be sure I’m labelling that correctly) it produced a clear background but clipped half of the speech sounds. I’d settle for something still somewhat muddy and distorted but understandable. With that, I could at least type up a transcript.

Kind regards

That’s not natural distortion or damage. Background noise, hiss, low volume, vocal competition, or overload distortion are all natural problems and it’s possible to find solutions although you may not be able to get theatrically perfect work.

“Underwater” or gargling usually means the work has been through processing or corrections already. Now it’s roll up your sleeves. You have to identify the type of correction, remove it specifically and then go on to correct any other more normal damage.

If you can’t do that and you can’t understand the words now, then this doesn’t look good.

How are you listening? Do you have top quality stereo headphones available? You can sometimes dig conversations out of the mud if damage is “flat” and different people are coming from different directions.

Koz

First, apologies for the delay in replying to these super helpful posts above. I have to fit this audio task around work and parenting so I have to do it in available space and that means I’m sometimes delayed.

Hi Koz

Thanks for your thoughts. I don’t have a super high quality pair of headphones (I’m just using ear buds or mid-range phones) but it’s certainly the case that I can catch a few more words if I use headphones rather than an external speaker. I don’t have a hi-fi setup here, and don’t need the recording to be crystal clear or perfect in any way. I’d settle for just being able to make out what is being said, even if it is at low quality.

The recording hasn’t been through any processing already. It’s an old (probably low quality) cassette tape. I simply digitised it in the hope of being able to optimise the sound file. At present, the original tape and the digitised copy are the same in terms of poor quality sound.

I think the recording issues were probably compounded by poor quality equipment during the consultation (including microphone), vibration on the desk, echoes and distance from microphone.

Mainly it’s two people talking, and not in competition. My son’s high voice (he was 8 at the time) and childish diction make him a little harder to understand than the measured tones of the much older doctor.

You have to identify the type of correction,

Absolutely, and that’s really the heart of the problem. I look at filters and various settings, and because I can’t identify the actual problem I’m simply guessing with the solution. If I had some handle on the type of issue I could hopefully experiment with trying to remove it with an appropriate filter (or filters). That’s why I was hoping that AI could magically substitute for my ignorance and pluck the words out. :o))) I only need to be able to understand the words enough to be able to type them up into a transcript.

I’m going to be experimenting with the suggestions from Mr B and Trebor this afternoon and will report back.

Thanks again for your help.

Kind regards

First, let me thank you all again for taking the time to offer advice and suggestions. I was just able to get time to experiment and can report the following:

  1. MrB - The speech enhancement plugin on OcenAudio was a slight improvement on the online AI tools but still clipped too much of the speech. But i could make out a few more words but not enough. I tried it several times on different clips but had to stop because the program froze each time after a few applications of the plugin and I had to keep restarting it. Obviously a conflict somewhere. Based on that test, I don’t think AI is going to help me as much as I’d hoped.

  2. Trebor - The Treble Boost definitely improved intelligibility! I could understand quite a few more words. I tried reducing the bass, as well, hoping that might take away some of the reverb/vibration, and that also seemed to improve things but only a tiny bit. I tried all the permutations I could think of but ran out of ideas. I can’t help feeling that combining Treble/Bass with other filters might be the way to a solution. I’m just not sure what would be the appropriate/optimal combination. I don’t need the recordings to be hi-fi quality, just intelligible enough to understand the words, even if there is background hum or other artefacts.

If anyone has any suggestions to improve things further I’d be grateful.

Kind regards