Title: Collaborative Steering of Microphone Array and Video Camera Toward Multi-Lingual Tele-Conference Through Speech-to-Speech Translation
Authors: Takanobu Nishiura, Rainer Gruhn, Satoshi Nakamura
Abstract:
It is very important for multi-lingual tele-conferencing through speech-to-speech translation to capture distant-talking speech with high quality. In addition, the speaker image is also needed to realize a natural communication in a such conference. A microphone array is an ideal candidate for capturing distant-talking speech. Uttered speech can be enhanced and speaker images can be captured by steering a microphone array and a video camera in the speaker direction. However, to realize automatic steering, it is necessary to localize the talker. To overcome this problem, we propose collaborative steering of the microphone array and the video camera in real-time for a multi-lingual tele-conference through speech-to-speech translation. We conducted experiments in a real room environment. The speaker localization rate was 97.7%, speech recognition rate was 90.0%, and TOEIC score was 530-540 points, subject to locating the speaker at a 2.0 meter distance from the microphone array.
|