An optimization method and system for sound source identification

By collecting videos and converting them into pinyin data, a neural network model is established, and combining handheld device cameras and microphones to identify pinyin syllables and match face positions, the problem of high cost of microphone arrays is solved and accurate sound source recognition on low-cost devices is achieved.

CN115910068BActive Publication Date: 2025-07-04FUJIAN TQ ONLINE INTERACTIVE INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211304430.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-24
Publication Date
2025-07-04
Estimated Expiration
2042-10-24

AI Technical Summary

Technical Problem

In the prior art, the microphone array cost is high or the single microphone device cannot accurately identify the sound source, resulting in high cost and poor effect of sound source recognition.

Method used

By collecting videos and converting them into pinyin data, a neural network model is established, and a handheld device camera and microphone can identify pinyin syllables and match face positions to achieve sound source recognition.

Benefits of technology

Accurately identifying the location of speakers in a conference on devices without microphone arrays reduces identification costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115910068B_ABST
    Figure CN115910068B_ABST
Patent Text Reader

Abstract

The present invention provides an optimized method for sound source recognition. The method includes: Step S1, collecting videos of different people speaking, and invoking speech recognition to capture text, thereby forming sample data of multiple people; Step S2, by establishing a neural network model, inputting the sample data into the neural network model to obtain optimal initial consonants, final vowels, and complete pinyin data, and forming a neural network model parameter group 1; Step S3, using the camera of a handheld device to photograph the faces of each member in the meeting, obtaining each person's face sequence through opencv, and substituting the face sequence into the neural network model parameter group 1 through the neural network model to output corresponding pinyin syllables; Step S4, obtaining the current audio through a microphone to get predicted pinyin syllables, comparing the predicted pinyin syllables with the pinyin syllables obtained from the face, and matching the corresponding face, then the current sound source is the position of the face; reducing the cost of sound source recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sound source recognition, and particularly to an optimized method and system for sound source recognition. Background Art

[0002] There are many methods for noise source identification, and one or several reasonable methods should be adopted according to the actual object and conditions during application. The development of noise source identification technology is closely linked to the progress of noise measurement technology. With the emergence and development of digital signal processing and computer technology, noise source identification technology has made great progress in the past few decades. New identification technologies and instrument devices have emerged continuously. Currently, sound source identification technology locates the direction of the sound source through a microphone array. However, the cost of the microphone array is high, or many devices themselves only have a single microphone, making it impossible to determine the sound source. Summary of the Invention

[0003] To overcome the above problems, the purpose of the present invention is to provide an optimized method for sound source recognition that can accurately identify the position of the speaker in a meeting.

[0004] The present invention is implemented as follows: An optimized method for sound source recognition, the method comprising the following steps:

[0005] Step S1: Collect videos of different people speaking, and call speech recognition to capture text. Convert the text into initials, finals, and complete pinyin, and mark the corresponding relationship between consecutive frames of pictures and complete pinyin in the video to form sample data of multiple people.

[0006] Step S2: By establishing a neural network model, input the sample data into the neural network model to obtain the optimal initials, finals, and complete pinyin data, and form a neural network model parameter group 1.

[0007] Step S3: Use the camera of a handheld device to capture the faces of each member in the meeting, obtain each person's face sequence through opencv, and substitute the face sequence into the neural network model parameter group 1 through the neural network model to output the corresponding pinyin syllables.

[0008] Step S4: Obtain the current audio through a microphone to get the predicted pinyin syllables, compare the predicted pinyin syllables with the pinyin syllables obtained from the face, match the corresponding face, and then the current sound source is the position of the face.

[0009] Further, in step S1, the corresponding relationship between consecutive frames of pictures and complete pinyin in the video is recorded through a file, where it is recorded that the consecutive frames of pictures from the i-th frame to the j-th frame in the video belong to a corresponding complete pinyin; where i and j are integers, j > i. After marking, adjust the brightness and color of the pictures, and finally form sample data of multiple people.

[0010] Further, the method collects a large amount of sample data, where the sample data includes face sequences and the corresponding pinyin for each face sequence; by establishing a convolutional neural network and an LSTM network, the parameters trained by the convolutional neural network for each frame of the face sequence are then input into the LSTM network, and the final result is compared with the pinyin data converted through speech recognition. The parameters in the convolutional neural network and the parameters in the LSTM network of the neural network model are optimized using the open-source deep learning framework tensorflow, that is, the neural network model parameter set 1 is obtained.

[0011] Further, in step S3, the position information of each frame of the face is obtained through opencv, and then the position information of each frame of the face is strung together according to the time series to obtain the face sequence.

[0012] The present invention provides an optimized sound source recognition system, which includes: a sample data generation module, a neural network training module, a face prediction pinyin module, and a sound source determination module;

[0013] The sample data generation module collects videos of different people speaking, and calls speech recognition to capture text, converts the text into pinyin initials, finals, and complete pinyin, and marks the corresponding relationship between consecutive frames of pictures and complete pinyin in the video to form sample data of multiple people;

[0014] The neural network training module, by establishing a neural network model, inputs the sample data into the neural network model to obtain the optimal initials, finals, and complete pinyin data, forming the neural network model parameter set 1;

[0015] The face prediction pinyin module captures the faces of each member in the meeting through the camera of a handheld device, obtains each face sequence through opencv, and substitutes the face sequence into the neural network model parameter set 1 through the neural network model to output the corresponding pinyin syllables;

[0016] The sound source determination module obtains the current audio through a microphone, gets the predicted pinyin syllables, compares the predicted pinyin syllables with the pinyin syllables obtained from the face, and matches the corresponding face, then the current sound source is the face position.

[0017] Further, in the sample data generation module, the corresponding relationship between consecutive frames of pictures and complete pinyin in the video is recorded through a file, where it is recorded that the consecutive frames of pictures from the i-th frame to the j-th frame in the video belong to a corresponding complete pinyin; where i and j are integers, j > i. After marking, the brightness and color of the pictures are adjusted, and finally the sample data of multiple people is formed.

[0018] Further, the system collects a large number of sample data, where the sample data includes face sequences and the corresponding pinyin for each face sequence; by establishing a convolutional neural network and an LSTM network, the parameters trained by the convolutional neural network for each frame of the face sequence are then input into the LSTM network, and the final result is compared with the pinyin data converted through speech recognition. The parameters in the convolutional neural network and the parameters in the LSTM network of the neural network model are optimized using the open-source deep learning framework tensorflow, that is, the neural network model parameter set 1 is obtained.

[0019] Further, in the face prediction pinyin module, the position information of each frame of the face is obtained through opencv, and then, the position information of each frame of the face is strung together according to the time series to obtain a face sequence.

[0020] The beneficial effects of the present invention are as follows: The sample data is trained through a neural network model. The face sequence is substituted into the neural network model through face recognition to output pinyin syllables. The current audio is obtained through a microphone to get the predicted pinyin syllables. The predicted pinyin syllables are compared with the pinyin syllables obtained from the face. If the corresponding face is matched, then the current sound source is the position of the face. In this way, in a device without a microphone array, the position of the speaker in the meeting can also be accurately identified, reducing the recognition cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a schematic flowchart of the method of the present invention.

[0022] Figure 2 is a schematic block diagram of the system principle of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] The present invention will be further described below with reference to the accompanying drawings.

[0024] Please refer to Figure 1 shown, an optimized method for sound source recognition of the present invention, the method includes the following steps:

[0025] Step S1, collect videos of different people speaking, and call speech recognition to capture text, convert the text into pinyin initials, finals, and complete pinyin, and mark the corresponding relationship between the continuous frames of the pictures in the video and the complete pinyin to form sample data of multiple people;

[0026] Step S2, by establishing a neural network model, input the sample data into the neural network model to obtain the optimal initials, finals, and complete pinyin data, forming the neural network model parameter set 1;

[0027] Step S3: Use the camera of the handheld device to capture the faces of all meeting members, obtain each face sequence through OpenCV, and substitute the face sequence into the neural network model parameter group 1 through the neural network model to output the corresponding pinyin syllables.

[0028] Step S4: Obtain the current audio through the microphone to get the predicted pinyin syllables. Compare the predicted pinyin syllables with the pinyin syllables obtained from the face. If the corresponding face is matched, then the current sound source is the position of the face.

[0029] The following further illustrates the present invention with a specific embodiment:

[0030] 1. By calling the existing speech recognition, capture the text from the speech, and then convert the text into the initials and finals of pinyin, and the key frame pictures of the complete pinyin. Mark the corresponding relationship between the consecutive frames of the pictures and the pinyin. Generate more samples by adjusting the brightness and color of the pictures.

[0031] For example: Call the iFlytek API to convert the speech of "Hello" into "Hello", and then convert it into the pinyin "ni hao".

[0032] Through the existing technology of face detection, extract the face sequence corresponding to "Hello" in the sample video. Perform data augmentation on the face sequence (change the brightness and color of each frame of the picture, and randomly adjust the speed of the face sequence), and the corresponding pinyin is "ni hao". Among them, a file is used to record the corresponding relationship between the consecutive frames of the pictures in the video and the complete pinyin. It records that the consecutive frames of the pictures from the i-th frame to the j-th frame in the video belong to a corresponding complete pinyin. Among them, i and j are integers, and j > i. After marking, adjust the brightness and color of the pictures.

[0033] Use a large amount of data marked as above to form sample data.

[0034] 2. By establishing a neural network model, input the sample data in step 1 into the neural network model to output the corresponding initials, finals, and complete pinyin. Finally, obtain the optimal neural network model parameter group 1; the technology for processing this neural network model is already existing technology. For reference, see the technology with the application number: 202011471841.2 and the patent name: Speech Synthesis Method, System, Device and Storage Medium Based on Neural Network.

[0035] The method collects a large number of sample data, where the sample data includes face sequences and the corresponding pinyin for each face sequence; by establishing a convolutional neural network and an LSTM network, the parameters trained by the convolutional neural network for each frame of the face sequence are input into the LSTM network, and the final result is compared with the pinyin data converted by speech recognition, and the parameters in the convolutional neural network and the parameters in the LSTM network of the neural network model are optimized using the open-source deep learning framework tensorflow, that is, the neural network model parameter set 1 is obtained.

[0036] 3. Use the mobile phone camera to capture the faces of each member in the meeting, obtain each face sequence through opencv, and substitute the face sequence into the neural network model parameter set 1 through the neural network model to output possible pinyin syllables.

[0037] Identify the face that is currently saying "hello", substitute it into the trained model, and obtain the most likely syllable "nihao".

[0038] 4. For the current audio obtained through the microphone, predict the pinyin syllables (wherein, through the model trained by oneself or directly call iFlytek's speech-to-text, and then convert the text into pinyin to obtain the predicted pinyin syllables), compare the pinyin syllables predicted by the audio with the pinyin syllables predicted by the face, and match the closest face (for example, face 1), then the current sound source is at the position of face 1.

[0039] Make predictions on several syllables predicted from the face sequence. For example, if the probability recognized by speech recognition is also 90% for "nihao", then this voice most likely corresponds to the person of this face sequence.

[0040] Please refer to Figure 2 As shown, an optimized sound source recognition system of the present invention, the system includes: a sample data generation module, a neural network training module, a face prediction pinyin module, and a sound source determination module;

[0041] The sample data generation module collects videos when different people speak, calls speech recognition to capture text, converts the text into pinyin initials, finals, and complete pinyin, and marks the corresponding relationship between consecutive frames of pictures and complete pinyin in the video to form sample data of multiple people;

[0042] The neural network training module, by establishing a neural network model, inputs the sample data into the neural network model, obtains the optimal initials, finals, and complete pinyin data, and forms the neural network model parameter set 1;

[0043] The face prediction pinyin module captures the faces of each member in the meeting through the camera of the handheld device, obtains each face sequence through OpenCV, and substitutes the face sequence into the neural network model parameter group 1 through the neural network model to output the corresponding pinyin syllables.

[0044] The sound source determination module obtains the current audio through the microphone, gets the predicted pinyin syllables, compares the predicted pinyin syllables with the pinyin syllables obtained from the face, and matches the corresponding face, then the current sound source is the face position.

[0045] In the sample data generation module, the corresponding relationship between the continuous frames of the pictures and the complete pinyin in the video is recorded through a file, where it is recorded that the continuous frames of the pictures from the i-th frame to the j-th frame in the video belong to a corresponding complete pinyin; where i and j are integers, j > i. After marking, the brightness and color of the pictures are adjusted, and finally the sample data of multiple people is formed.

[0046] The system collects a large amount of sample data, and the sample data includes face sequences and the pinyin corresponding to each face sequence; by establishing a convolutional neural network and an LSTM network, the parameters trained by the convolutional neural network for each frame of the face sequence are then passed into the LSTM network, and the final result is compared with the pinyin data converted through speech recognition. The parameters in the convolutional neural network and the parameters in the LSTM network of the neural network model are taken to the optimal values using the open-source deep learning framework TensorFlow, that is, the neural network model parameter group 1 is obtained.

[0047] Among them, in the face prediction pinyin module, the position information of each frame of the face is obtained through OpenCV, and then the position information of each frame of the face is strung together according to the time sequence to obtain the face sequence.

[0048] In short, the present invention trains the sample data through the neural network model, substitutes the face sequence into the neural network model through face recognition to output pinyin syllables, obtains the current audio through the microphone to get the predicted pinyin syllables, compares the predicted pinyin syllables with the pinyin syllables obtained from the face, and matches the corresponding face, then the current sound source is the face position; in this way, in a device without a microphone array, the position of the speaker in the meeting can also be accurately identified; the recognition cost is reduced.

[0049] The above are only the preferred embodiments of the present invention, and all equivalent changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope covered by the present invention.

Claims

1. An optimized method for sound source identification, characterized in that: The method includes the following steps: Step S1: Collect videos of different people speaking, call speech recognition to capture text, convert the text into initials, finals, and complete pinyin, and mark the correspondence between consecutive frames of pictures in the video and the complete pinyin to form sample data of multiple people; Step S2: By establishing a neural network model, input the sample data into the neural network model to obtain the optimal initials, finals, and complete pinyin data, and form neural network model parameter group 1; Step S3: Use the camera of a handheld device to capture the faces of each member in the meeting, obtain each person's face sequence through opencv, and substitute the face sequence into neural network model parameter group 1 through the neural network model to output the corresponding pinyin syllables; Step S4: Obtain the current audio through a microphone to get the predicted pinyin syllables, compare the predicted pinyin syllables with the pinyin syllables obtained from the face, match the corresponding face, and then the current sound source is the face position.

2. The optimized method for sound source recognition according to claim 1, characterized in that: In step S1, a file is used to record the correspondence between consecutive frames of pictures in the video and the complete pinyin, where it is recorded that the consecutive frames of pictures from the i-th frame to the j-th frame in the video belong to a corresponding complete pinyin; where i and j are integers, j > i. After marking, adjust the brightness and color of the pictures, and finally form sample data of multiple people.

3. The method for optimizing sound source recognition according to claim 1, wherein: The method collects a large amount of sample data, where the sample data includes face sequences and the pinyin corresponding to each face sequence; by establishing a convolutional neural network and an LSTM network, the parameters trained by the convolutional neural network for each frame of the face sequence are then input into the LSTM network, and the final result is compared with the pinyin data converted by speech recognition. The parameters in the convolutional neural network and the parameters in the LSTM network are processed using the open-source deep learning framework tensorflow through the neural network model to obtain the optimal values, that is, neural network model parameter group 1.

4. The optimized method for sound source identification according to claim 1, characterized in that: In step S3, the position information of each frame of the face is obtained through opencv, and then, the position information of each frame of the face is strung together according to the time series to obtain the face sequence.

5. An optimized system for sound source identification, characterized in that: The system includes: a sample data generation module, a neural network training module, a face prediction pinyin module, and a sound source determination module; The sample data generation module collects videos of different people speaking, calls speech recognition to capture text, converts the text into initials, finals, and complete pinyin, and marks the correspondence between consecutive frames of pictures in the video and the complete pinyin to form sample data of multiple people; The neural network training module, by establishing a neural network model, inputs the sample data into the neural network model to obtain the optimal initials, finals, and complete pinyin data, and forms neural network model parameter group 1; The face prediction pinyin module uses the camera of a handheld device to capture the faces of each member in the meeting, obtains each person's face sequence through opencv, and substitutes the face sequence into neural network model parameter group 1 through the neural network model to output the corresponding pinyin syllables; The sound source determination module obtains the current audio through a microphone, gets the predicted pinyin syllables, compares the predicted pinyin syllables with the pinyin syllables obtained from the face, and if a corresponding face is matched, the current sound source is the face position.

6. The optimized sound source identification system according to claim 5, characterized in that: In the sample data generation module, the corresponding relationship between the continuous frames of pictures and the complete pinyin in the video is recorded through a file, where it is recorded that the continuous frames of pictures from the i-th frame to the j-th frame in the video belong to a corresponding complete pinyin; where i and j are integers, j > i. After marking, the brightness and color of the pictures are adjusted, and finally the sample data of multiple people are formed.

7. An optimized sound source recognition system according to claim 5, characterized in that: The system collects a large amount of sample data, where the sample data includes face sequences and the corresponding pinyin for each face sequence; by establishing a convolutional neural network and an LSTM network, the parameters trained by the convolutional neural network for each frame of the face sequence are then passed into the LSTM network, and the final result is compared with the pinyin data converted through speech recognition. The parameters in the convolutional neural network and the parameters in the LSTM network are processed using the open-source deep learning framework tensorflow through the neural network model to obtain the optimal values, that is, the neural network model parameter group 1 is obtained.

8. An optimized sound source identification system according to claim 5, characterized in that: In the face predicted pinyin module, the position information of each frame of the face is obtained through opencv, and then the position information of each frame of the face is strung together according to the time series to obtain the face sequence.

Citation Information

Patent Citations

  • Speech synthesis method, system and device based on neural network and storage medium

    CN112652291A

  • Speech recognition method, device and computer readable storage medium

    CN107945789A

  • Chinese lip language recognition method based on tone cascade sequence-to-sequence model

    CN111178157A