Sound data processing methods, apparatus, equipment, storage media and software products
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-28
- Publication Date
- 2026-08-14
AI Technical Summary
然而,目前智能机器人的语音交互过程中,其对于外部噪声的降噪效果较差,导致存在语音指令识别的准确性较低的问题
[0020]本申请实施例提供的声音数据处理方法、装置、设备、存储介质及程序产品,通过提取采集到的声音数据的声纹特征,噪声声音数据中包括第一语音指令和噪声。将噪声声纹特征与预构建的声纹数据库中的预设声纹进行匹配,确定噪声声音数据中第一语音指令的频段、第一噪声的频段,噪声预设声纹包括预设噪声声纹、预设指令声纹,噪声第一噪声与噪声预设噪声声纹匹配。根据噪声第一噪声的频段的第一噪声声纹特征,确定目标降噪参数,噪声目标降噪参数用于降低噪声第一噪声。根据噪声目标降噪参数调整声音采集采用的滤波器参数,以通过基于指令声纹的动态降噪,从而提高识别语音指令时的降噪效果,进而提升语音指令的识别准确率。
Smart Images

Figure CN122575373A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent robot technology, and in particular to a sound data processing method, apparatus, device, storage medium, and program product. Background Technology
[0002] With the continuous advancement of artificial intelligence technology, intelligent robots have been widely applied in industries, manufacturing, and product sales. Within the functional system of intelligent robots, voice interaction, as the interaction method closest to natural human communication habits, has its core value in achieving machine control and information acquisition through voice commands. This not only significantly reduces the learning cost for users but also promotes the improvement of robot intelligence, becoming a key factor in measuring the performance of intelligent robots. However, currently, the voice interaction process of intelligent robots suffers from poor noise reduction, resulting in low accuracy in voice command recognition.
[0003] Therefore, improving the noise reduction effect and the accuracy of voice command recognition in intelligent robots is an urgent problem to be solved. Summary of the Invention
[0004] The sound data processing methods, apparatus, devices, storage media, and program products provided in this application are used to improve the noise reduction effect and increase the recognition accuracy of voice commands when intelligent robots recognize voice commands.
[0005] In a first aspect, embodiments of this application provide a sound data processing method, applied to a target device, comprising:
[0006] Extract the voiceprint features from the collected sound data, which includes a first voice command and noise;
[0007] The voiceprint features are matched with preset voiceprints in a pre-constructed voiceprint database to determine the frequency band of the first voice command and the frequency band of the first noise in the sound data. The preset voiceprint includes a preset noise voiceprint and a preset command voiceprint. The first noise is matched with the preset noise voiceprint.
[0008] Based on the first noise acoustic signature characteristics of the first noise frequency band, target noise reduction parameters are determined, and the target noise reduction parameters are used to reduce the first noise;
[0009] Adjust the filter parameters used for sound acquisition according to the target noise reduction parameters.
[0010] Secondly, embodiments of this application provide a sound data processing apparatus, applied to a target device, comprising:
[0011] The first processing module is used to extract the voiceprint features of the collected sound data, wherein the sound data includes a first voice command and noise.
[0012] The second processing module is used to match the voiceprint features with preset voiceprints in a pre-built voiceprint database, and determine the frequency band of the first voice command and the frequency band of the first noise in the sound data. The preset voiceprint includes a preset noise voiceprint and a preset command voiceprint, and the first noise is matched with the preset noise voiceprint.
[0013] The determining module is used to determine target noise reduction parameters based on the first noise acoustic characteristics of the frequency band of the first noise, wherein the target noise reduction parameters are used to reduce the first noise;
[0014] The control module is used to adjust the filter parameters used for sound acquisition according to the target noise reduction parameters.
[0015] Thirdly, embodiments of this application provide an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0016] The memory stores computer-executed instructions;
[0017] The processor executes computer execution instructions stored in the memory to implement various possible implementations as described in the first aspect above.
[0018] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement various possible implementations as described in the first aspect above.
[0019] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements various possible implementations as described in the first aspect above.
[0020] The sound data processing method, apparatus, device, storage medium, and program product provided in this application extract the voiceprint features of the collected sound data. The noise sound data includes a first voice command and noise. The noise voiceprint features are matched with preset voiceprints in a pre-constructed voiceprint database to determine the frequency band of the first voice command and the frequency band of the first noise in the noise sound data. The preset noise voiceprint includes a preset noise voiceprint and a preset command voiceprint. The first noise is matched with the preset noise voiceprint. Based on the first noise voiceprint features of the frequency band of the first noise, target noise reduction parameters are determined. These target noise reduction parameters are used to reduce the first noise. The filter parameters used for sound acquisition are adjusted according to the target noise reduction parameters to improve the noise reduction effect when recognizing voice commands through dynamic noise reduction based on the command voiceprint, thereby improving the recognition accuracy of voice commands. Attached Figure Description
[0021] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0022] Figure 1 This is a schematic diagram of the structure of an intelligent robot provided in an embodiment of this application;
[0023] Figure 2 A schematic flowchart illustrating a sound data processing method provided in an embodiment of this application;
[0024] Figure 3 A flowchart illustrating another sound data processing method provided in an embodiment of this application;
[0025] Figure 4 A flowchart illustrating another sound data processing method provided in an embodiment of this application;
[0026] Figure 5 A flowchart illustrating another sound data processing method provided in an embodiment of this application;
[0027] Figure 6 A flowchart illustrating another sound data processing method provided in an embodiment of this application;
[0028] Figure 7 A flowchart illustrating another sound data processing method provided in an embodiment of this application;
[0029] Figure 8 A flowchart illustrating another sound data processing method provided in an embodiment of this application;
[0030] Figure 9 This is a schematic diagram of the structure of a sound data processing device provided in an embodiment of this application;
[0031] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0032] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0033] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0034] Currently, existing technologies mainly achieve speech noise reduction for intelligent robots in the following ways: relying on noise suppression algorithms with fixed parameters to perform basic noise reduction on the input sound signal through frequency domain filtering or time domain thresholding. Alternatively, combining beamforming technology to enhance the speech in the target direction by utilizing the time difference or phase difference of the sound source arrival, while suppressing lateral and rear noise. Another approach is to use a static noise model library, matching the current environmental noise characteristics with pre-stored noise spectrum templates, and adjusting filter parameters to achieve targeted noise reduction.
[0035] However, these existing noise reduction methods generally suffer from the following problems:
[0036] Noise suppression algorithms with fixed parameters are difficult to adjust the filtering threshold or frequency domain processing range flexibly in non-steady noise or multi-source superposition scenarios due to the fixed parameters. This can easily lead to over-suppression or under-suppression of the target audio segment, resulting in loss of speech details or noise residue.
[0037] While beamforming technology can achieve basic noise reduction by enhancing the direction of the sound source, it relies solely on spatial direction information without combining it with the characteristics of the speech content. In complex sound fields, it is easy to confuse the user's speech with noise and reverberation signals that have similar frequencies, resulting in the effective speech components being filtered out incorrectly or the noise signal being enhanced incorrectly.
[0038] The static noise model library matches based on pre-stored noise spectrum templates, which cannot track the dynamic changes in the spectrum of environmental noise in real time. In scenarios where noise characteristics change rapidly, the template matching results deviate from the actual noise characteristics, causing filter parameter adjustments to lag or become inaccurate. Ultimately, this leads to the target speech signal being covered by noise, distorted, or misjudged, directly reducing the parsing accuracy of speech commands.
[0039] In view of this, embodiments of this application provide a sound data processing method. This method extracts voiceprint features from the collected sound data and matches the voiceprint features of voice commands and noise in the sound data based on a pre-built voiceprint database to lock the voiceprint features of the voice commands. Then, it dynamically determines the target noise reduction parameters for the current environment based on the noise countermeasure strategy corresponding to the noise's voiceprint features. Finally, it adjusts the parameters of the filter used for sound data acquisition according to the target noise reduction parameters, thereby improving the noise reduction effect when the target device recognizes voice commands through dynamic noise reduction, and ultimately improving the accuracy of voice command recognition.
[0040] The subject of this application is the target device, which can be applied in industrial scenarios (such as industrial workshops), manufacturing scenarios (such as production workshops), and product sales scenarios (such as car dealerships and electronic product experience areas). The target device can be, for example, an electronic device with voice interaction capabilities, such as an intelligent robot, a smart speaker, or an in-vehicle voice interaction system. For ease of explanation, subsequent embodiments of this application will use an intelligent robot as an example of the target device.
[0041] Figure 1 This is a structural schematic diagram of an intelligent robot provided in an embodiment of this application. Figure 1 As shown, the intelligent robot includes a microphone array, a vision sensor, a filter connected to the microphone array, and a pose sensor.
[0042] The microphone array is used to collect sound data from the environment surrounding the intelligent robot. This microphone array can be, for example, a ring microphone array, a matrix microphone array, a triangular microphone array, or a micro-electro-mechanical system (MEMS) microphone array. The collected sound data can include voice commands, ambient noise, and noise generated by the intelligent robot itself during operation. Optionally, the intelligent robot can include both an ear ring microphone array and a chest microphone array, where the ear ring microphone array is primarily used to collect voice commands, and the chest microphone array is primarily used to collect noise generated by the intelligent robot itself during operation. The microphone array allows for the preliminary determination of the sound source location of the collected sound data.
[0043] The filter is used to reduce noise in the audio data acquired by the microphone array. This application does not limit the specific form of the filter, as long as it can perform noise reduction on the audio data acquired by the microphone array.
[0044] Visual sensors are used to collect visual data from the environment surrounding the intelligent robot. Examples include image acquisition sensors such as cameras, video cameras, and depth cameras. This application does not limit the number of visual sensors deployed on the intelligent robot, as long as they are capable of collecting visual data from the robot's surrounding environment.
[0045] A pose sensor is used to collect motion perception data for intelligent robots. This pose sensor may include, for example, an accelerometer, a gyroscope, or an inertial measurement unit (IMU).
[0046] Intelligent robots can collect nearby sound data in real time. When a preset voice command is detected, the robot can extract voiceprint features from the collected sound data and match the voiceprint features of the voice command and noise in the sound data based on a pre-built voiceprint database to pinpoint the voiceprint features of the voice command. Then, based on the noise countermeasure strategy corresponding to the noise voiceprint features, the robot dynamically determines the target noise reduction parameters for the current environment. The robot then adjusts the filter used for sound data acquisition according to the target noise reduction parameters, improving the noise reduction effect when the target device recognizes voice commands through dynamic noise reduction, thereby enhancing the accuracy of voice command recognition.
[0047] Optionally, the intelligent robot can also perform comprehensive analysis based on motion perception data and visual data synchronized with the above-mentioned sound data to more accurately determine the sound source location of the voice command, and further adjust the target noise reduction parameters based on the sound source location to perform targeted noise reduction on sounds (i.e. noise) from other locations besides the sound source location of the voice command, so as to improve the noise reduction effect on other noises that cannot be matched with the noise in the voiceprint database.
[0048] The following is based on Figure 1 The target device shown is the execution subject. Specific embodiments are used to describe in detail the technical solution of this application and how the technical solution solves the aforementioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0049] Figure 2 This is a schematic flowchart illustrating a sound data processing method provided in an embodiment of this application. Figure 2 As shown, the method may include:
[0050] S201. Extract the voiceprint features from the collected sound data.
[0051] The audio data includes the first voice command and noise.
[0052] The first voice command can control the target device to perform a specific operation, such as "Please guide the audience to exhibition area A". These voice commands that can control the target device to perform specific operations are pre-stored in the target device. The target device can collect sound data in the surrounding environment in real time. When it detects that the sound data includes the same content as the pre-stored voice command, it can use the sound data including the first voice command as the sound data for voiceprint feature extraction.
[0053] Noise in the audio data can include other sounds besides the first voice command, such as the sounds of equipment operating in the environment, people talking, music, etc.
[0054] In this step, the voiceprint features of the sound data can be extracted using methods such as Mel Frequency Cepstral Coefficients (MFCC) and Short-Time Fourier Transform (STFT). Alternatively, the sound data can be input into a pre-trained voiceprint feature extraction model to extract the voiceprint features of the sound data.
[0055] For example, taking the extraction of speaker features from audio data using MFCC as an example, the audio data can first be preprocessed. This preprocessing can be done through methods such as pre-emphasis (enhancing high-frequency information), framing (dividing continuous speech into short frames, such as 20-30ms / frame), and windowing (such as Hamming windowing to reduce spectral leakage) to obtain preprocessed audio data. Then, a Fast Fourier Transform (FFT) is performed on each frame of audio data to obtain the spectrum, and the linear spectrum is converted to a Mel spectrum using a Mel filter bank. Finally, a Discrete Cosine Transform (DCT) is used to decorrelate and compress the data to obtain the MFCC coefficient sequence, which is then used as the speaker features.
[0056] S202. Match the voiceprint features with the preset voiceprints in the pre-constructed voiceprint database to determine the frequency band of the first voice command and the frequency band of the first noise in the sound data.
[0057] The preset voiceprint includes a preset noise voiceprint and a preset command voiceprint, and the first noise is matched with the preset noise voiceprint.
[0058] The voiceprint database is pre-established during the development or use of the target device, storing a large number of preset noise voiceprints and preset command voiceprints. Preset noise voiceprints are obtained by collecting various common noise samples and extracting their voiceprint features using the same method as the voiceprint feature extraction described above. Preset command voiceprints are obtained by collecting a large number of commonly used user voice command samples, extracting their voiceprint features, and storing them in the database.
[0059] In this step, various matching algorithms can be used to match voiceprint features with preset voiceprints. For example, a matching algorithm based on Euclidean distance or a matching algorithm based on cosine similarity can be used. For each extracted voiceprint feature and each preset voiceprint in the database, the matching degree (such as cosine similarity, Euclidean distance, etc.) between them is calculated. The higher the matching degree, the more similar the two voiceprint features are.
[0060] Specifically, the matching degree between voiceprint features and preset noise voiceprints, and the matching degree between preset command voiceprints can be calculated separately. If a voiceprint feature in the sound data successfully matches the preset noise voiceprint, it indicates that there is a first noise in the sound data corresponding to the preset noise voiceprint, and the frequency band of the first noise corresponding to this part of the voiceprint feature can be determined. If a voiceprint feature in the sound data successfully matches the preset command voiceprint, it indicates that there is a first speech command in the sound data corresponding to the preset command voiceprint, and the frequency band of the first speech command corresponding to this part of the voiceprint feature can be determined.
[0061] S203. Determine the target noise reduction parameters based on the first noise acoustic characteristics of the first noise frequency band.
[0062] The target noise reduction parameter is used to reduce the first noise.
[0063] In this step, noise reduction parameters corresponding to each preset noise voiceprint can be generated in advance based on preset noise voiceprints in the voiceprint database to effectively suppress the noise corresponding to each preset noise voiceprint. For example, the statistical characteristics such as energy distribution, mean, and variance of each preset noise voiceprint at different frequencies can be analyzed in advance. Then, based on these statistical characteristics, different noise reduction algorithms can be used to generate noise reduction parameters that can effectively suppress the noise. For example, an adaptive filtering algorithm can be used for steady-state noise, and a threshold noise reduction algorithm can be used for transient noise. How to generate corresponding noise reduction parameters based on specific noise can be found in existing technologies and will not be elaborated here.
[0064] In practical applications, noise reduction parameters can be optimized through experimentation and debugging. For example, tests can be conducted in different noise environments to observe the noise reduction effect, and the noise reduction parameters corresponding to each preset noise soundprint can be adjusted based on the test results until a satisfactory noise reduction effect is achieved.
[0065] Since the first noise is noise that matches the preset noise voiceprint in the voiceprint database, the target noise reduction parameters corresponding to the first noise can be determined directly based on the mapping relationship between the preset noise voiceprint and the noise reduction parameters, as well as the first noise.
[0066] S204. Adjust the filter parameters used for sound acquisition according to the target noise reduction parameters.
[0067] The parameters of the filter can include its cutoff frequency, gain, etc. These filter parameters can be adjusted according to the target noise reduction parameters to effectively suppress the primary noise. For example, if the target noise reduction parameters indicate that noise reduction needs to be performed within a specific frequency range, the filter's cutoff frequency can be adjusted to better filter noise within that frequency range.
[0068] After adjusting the filter parameters, the sound acquisition device can filter the sound signal according to the adjusted filter parameters during subsequent sound acquisition, thereby improving the quality of the acquired sound data and reducing the interference of primary noise on voice command recognition.
[0069] The method provided in this application extracts the voiceprint features of the collected sound data, which includes a first voice command and noise. The noise voiceprint features are matched with preset voiceprints in a pre-constructed voiceprint database to determine the frequency band of the first voice command and the frequency band of the first noise in the noise sound data. The preset noise voiceprint includes a preset noise voiceprint and a preset command voiceprint. The first noise is matched with the preset noise voiceprint. Based on the first noise voiceprint features of the frequency band of the first noise, target noise reduction parameters are determined. These target noise reduction parameters are used to reduce the first noise. The filter parameters used for sound acquisition are adjusted according to the target noise reduction parameters to improve the noise reduction effect when recognizing voice commands through dynamic noise reduction based on the command voiceprint, thereby improving the recognition accuracy of voice commands.
[0070] The following section provides a detailed explanation of how, in step S202, the voiceprint features are matched with preset voiceprints in the pre-constructed voiceprint database to determine the frequency band of the first voice command and the frequency band of the first noise in the sound data. Figure 3 A flowchart illustrating another sound data processing method provided in this application embodiment is shown. Figure 3 As shown, the aforementioned step S202 may specifically include:
[0071] S301. Determine the frequency band of the first voice command based on the voiceprint feature portion that matches the voiceprint feature of the preset command voiceprint with a degree greater than or equal to a first preset matching degree threshold.
[0072] The first preset matching degree threshold can be determined according to actual needs. Taking the matching degree determined by cosine similarity as an example, the first preset matching degree threshold can be set to 0.95. If the voiceprint features of the voice data include a voiceprint feature part with a cosine similarity greater than or equal to 0.95 with the voiceprint features of the preset command voiceprint, then the voice data corresponding to the voiceprint feature part is the frequency band of the first voice command.
[0073] Specifically, when calculating the matching degree, the extracted voiceprint features can be compared frame by frame with the preset command voiceprint. For each frame of voiceprint features, its matching degree with the preset command voiceprint is calculated. If the matching degree of a certain frame of voiceprint features with the preset command voiceprint is greater than or equal to a first preset matching degree threshold, then the voiceprint features of that frame are considered to belong to the voiceprint features of the first speech command.
[0074] Based on these voiceprint feature frames belonging to the first voice command, their corresponding frequency bands are determined. By analyzing the distribution of these voiceprint feature frames in the frequency domain, it can be determined which frequency ranges the first voice command is mainly concentrated in, thereby determining the frequency band of the first voice command.
[0075] For example, assuming the first preset matching threshold is 0.95, for each frame of extracted voiceprint feature A, its cosine similarity with the preset command voiceprint B is calculated. If the cosine similarity between a frame feature and B is greater than or equal to 0.95, then that frame feature is marked as a feature of the first voice command. Then, the frequency of these marked features at different frequencies is counted, and the frequency range with the most frequent occurrences is the frequency band of the first voice command.
[0076] S302. Determine the frequency band of the first noise based on the portion of the voiceprint features whose matching degree with the preset noise voiceprint is greater than or equal to the second preset matching degree threshold.
[0077] The second preset matching degree threshold can be determined according to actual needs. Taking the matching degree determined by cosine similarity as an example, the second preset matching degree threshold can be set to 0.95. If the voiceprint features of the sound data include a voiceprint feature part with a cosine similarity greater than or equal to 0.95 with the voiceprint features of the preset noise voiceprint, then the sound data corresponding to the voiceprint feature part is the frequency band of the first noise.
[0078] Specifically, when calculating the matching degree, the extracted voiceprint features can be compared frame by frame with the preset noise voiceprint. For each frame of voiceprint features, its matching degree with the preset noise voiceprint is calculated. If the matching degree of a certain frame's voiceprint features with the preset noise voiceprint is greater than or equal to a second preset matching degree threshold, then the voiceprint features of that frame are considered to belong to the voiceprint features of the first noise.
[0079] Based on these voiceprint feature frames belonging to the first noise, their corresponding frequency bands are determined. By analyzing the distribution of these voiceprint feature frames in the frequency domain, it can be determined which frequency ranges the first noise is mainly concentrated in, thereby determining the frequency band of the first noise.
[0080] The following section provides a detailed explanation of how the target noise reduction parameters are determined based on the first noise acoustic signature characteristics of the first noise frequency band in step S203. Figure 4 A flowchart illustrating another sound data processing method provided in this application embodiment. (See attached diagram.) Figure 4 As shown, the aforementioned step S203 may specifically include:
[0081] S401. Determine the target sound source location of the first voice command based on the sound data, the visual data synchronized with the sound data in time, and the motion perception data of the target device synchronized with the sound data in time.
[0082] The visual data includes the target object from which the first voice command is issued. This target object can be, for example, a worker in a workshop, a salesperson in a sales scenario, a customer, or an electronic device capable of issuing voice commands, such as a mobile phone, speaker, or smart speaker.
[0083] In real-world scenarios, in addition to collecting audio data, the target device also collects visual and motion sensing data simultaneously. Visual data can be acquired through image acquisition devices, and includes image information of the target device's surrounding environment, which may include the target object that issued the first voice command. Motion sensing data can be acquired through sensors such as accelerometers and gyroscopes, reflecting the target device's own motion state.
[0084] In this step, based on the time tags of the collected sound data, visual data that is time-synchronized with the sound data (i.e., data segments with the same time tags in the visual data) and motion perception data that is time-synchronized with the sound data (i.e., data segments with the same time tags in the motion perception data) can be determined.
[0085] Motion sensing data can be used to update the target device's coordinate system in real time based on its real-time pose. Taking an intelligent robot as an example, when the robot is standing, it can determine its current pose using motion sensing data, establishing a first coordinate system (e.g., a spatial coordinate system with the robot's center as the origin). When the robot is moving, it can also determine its current pose using motion sensing data, establishing a second coordinate system (different from the first). Similarly, when the robot is crouching, it can determine its current pose using motion sensing data, establishing a third coordinate system (e.g., a spatial coordinate system with the robot's body center as the origin). By using motion sensing data, a coordinate system corresponding to the target device's real-time pose can be determined, providing a reference coordinate system for subsequently determining the location of the target sound source and improving the accuracy of the determined location.
[0086] For audio data, the initial sound source location of the first voice command can be determined by utilizing the spatial information of the microphone array and the characteristics of sound propagation. A microphone array can simultaneously collect sound signals through multiple microphones, and the initial sound source location can be calculated based on information such as the time difference and intensity difference of the sound signals received by different microphones.
[0087] For visual data, the spatial location of objects (such as people, devices, etc.) in the visual data can be detected by deep learning-based object detection algorithms (such as YOLO, Faster R-CNN, etc.).
[0088] The initial sound source location and spatial location mentioned above can both be in a coordinate system corresponding to the real-time pose of the target device, so as to reduce the data processing caused by coordinate system transformation and thus improve the efficiency and accuracy of location determination.
[0089] One possible implementation is to determine the spatial position of the target object in the visual data using visual data and motion perception data, and to determine the initial sound source position of the target object using sound data and motion perception data. Then, the initial sound source position is adjusted based on the spatial position of the target object to obtain the target sound source position of the first voice command.
[0090] Another possible approach is to fuse sound data, visual data, and motion perception data, such as by using a pre-trained acoustic-visual-motion offset model to fuse these three types of data, to obtain the target sound source location of the first voice command.
[0091] For example, taking the person issuing the first voice command as the target object, the location of the target sound source of the first voice command can be determined through the following sub-steps:
[0092] S4011. In the coordinate system of the target device determined based on motion perception data, determine the head position and lip movement data of each candidate object in the visual data, and the initial sound source position of the first speech command in the sound data.
[0093] Candidate objects can be any object included in the visual data, such as all people appearing in the visual data. Within the coordinate system of the target device determined based on motion-sensing data, target detection and tracking algorithms, such as deep learning-based target detection algorithms, are used to detect candidate objects in the visual data and track their head positions. Simultaneously, by analyzing the lip image features of the candidate objects, their lip movement data is determined; this lip movement data can reflect whether the candidate object is speaking.
[0094] The initial sound source location can be calculated using information such as the time difference and intensity difference of the sound signals received by different microphones in the microphone array. Since this initial sound source location is calculated based on the sound signal, its accuracy is poor. Therefore, the initial sound source location can be further adjusted based on visual data to improve the accuracy of determining the sound source location of the first voice command.
[0095] S4012. Determine the individual attribute characteristics of the target object based on the first voiceprint feature in the sound data.
[0096] Individual attribute characteristics include at least one of gender and age characteristics.
[0097] In this step, a pre-trained classification model can be used to identify individual attribute features. For example, convolutional neural network models, decision tree models, etc., can be used to identify the gender and age features of the target object. Specifically, for example, the first voiceprint feature can be input into a trained convolutional neural network model, and the model will output the individual attribute features of the target object, such as gender and age.
[0098] For example, for the first voice command issued by user A, after inputting into a pre-trained classification model, it can output that the user's gender is male and the age is 25-30 years old; for the first voice command issued by user B, after inputting into a pre-trained classification model, it can output that the user's gender is female and the age is 40-45 years old, etc.
[0099] S4013. Based on individual attribute characteristics, initial sound source location, head position of each candidate object, and lip movement data of each candidate object, determine the target object from the candidate objects.
[0100] In this step, a reasonable distance range can be set within the coordinate system of the target device determined based on motion sensing data, centered on the initial sound source location. This distance range can be set by comprehensively considering the characteristics of sound propagation and environmental factors in the actual scene. For example, in a relatively open indoor scene, sound propagation is relatively stable, and the distance range can be appropriately widened; while in a noisy environment with a complex spatial layout, the distance range needs to be narrowed. Candidate objects whose distance from the initial sound source location is within this range (e.g., whether they are within this range is determined by their head position) are selected to form a preliminary screening set. This is because the target object issuing the first voice command is highly likely to be physically close to the initial location where the sound was emitted.
[0101] Next, lip movement data is used to further refine the candidate pool in the initial screening. Lip movement data can directly reflect whether a candidate is speaking. By analyzing the lip image features of the candidate, such as the degree of lip opening and closing, movement trajectory, and frequency, it is determined whether they are speaking. Some judgment criteria can be preset, such as identifying a candidate as speaking when the degree of lip opening and closing reaches a certain threshold and the movement frequency is within a specific range. Priority is given to candidates in the initial screening who are speaking, because those who are speaking are more likely to be the target of the first voice command.
[0102] Next, the final target object is determined using individual attribute features. The individual attribute features of the initially screened candidate objects who are currently speaking are meticulously compared with the individual attribute features of the target object determined based on the first voiceprint feature. For gender, if the gender of the candidate object matches that of the target object, a higher matching score is given. For age, the overlap between the age range of the candidate object and the age range of the target object is compared; the higher the overlap, the higher the matching score. To comprehensively consider all features, different weights can be assigned to gender and age features. The weight allocation should be adjusted according to the actual situation. For example, in some scenarios, gender features have high discriminative power and can be given a higher weight; while in other scenarios, age features are more critical, so the weight of age features is increased accordingly. The matching scores of all features are weighted and summed, and the candidate object with the highest score is the target object.
[0103] S4014. Determine the target sound source location based on the initial sound source location and the head position of the target object in the visual data.
[0104] After identifying the target object, it can be assumed that the head position of the target object in the visual data is more accurate than the initial sound source position. Therefore, the initial sound source position can be adjusted based on the head position to obtain the target sound source position.
[0105] Alternatively, the head position of the target object can be fused with the initial sound source position, such as by weighted averaging or probability distribution-based fusion, to improve the accuracy of the target sound source position determination.
[0106] S402. Based on the location of the target sound source and the characteristics of the first noise soundprint, determine the target noise reduction parameters.
[0107] The target noise reduction parameters are also used to reduce the noise of the second noise, the location of which is different from that of the target noise source.
[0108] Optionally, the second noise may also include noise whose voiceprint features differ from those of the first voiceprint.
[0109] Since the first noise only corresponds to the preset noise in the voiceprint database, if the corresponding preset noise cannot be matched in the voiceprint database, the noise that is different from the target voiceprint location and / or has different voiceprint features can be determined as the second noise based on the comparison of the voiceprint location and voiceprint features.
[0110] Specifically, we can first identify the sounds from each sound source location in the sound data, then lock the target sound source location, and consider all sounds from other locations outside the target sound source location as second noise. Secondly, for sounds also from the target sound source location, we can compare all the voiceprint features of these sounds with the voiceprint features of the first noise. If there are sounds whose similarity to the voiceprint features of the first noise is lower than a preset value, these sounds can also be considered as second noise.
[0111] After determining the second noise, target noise reduction parameters can be directly generated based on the first and second noises. The noise reduction parameters for the first noise are determined as described in step S203 above. The noise reduction parameters for the second noise can be determined, for example, based on the noise type of the second noise, or by performing frequency band analysis on the second noise (e.g., performing spectral analysis on the second noise to determine the frequency bands where its energy is concentrated; setting higher noise reduction gain for high-energy frequency bands and lower noise reduction gain for low-energy frequency bands). Then, the noise reduction parameters for the first and second noises are fused to determine the target noise reduction parameters.
[0112] For example, a weighted fusion method can be used to fuse the noise reduction parameters of the first noise and the second noise to determine the target noise reduction parameters. The degree of interference can be judged by analyzing factors such as the energy magnitude and duration of the two noises. For instance, if the first noise has high energy and long duration, causing severe interference to the audio data, its noise reduction parameter can be assigned a weight of 0.6; the second noise has relatively less interference, so its noise reduction parameter can be assigned a weight of 0.4. In the filter parameter space of the target device, for each parameter dimension (such as noise reduction gain in different frequency bands, filter coefficients, etc.), the noise reduction parameters of the first and second noises are multiplied by their respective weights, and then the results are added together. For example, in the 1500Hz frequency band, the noise reduction gain of the first noise is 10dB with a weight of 0.6; the noise reduction gain of the second noise is 8dB with a weight of 0.4. Then, the fused noise reduction gain for this frequency band is 10×0.6+8×0.4=9.2dB. The same weighted calculation is performed on all parameter dimensions to finally obtain the fused target noise reduction parameters.
[0113] For example, taking the determination of noise reduction parameters for the second noise based on its noise type as an example, the target noise reduction parameters can be determined through the following sub-steps:
[0114] S4021. Based on the location of the target sound source and the voiceprint characteristics of the target object, determine the frequency band of the second noise.
[0115] The system identifies sounds originating from various sound sources within the sound data, then pinpoints the target sound source location. Sounds from locations other than the target sound source are considered secondary noise. Next, for sounds also originating from the target sound source location, their voiceprint features are compared with those of the first noise. If a sound exists with a similarity to the first noise's voiceprint features below a preset value, that sound can also be considered secondary noise. The frequency band of the secondary noise is then determined accordingly.
[0116] S4022. Determine the noise type of the second noise based on its frequency band.
[0117] The noise type includes at least one of steady-state noise, transient noise, and impulse noise.
[0118] To determine steady-state noise, the long-term stability of the sound signal within the second noise frequency band can be analyzed. Steady-state noise is characterized by its relatively stable spectrum and intensity over a relatively long period. By recording the signal, the spectral changes within the second noise frequency band can be analyzed. If, over a continuous period, the main frequency components and energy distribution of the spectrum remain essentially unchanged, and the fluctuation range of the sound intensity is within a small threshold, then the second noise can be determined to be steady-state noise.
[0119] Transient noise can be identified by observing the sudden changes in the sound signal. Transient noise is typically noise that appears abruptly and lasts for a short period. It can be identified by analyzing the time-domain waveform of the sound signal within a second noise frequency band. When a rapid increase and decrease in amplitude is observed in the sound signal within an extremely short time (e.g., a few milliseconds to tens of milliseconds), and the spectrum changes significantly during this process, the noise can be identified as transient noise.
[0120] The identification of impulse noise is primarily based on its high intensity and short duration. Impulse noise is typically generated by sudden impacts or collisions and possesses very high instantaneous energy. Within the second noise frequency band, if an extremely high intensity sound signal is detected momentarily, and its duration is very short, it can be identified as impulse noise.
[0121] Optionally, the noise type of the second noise can be determined based on a pre-built noise classification model, noise classification algorithm, etc. For example, the noise classification model can be a convolutional neural network model, and the noise classification algorithm can be a support vector machine algorithm, etc.
[0122] S4023. Using preset noise countermeasure strategies corresponding to the noise type of the second noise and the voiceprint characteristics of the first noise respectively, determine the target noise reduction parameters.
[0123] For steady-state noise, adaptive filtering algorithms (such as the least mean square algorithm or the recursive least square algorithm) can be used for noise reduction. Adaptive filtering algorithms can automatically adjust the filter coefficients according to the input noise signal to achieve the best noise reduction effect.
[0124] For transient noise, threshold noise reduction algorithms can be used. These algorithms monitor and judge audio signals in real time based on a preset threshold. When the signal amplitude exceeds the threshold, it is identified as transient noise and processed accordingly to reduce its impact. The threshold noise reduction algorithm first performs a detailed analysis of the audio signal within a second noise frequency band to determine the signal amplitude distribution. Then, based on the sudden and short-duration characteristics of transient noise, a reasonable threshold is set. During actual operation, the algorithm continuously scans the audio signal. Once it detects that the signal amplitude exceeds the threshold, it quickly activates the noise reduction mechanism, such as attenuating or filtering the signal exceeding the threshold. This effectively removes transient noise while minimizing interference with normal audio signals.
[0125] For impulse noise, peak suppression algorithms can be used for noise reduction. Peak suppression algorithms automatically adjust the degree of noise suppression based on the peak intensity of the input noise signal to achieve optimal noise reduction. Peak suppression algorithms can set a peak threshold to address the characteristic of impulse noise generating instantaneous high-energy peaks. When monitoring the sound signal within the second noise frequency band, once the peak intensity of the signal is detected to exceed this threshold, the suppression mechanism is immediately activated. The suppression process can attenuate the portion exceeding the threshold in a linear or non-linear manner, reducing the peak intensity of the signal to an acceptable range.
[0126] Then, the denoising strategies for the second noise type and the denoising parameters for the acoustic signature characteristics of the first noise are comprehensively considered to determine the final target denoising parameters. During the synthesis process, each parameter can be adjusted and optimized according to the actual situation. For example, assuming that the second noise is steady-state noise and the acoustic signature characteristics of the first noise indicate that there is strong noise energy at certain specific frequencies, the denoising gain at these specific frequencies can be appropriately increased based on the adaptive filtering algorithm.
[0127] In one possible implementation, Figure 5 A flowchart illustrating another sound data processing method provided in this application embodiment is shown. Figure 5 As shown, the aforementioned step S4023 may specifically include:
[0128] S501. Using preset noise countermeasure strategies corresponding to the noise type of the second noise and the voiceprint characteristics of the first noise respectively, determine the initial noise reduction parameters.
[0129] In this step, the method for determining the initial noise reduction parameters can refer to the content of step S4023 above. The only difference is that the target noise reduction parameters determined in S4023 are used as the initial noise reduction parameters. After further adjusting the initial noise reduction parameters, the final target noise reduction parameters are generated.
[0130] S502. Based on the operating sound data of the target device, determine the noise reduction parameter adjustment parameters for the operating sound data.
[0131] The target device generates its own operational sound data during operation, which may interfere with sound acquisition and voice command recognition. Therefore, it is necessary to determine and adjust the noise reduction parameters based on the target device's operational sound data. The operational sound data can be, for example, as described above. Figure 1 The data was acquired using the chest microphone array mentioned in the text.
[0132] Feature extraction is performed on the collected operating sound data, such as extracting its power spectrum and energy distribution. Based on the characteristics of the operating sound data, noise reduction parameters are determined and adjusted.
[0133] For example, noise reduction parameter adjustment parameters can be pre-configured for different types of running sound data. Based on the characteristics of the running sound data, it can be determined which type of running sound data it belongs to, and then the corresponding noise reduction parameter adjustment parameters can be determined.
[0134] Alternatively, the noise reduction parameters corresponding to the characteristics of the running sound data can be determined using a specific noise reduction algorithm. For example, the noise reduction parameters corresponding to the characteristics of the running sound data can be determined through comparative analysis.
[0135] S503. Adjust the initial noise reduction parameters using the noise reduction parameter adjustment parameters to obtain the target noise reduction parameters.
[0136] In this step, adjustments can be made, for example, by weighted summation. The adjusted noise reduction parameters can be weighted and added to the initial noise reduction parameters to obtain the adjusted target noise reduction parameters.
[0137] The adjusted target noise reduction parameters can better adapt to the operating environment and noise conditions of the target equipment, thereby improving the noise reduction effect. In practical applications, the target noise reduction parameters can be further optimized through continuous monitoring and adjustment to achieve the best noise reduction performance.
[0138] Figure 6 A flowchart illustrating another sound data processing method provided in this application embodiment is shown. Figure 6 As shown, the method may further include:
[0139] S601. Based on the voiceprint features of the first voice command and the mapping relationship between the voiceprint features and voiceprint permissions, determine the first voiceprint permission of the voiceprint features of the first voice command.
[0140] The target device has a pre-established mapping relationship between voiceprint features and voiceprint permissions, which can be stored in a database. Voiceprint permissions represent the user's permission level to operate the target device. Different users can have different voiceprint permissions. For example, in a workshop scenario, the voiceprint permissions of the workshop manager are higher than those of ordinary staff; similarly, in a product sales scenario, the voiceprint permissions of sales personnel are higher than those of customers.
[0141] Once the voiceprint features of the first voice command are obtained, these features are compared with voiceprint features in a mapping database. For example, feature matching algorithms, such as the Euclidean distance method or cosine similarity method mentioned above, can be used to find the best-matching voiceprint feature record. Based on the matching record, the first voiceprint permission corresponding to the voiceprint features of the first voice command is determined. For example, the database records that some voiceprint features correspond to high-level permissions, allowing the execution of all device functions; while some voiceprint features correspond to low-level permissions, allowing only the execution of some basic functions.
[0142] S602. If the first voiceprint permission meets the voiceprint permission conditions required to execute the first voice command, then the first operation corresponding to the first voice command is executed based on the first voice command.
[0143] After determining the initial voiceprint permission, it is necessary to determine whether this permission meets the voiceprint permission conditions required to execute the first voice command. The target device can pre-set the minimum voiceprint permission standards required for different voice commands. For example, a simple command like checking the weather may only require low-level voiceprint permission; while sensitive operations such as modifying device system settings require high-level voiceprint permission.
[0144] The first voiceprint permission is compared with the voiceprint permission conditions required for the first voice command. If the first voiceprint permission meets or exceeds the required permission conditions, the target device can execute the corresponding first operation based on the first voice command. For example, if the first voice command is "play music", and the first voiceprint permission meets the permissions required for the music playback command, the target device can call the music playback module to search for and play the corresponding music.
[0145] Optionally, a secondary verification of voiceprint permissions can be performed based on the target object in the visual data to further improve the security and accuracy of voiceprint permissions. Specifically, if the first voiceprint permission meets the voiceprint permission conditions required to execute the first voice command, and the target object in the visual data passes the permission verification, then the first operation corresponding to the first voice command is executed based on the first voice command.
[0146] Specifically, facial recognition technology can be used to identify target objects in visual data. For example, the target device may pre-store facial image features of users with different voiceprint permissions. When the first voiceprint permission meets the permission conditions required to execute the first voice command, the facial image features of the target object in the visual data are extracted and compared with the stored facial image features. If the facial features match successfully, that is, the target object passes the permission verification, then the target device will execute the corresponding first operation based on the first voice command. For example, the first voiceprint permission meets the permission conditions for the "enable device privacy mode" command, but facial recognition is still needed to verify whether the target object is a user with that permission. Only when facial recognition also passes will the device's privacy mode be enabled.
[0147] The method provided in this application determines the first voiceprint permission of the first voice command based on the voiceprint features of the first voice command and the mapping relationship between the voiceprint features and voiceprint permissions. If the first voiceprint permission meets the voiceprint permission conditions required to execute the first voice command, then the first operation corresponding to the first voice command is executed based on the first voice command. This ensures that only voice commands issued by users with legitimate permissions can trigger the corresponding operation, avoiding misoperation or malicious operation by unauthorized personnel, thereby improving the security and reliability of target device control.
[0148] Figure 7 A flowchart illustrating another sound data processing method provided in this application embodiment is shown. Figure 7 As shown, the method may further include:
[0149] S701. Upon receiving a second voice command, determine the second voiceprint permission of the second voice command based on the voiceprint characteristics of the second voice command and the mapping relationship between the voiceprint characteristics and voiceprint permissions.
[0150] When the target device receives the second voice command, it extracts the voiceprint features of the second voice command using the same method as extracting the voiceprint features of the first voice command. Then, it compares the extracted voiceprint features of the second voice command with a pre-established mapping database of voiceprint features and voiceprint permissions. Using a feature matching algorithm, it finds the record that best matches the voiceprint features of the second voice command, thereby determining the second voiceprint permission corresponding to the voiceprint features of the second voice command. For example, the second voice command might be a new control command; by comparing it with the mapping database, it might be found that its voiceprint features correspond to intermediate voiceprint permissions.
[0151] S702. If the target device is performing the first operation and the second voiceprint permission is higher than the first voiceprint permission, then stop the first operation and execute the second operation corresponding to the second voice command.
[0152] While the target device is performing the first operation, once the second voiceprint permission is determined, it is compared with the first voiceprint permission. If the second voiceprint permission is higher than the first voiceprint permission, it means that the user who issued the second voice command has higher operation permissions, and the target device can stop the currently executing first operation.
[0153] For example, during the process of the intelligent robot handling the handling task (first operation) given by the voice instructions of ordinary staff, if it receives a second voice command "shut down equipment B" from the workshop supervisor, the intelligent robot will stop the handling task and immediately execute the operation of shutting down equipment B, since the second voiceprint permission corresponding to the second voice command is higher than the first voiceprint permission (the voiceprint permission of the ordinary staff's voice instruction) based on the judgment.
[0154] Optionally, after completing the second operation, the intelligent robot may choose not to continue performing the unfinished first operation, or it may perform the following step S703.
[0155] S703. After the target device completes the second operation, continue to execute the first operation.
[0156] Once the target device completes the second operation corresponding to the second voice command, if the first operation is not yet completed and needs to be continued, the target device can continue to execute the first operation. For example, during the process of an intelligent robot handling a material handling task instructed by a regular worker (the first operation), if it receives a second voice command "Close device B" from the workshop supervisor, based on the judgment that the second voiceprint permission corresponding to the second voice command is higher than the first voiceprint permission (the voiceprint permission of the regular worker's voice instruction), the intelligent robot will stop the material handling task and immediately execute the operation of closing device B. After completing the operation of closing device B, it can continue to execute the unfinished first operation (i.e., the material handling task instructed by the regular worker's voice).
[0157] The method provided in this application, upon receiving a second voice command, determines the second voiceprint permission of the second voice command based on its voiceprint features and the mapping relationship between voiceprint features and voiceprint permissions. If the target device is performing a first operation and the second voiceprint permission is higher than the first voiceprint permission, the first operation is stopped and the second operation corresponding to the second voice command is executed. After the target device completes the second operation, the first operation continues. This achieves flexible and prioritized control of device operation based on different voiceprint permissions, ensuring that high-permission voice commands can interrupt low-permission operations in a timely manner and be executed first, while also ensuring that low-permission operations can continue after high-permission operations are completed. This avoids inconvenience caused by operation interruption, ensures the continuity and efficiency of device operation processes, thereby improving the user experience and convenience of using the device, enhancing the rationality and accuracy of the device's response to different user commands, and better meeting the needs of device operation management in multi-user scenarios.
[0158] Figure 8 A flowchart illustrating another sound data processing method provided in this application embodiment is shown. Figure 8 As shown, the method may further include:
[0159] S801. When multiple voice commands are received within a preset time period, the identity of the object corresponding to each voice command is determined based on the voiceprint characteristics of each voice command.
[0160] Within a preset time period, the target device may receive multiple voice commands. For each received voice command, its voiceprint features can be extracted using the voiceprint feature extraction method described earlier.
[0161] One possible implementation involves pre-storing voiceprint feature templates for different object identities in the target device. The voiceprint features of each voice command are compared with these stored templates. A feature matching algorithm is used to find the best-matching voiceprint feature template, thus determining the object identity corresponding to each voice command. For example, if three voice commands are received within one minute, voiceprint feature comparison can determine that the first voice command was issued by object A, the second by object B, and the third by object A again.
[0162] Another possible implementation is to determine the individual attribute features corresponding to each voiceprint feature based on the voiceprint characteristics of each voice command. The method for determining individual attribute features based on voiceprint features can be referred to in the aforementioned embodiments and will not be repeated here. Then, based on the individual attribute features corresponding to each voice command, the object's identity is classified. For example, if four voice commands are received within thirty seconds, analysis of the voiceprint features reveals that the first and third voice commands correspond to adult males with relatively deep and broad voices; the second voice command corresponds to adult females with relatively thin and bright voices; and the fourth voice command corresponds to a child with immature voices. Based on these individual attribute features, combined with preset classification rules, the first and third voice commands are classified as being issued by object C (adult male), the second voice command as being issued by object D (adult female), and the fourth voice command as being issued by object E (child).
[0163] S802. Cache the timing data of voice commands corresponding to each object identity according to the object identity.
[0164] Among them, the timing data of voice commands is used for multi-turn dialogue context association analysis.
[0165] After determining the object identity corresponding to each voice command, the voice commands corresponding to each object identity are cached in chronological order. Specifically, data structures such as queues or lists can be used to store the chronological data of these voice commands.
[0166] For example, voice commands issued by object A can be stored in the cache queue corresponding to object A in the order they are issued; voice commands issued by object B can be stored in the cache queue corresponding to object B.
[0167] Voice command timing data contains information such as the content and timing of the voice command, which can be used for contextual analysis in subsequent multi-turn dialogues. For example, in later interactions, this timing data can be used to understand the previous commands given by object A or object B, thereby better understanding the meaning of the current command.
[0168] S803. When a new voice command is received, the identity of the object corresponding to the new voice command is determined based on the voiceprint characteristics of the new voice command.
[0169] When the target device receives a new voice command, it also extracts the voiceprint features of the new voice command. Then, it uses the extracted voiceprint features to determine the identity of the object corresponding to the new voice command. For example, after a new voice command is issued, voiceprint feature comparison determines that the command was issued by object C.
[0170] S804. If the object identity corresponding to the new voice command belongs to the object identity of the cached voice command timing data, then based on the voice command timing data of the object identity corresponding to the new voice command, determine the association between the new voice command and the historical commands in the voice command timing data.
[0171] If the object identity corresponding to the new voice command is an object identity for which voice command timing data has already been cached, then the cached voice command timing data for that object identity can be retrieved, and contextual association analysis can be performed on the historical commands in the voice command timing data. For example, semantic analysis, keyword matching, and other methods can be used to determine the contextual association. For instance, suppose the new voice command is "continue playing," and the object's previous voice command was "play a song by a certain singer." Through semantic analysis, it can be determined that the new voice command is to continue playing the singer's song; that is, there is a relationship of continued execution between the new voice command and the historical command.
[0172] S805: Based on the new voice command and its associated relationships, execute the third operation corresponding to the new voice command.
[0173] After establishing the association between the new voice command and historical commands, the target device can execute the corresponding third operation based on the new voice command and the association. Using the example above, since the new voice command "continue playing" and the historical command "play a song by a certain artist" have a continuation relationship, the target device can continue playing that artist's song, completing the third operation corresponding to the new voice command.
[0174] The method provided in this application determines the object identity corresponding to each voice command based on the voiceprint features of each voice command when multiple voice commands are received within a preset time period. The timing data of the voice commands corresponding to each object identity is cached, and this timing data is used for multi-turn dialogue context association analysis. When a new voice command is received, the object identity corresponding to the new voice command is determined based on the voiceprint features of the new voice command. If the object identity corresponding to the new voice command belongs to an object identity whose timing data has been cached, the association relationship between the new voice command and historical commands in the timing data is determined based on the timing data of the voice command corresponding to the object identity. Then, based on the new voice command and its associated relationships, the third operation corresponding to the new voice command is executed to accurately capture the semantic logic of different objects in multi-turn dialogues. The historical command information of each object is distinguished and integrated according to their identity, so that the new voice command can closely follow the previous dialogue content of the object. This ensures that in multi-person interaction scenarios, the multi-turn voice commands of each object can form a coherent and orderly contextual association, thereby improving the accuracy and coherence of voice command processing in multi-person interaction scenarios. This provides users with a more natural, smooth voice interaction experience that conforms to their personal communication habits, and effectively avoids the problem of command confusion and context breakage caused by multiple people interacting at the same time.
[0175] Figure 9 This is a schematic diagram of the structure of a sound data processing device provided in an embodiment of this application. Figure 9 As shown, the sound data processing device may include: a first processing module 11, a second processing module 12, a determination module 13, and a control module 14.
[0176] The first processing module 11 is used to extract the voiceprint features of the collected sound data, which includes the first voice command and noise.
[0177] The second processing module 12 is used to match the voiceprint features with the preset voiceprints in the pre-built voiceprint database, determine the frequency band of the first voice command and the frequency band of the first noise in the sound data, the preset voiceprints include preset noise voiceprints and preset command voiceprints, and the first noise is matched with the preset noise voiceprints.
[0178] The determination module 13 is used to determine the target noise reduction parameters based on the first noise acoustic characteristics of the first noise frequency band. The target noise reduction parameters are used to reduce the first noise.
[0179] The control module 14 is used to adjust the filter parameters used for sound acquisition according to the target noise reduction parameters.
[0180] Optionally, the second processing module 12 is specifically used to determine the frequency band of the first voice command based on the portion of the voiceprint features that match the voiceprint features of the preset command voiceprint with a degree greater than or equal to a first preset matching threshold. The frequency band of the first noise is determined based on the portion of the voiceprint features that match the voiceprint features of the preset noise voiceprint with a degree greater than or equal to a second preset matching threshold.
[0181] Optionally, module 13 is specifically used to determine the target sound source location of the first voice command based on sound data, visual data synchronized with the sound data in time, and motion perception data of the target device synchronized with the sound data in time. Based on the target sound source location and the first noise voiceprint characteristics, target noise reduction parameters are determined. The visual data includes the target object issuing the first voice command, and the target noise reduction parameters are also used to reduce noise in the second noise, the sound source location of which is different from the target sound source location.
[0182] Optionally, the determining module 13 is specifically used to determine, within the coordinate system of the target device determined based on motion perception data, the head position and lip movement data of each candidate object in the visual data, and the initial sound source position of the first voice command in the sound data. Based on the first voiceprint feature in the sound data, the individual attribute features of the target object are determined. Based on the individual attribute features, the initial sound source position, the head position of each candidate object, and the lip movement data of each candidate object, the target object is determined from the candidate objects. Based on the initial sound source position and the head position of the target object in the visual data, the target sound source position is determined. The individual attribute features include at least one of gender and age features.
[0183] Optionally, module 13 is specifically used to determine the frequency band of the second noise based on the location of the target sound source and the acoustic signature characteristics of the target object. Based on the frequency band of the second noise, the noise type of the second noise is determined. Preset noise countermeasure strategies corresponding to the noise type of the second noise and the acoustic signature characteristics of the first noise are used to determine the target noise reduction parameters. The noise type includes at least one of steady-state noise, transient noise, and impulse noise.
[0184] Optionally, module 13 is specifically used to determine initial noise reduction parameters by employing preset noise countermeasure strategies corresponding to the noise type of the second noise and the acoustic signature characteristics of the first noise, respectively. Based on the operating sound data of the target device, noise reduction parameter adjustment parameters are determined for the operating sound data. The initial noise reduction parameters are adjusted using the noise reduction parameter adjustment parameters to obtain the target noise reduction parameters.
[0185] Optionally, the determining module 13 is further configured to determine the first voiceprint permission of the first voice command based on the voiceprint features of the first voice command and the mapping relationship between the voiceprint features and voiceprint permissions. The controlling module 14 is further configured to execute a first operation corresponding to the first voice command based on the first voice command if the first voiceprint permission meets the voiceprint permission conditions required to execute the first voice command, or execute a first operation corresponding to the first voice command based on the first voice command if the first voiceprint permission meets the voiceprint permission conditions required to execute the first voice command and the target object in the visual data passes the permission verification.
[0186] Optionally, the determining module 13 is further configured to, upon receiving the second voice command, determine the second voiceprint permission of the second voice command based on the voiceprint characteristics of the second voice command and the mapping relationship between the voiceprint characteristics and voiceprint permissions. The control module 14 is further configured to, if the target device is performing the first operation and the second voiceprint permission is higher than the first voiceprint permission, stop the first operation and execute the second operation corresponding to the second voice command.
[0187] Optionally, the determining module 13 is further configured to, when multiple voice commands are received within a preset time period, determine the object identity corresponding to each voice command based on the voiceprint features of each voice command. The control module 14 is further configured to cache the voice command timing data corresponding to each object identity according to the object identity, and the voice command timing data is used for multi-turn dialogue context association analysis. The determining module 13 is further configured to, when a new voice command is received, determine the object identity corresponding to the new voice command based on the voiceprint features of the new voice command. If the object identity corresponding to the new voice command belongs to the object identity of the cached voice command timing data, then based on the voice command timing data of the object identity corresponding to the new voice command, determine the association relationship between the new voice command and the historical commands in the voice command timing data. The control module 14 is further configured to, based on the new voice command and the association relationship, execute the third operation corresponding to the new voice command.
[0188] The sound data processing apparatus provided in this application embodiment can execute the sound data processing method in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described again here.
[0189] Figure 10 This is a schematic diagram of an electronic device provided in an embodiment of this application. The electronic device is used to execute the aforementioned sound data processing method. Figure 10 As shown, the electronic device 1000 may include at least one processor 1001, a memory 1002, and a communication interface 1003.
[0190] The memory 1002 is used to store programs. Specifically, the program may include program code, which includes computer operation instructions.
[0191] The memory 1002 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0192] The processor 1001 is used to execute computer execution instructions stored in the memory 1002 to implement the method described in the foregoing method embodiments. The processor 1001 may be a CPU, an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.
[0193] The processor 1001 can communicate and interact with external devices through the communication interface 1003. These external devices can be, for example, the aforementioned database. In specific implementations, if the communication interface 1003, memory 1002, and processor 1001 are implemented independently, they can be interconnected via a bus to complete communication. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc., but this does not imply that there is only one bus or one type of bus.
[0194] Optionally, in a specific implementation, if the communication interface 1003, memory 1002 and processor 1001 are integrated on a single chip, then the communication interface 1003, memory 1002 and processor 1001 can communicate through an internal interface.
[0195] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0196] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.
[0197] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
Claims
1. A method for processing sound data, characterized in that, Applied to a target device, the method includes: Extract the voiceprint features from the collected sound data, which includes a first voice command and noise; The voiceprint features are matched with preset voiceprints in a pre-constructed voiceprint database to determine the frequency band of the first voice command and the frequency band of the first noise in the sound data. The preset voiceprint includes a preset noise voiceprint and a preset command voiceprint. The first noise is matched with the preset noise voiceprint. Based on the first noise acoustic signature characteristics of the first noise frequency band, target noise reduction parameters are determined, and the target noise reduction parameters are used to reduce the first noise; Adjust the filter parameters used for sound acquisition according to the target noise reduction parameters.
2. The method according to claim 1, characterized in that, The step of matching the voiceprint features with preset voiceprints in a pre-constructed voiceprint database to determine the frequency band of the first voice command and the frequency band of the first noise in the sound data includes: The frequency band of the first voice command is determined based on the portion of the voiceprint features that match the voiceprint features of the preset command voiceprint with a degree greater than or equal to a first preset matching threshold. The frequency band of the first noise is determined based on the portion of the voiceprint features that match the preset noise voiceprint features with a degree greater than or equal to a second preset matching threshold.
3. The method according to claim 1, characterized in that, The step of determining the target noise reduction parameters based on the first noise acoustic signature characteristics of the first noise frequency band includes: Based on the sound data, visual data synchronized with the sound data in time, and motion perception data of the target device synchronized with the sound data in time, the target sound source location of the first voice command is determined, wherein the visual data includes the target object that issued the first voice command. Based on the location of the target sound source and the voiceprint characteristics of the target object, the frequency band of the second noise is determined, wherein the location of the second noise source is different from the location of the target sound source; Based on the frequency band of the second noise, the noise type of the second noise is determined, and the noise type includes at least one of steady-state noise, transient noise, and impulse noise; The target noise reduction parameters are determined by using preset noise countermeasure strategies corresponding to the noise type of the second noise and the voiceprint features of the first noise, respectively. The target noise reduction parameters are also used to reduce the noise of the second noise.
4. The method according to claim 3, characterized in that, Determining the target sound source location corresponding to the first voice command based on the sound data, visual data synchronized in time with the sound data, and motion perception data of the target device synchronized in time with the sound data includes: In the coordinate system of the target device determined based on the motion perception data, the head position and lip movement data of each candidate object in the visual data, and the initial sound source position of the first voice command in the sound data are determined. Based on the first voiceprint feature in the sound data, the individual attribute features of the target object are determined, and the individual attribute features include at least one of gender features and age features; Based on the individual attribute features, the initial sound source location, the head position of each candidate object, and the lip movement data of each candidate object, the target object is determined from the candidate objects; The target sound source location is determined based on the initial sound source location and the head position of the target object in the visual data.
5. The method according to claim 3, characterized in that, The step of determining the target noise reduction parameters by employing preset noise countermeasure strategies corresponding to the noise type of the second noise and the voiceprint features of the first noise, respectively, includes: Initial noise reduction parameters are determined by using preset noise countermeasures corresponding to the noise type of the second noise and the voiceprint features of the first noise, respectively. Based on the operating sound data of the target device, determine the noise reduction parameter adjustment parameters for the operating sound data; The initial noise reduction parameters are adjusted using the noise reduction parameter adjustment parameters to obtain the target noise reduction parameters.
6. The method according to claim 1, characterized in that, Also includes: Based on the voiceprint features of the first voice command and the mapping relationship between voiceprint features and voiceprint permissions, the first voiceprint permission of the voiceprint features of the first voice command is determined. If the first voiceprint permission meets the voiceprint permission conditions required to execute the first voice command, then the first operation corresponding to the first voice command is executed based on the first voice command. Alternatively, if the first voiceprint permission meets the voiceprint permission conditions required to execute the first voice command, and the target object in the visual data passes the permission verification, then the first operation corresponding to the first voice command is executed based on the first voice command.
7. The method according to claim 1, characterized in that, Also includes: If multiple voice commands are received within a preset time period, the identity of the object corresponding to each voice command is determined based on the voiceprint features of each voice command. According to the object identity, the voice command timing data corresponding to each object identity is cached respectively, and the voice command timing data is used for multi-turn dialogue context association analysis; Upon receiving a new voice command, the identity of the object corresponding to the new voice command is determined based on the voiceprint characteristics of the new voice command; If the object identity corresponding to the new voice command belongs to the object identity of the cached voice command timing data, then based on the voice command timing data of the object identity corresponding to the new voice command, the association relationship between the new voice command and the historical commands in the voice command timing data is determined. Based on the new voice command and the association relationship, execute the third operation corresponding to the new voice command.
8. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-7.
10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method described in any one of claims 1-7.