Robot beam locking and speech separation method, device, equipment and storage medium
By acquiring multi-channel voice signals and robot state information, calculating beamwidth and performing motion compensation, the problem of inaccurate target sound source recognition in robot voice interaction is solved, improving voice recognition accuracy and user experience, and is suitable for robot voice interaction in complex environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN MINRRAY IND CORP LTD
- Filing Date
- 2026-05-20
- Publication Date
- 2026-07-21
Smart Images

Figure CN122435944A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robot voice interaction technology, and in particular to a robot beamlocking and voice separation method, apparatus, device and storage medium. Background Technology
[0002] With the rapid development of robotics technology, the demand for intelligent and precise human-computer interaction is increasing, especially in complex environments with multiple sound sources and strong noise. Accurate recognition, targeting, and separation of target user voice by robots has become crucial for improving the interactive experience. Most existing robot voice interaction technologies employ omnidirectional sound pickup, which makes it difficult to effectively distinguish between the target user and interfering sound sources, leading to problems such as low voice recognition accuracy, unstable beam locking, and chaotic response during multi-user interaction. Summary of the Invention
[0003] This application provides a method, apparatus, device, and storage medium for robot beam locking and voice separation, which can solve at least one of the technical problems in the background art to a certain extent.
[0004] To achieve the above objectives, this application adopts the following technical solution: In a first aspect, a method for robot beam locking and voice separation is provided, the method comprising: Acquire multi-channel voice signals and robot status information; Based on the multi-channel speech signal, the distance information d and azimuth information of the target sound source are determined; Based on the distance information d and the objective function model, the beamwidth W is calculated. The objective function model is: W = 2° + 0.5° × (d / 1m); Based on the beamwidth W, the distance information, and the orientation information, motion compensation is performed in conjunction with the robot state information to obtain the beam parameters of the target beam. The initial speech signal is acquired based on the target beam, and the initial speech signal is processed to obtain the target speech signal; The robot is controlled based on the target voice signal.
[0005] Secondly, a robot beamlocking and voice separation device is provided, comprising: The acquisition module is used to acquire multi-channel voice signals and robot status information; The determination module is used to determine the distance information d and azimuth information of the target sound source based on the multi-channel speech signal; The calculation module is used to calculate the beamwidth W based on the distance information d and the objective function model, wherein the objective function model is: W = 2° + 0.5° × (d / 1m); The motion compensation module is used to perform motion compensation based on the beamwidth W, the distance information, and the orientation information, combined with the robot state information, to obtain the beam parameters of the target beam. The processing module is used to acquire an initial speech signal based on the target beam and process the initial speech signal to obtain a target speech signal; The control module is used to control the robot based on the target voice signal.
[0006] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the robot beamlocking and voice separation method as described in any one of the first aspects above.
[0007] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the robot beamlocking and voice separation method as described in any one of the first aspects above.
[0008] Fifthly, embodiments of this application provide a computer program product that, when run on an electronic device, causes the electronic device to execute the robot beamlocking and voice separation method described in any of the first aspects above.
[0009] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here.
[0010] In this embodiment, firstly, multi-channel speech signals and robot state information are acquired. Based on the multi-channel speech signals, the distance information d and azimuth information of the target sound source are determined. Based on the distance information d and the objective function model, the beamwidth W is calculated. The objective function model is: W = 2° + 0.5° × (d / 1m). Based on the beamwidth W, the distance information, and the azimuth information, motion compensation is performed in conjunction with the robot state information to obtain the beam parameters of the target beam. Based on the target beam, an initial speech signal is acquired and processed to obtain the target speech signal. Based on the target speech signal, the robot is controlled. This solves the problems of existing robot voice control being susceptible to interference and inaccurate following, improving voice recognition accuracy and user experience. It is suitable for robot voice interaction scenarios in indoor, multi-person, high-noise, and high-reverberation environments, forming a fully acoustic closed-loop collaboration. The quality of the target voice output by the beam dynamic locking module directly affects the accuracy of the triple fusion verification; the verification results in turn adjust the sensitivity of beam switching. The separation results of the multi-source separation engine (such as whether there is an interruption) serve as direct input for state management, and the current state of the state management module (such as "listening" or "switching") determines whether the beam re-locks to a new target. There are non-obvious interdependencies and real-time feedback among the modules, which together solve the three-layer coupling dilemma that cannot be addressed by a single technology or simple combination.
[0011] The above description is merely an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below. Attached Figure Description
[0012] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiments below. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 This is a flowchart illustrating the robot beam locking and voice separation method provided in an embodiment of this application; Figure 2 System architecture diagram of the multi-channel voice interaction system provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of the robot beam locking and voice separation device provided in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0013] The embodiments of the technical solutions of this application will now be described in detail with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of this application, and are therefore merely examples and should not be used to limit the scope of protection of this application. When the following description relates to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. Various changes, modifications, and equivalents of the methods, apparatus, and / or systems described herein will become apparent upon understanding this disclosure. For example, the order of operations described herein is merely illustrative and is not limited to those orders set forth herein, but can be changed as will become apparent upon understanding this disclosure, except for operations that must be performed in a specific order. Furthermore, for clarity and conciseness, descriptions of features known in the art may be omitted.
[0014] The embodiments described in the following examples of this disclosure are not representative of all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0015] The following describes the scenarios involved in the embodiments of this application. In near-field natural dialogue scenarios such as service robots, users are often in the same direction (e.g., sitting side by side), alternating, or interrupting states, and the environment often has visual limitations (backlighting, occlusion) or privacy requirements. In near-field multi-person robot interaction scenarios, the following problems may occur: 1. The contradiction between the limited spatial resolution of the beam and the aliasing of the same-direction sound sources leads to a low input signal-to-noise ratio for the separation algorithm.
[0016] 2. In environments with densely overlapping sound sources, the reliability of identification based on a single biometric feature (such as voiceprint) decreases.
[0017] 3. There is a huge gap between the transience of acoustic events and the coherence of dialogue intentions, making it difficult for interaction logic to be synchronized with physical signal events in real time.
[0018] Currently, even when combining existing beamforming technology, speech separation algorithms, single voiceprint recognition, and simple state machines, the lack of collaborative optimization among the various layers of technology means that usable human-computer dialogue interaction can only be achieved in near-field multi-person interaction scenarios characterized by "near field, unidirectional, multi-person, and no vision."
[0019] To address this, this application provides a robot beamlocking and speech separation method. Under the constraints of "no visual conditions," "target locking and separation coordination for sound sources in the same direction within the beam," and "acoustic event-driven multi-state interaction that adapts to natural dialogue and interjections among multiple people," this method solves the problems of robot voice control being easily interfered with and inaccurate following, thereby improving the accuracy of speech recognition and user experience.
[0020] See Figure 1 This is a flowchart illustrating the robot beam locking and voice separation method provided in this application embodiment. The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0021] like Figure 1 As shown, the robot beam locking and voice separation method provided in this embodiment includes the following steps: Step 101: Acquire multi-channel voice signals and robot status information.
[0022] The robot's state information may include robot posture, movement speed, body vibration information, etc., which are not limited here.
[0023] Among them, multi-channel voice signal refers to multiple voice data collected synchronously by multiple microphone units in a microphone array at the same time and in the same spatial environment.
[0024] In acquiring multi-channel voice signals, a ring microphone array can be used to collect the voice signals.
[0025] For example, the relevant settings parameters for a circular microphone array could be: 16 channels, 8cm diameter, 48kHz, 24bit, 360° coverage, and positioning accuracy ±1°, without any specific limitations.
[0026] Optionally, multi-channel voice signals and robot state information can be obtained by calling the function "beam_lock_voice_separation(mic_array_data, robot_pose)".
[0027] Here, mic_array_data can be the raw data of a 16-channel microphone array (48kHz, 24bit), and robot_pose can be the robot's current pose, which is not limited here.
[0028] Step 102: Based on the multi-channel speech signal, determine the distance information d and the orientation information of the target sound source.
[0029] Specifically, the first step is to extract acoustic features from the multi-channel speech signal, including both time-domain and frequency-domain features, such as energy, zero-crossing rate, and MFCC. Then, the extracted acoustic features can be used for sound source localization to obtain the location and distance information of the target sound source.
[0030] The distance information ranges from 0.3 to 5 meters, with a distance error of less than 5 centimeters.
[0031] Optionally, the environmental noise baseline is dynamically updated with a period of 100 milliseconds, and the noise suppression depth is adaptively adjusted according to the environmental noise baseline to ensure that the noise suppression amount is not less than 35 dB.
[0032] Step 103: Based on the distance information d and the objective function model, calculate the beamwidth W. The objective function model is: W = 2° + 0.5° × (d / 1m).
[0033] It should be noted that when calculating the beamwidth based on distance information, it can be calculated using a pre-built objective function model. The pre-built objective function model can be: W = f(d) = 2° + 0.5° × (d / 1m) Where W is the beamwidth and d is the distance to the sound source. When d < 1m, an ultra-narrow beam is formed (W ≤ 2°).
[0034] For example, when d = 0.5m (near field), the beamwidth W = 2° + 0.5° × 0.5 = 2.25°, which meets the requirement of ultra-narrow beamwidth in the near field. When d = 5m (far field), the beamwidth W = 2° + 0.5° × 5 = 4.5°, achieving moderate beamwidth in the far field.
[0035] Step 104: Based on the beamwidth W, distance information, and orientation information, motion compensation is performed in conjunction with the robot's state information to obtain the beam parameters of the target beam.
[0036] The beam parameters include one or more of the following: beam pointing angle, beamwidth, beam gain, and microphone channel weighting coefficient.
[0037] Understandably, by combining the robot's current posture information, motion compensation can be performed on the beam pointing to correct beam offset caused by changes in robot posture, ultimately calculating the beam parameters of the target beam. These beam parameters include the corrected beam pointing angle, beamwidth, beam gain, and microphone channel weighting coefficients.
[0038] Step 105: Acquire the initial speech signal based on the target beam and process the initial speech signal to obtain the target speech signal.
[0039] Specifically, beamforming processing can be performed on multi-channel speech signals based on the beam parameters of the target beam to obtain an initial speech signal after directional enhancement. Beamforming processing involves delaying, weighting, and superimposing the multi-channel speech signals to create a pickup gain in the direction of the target sound source and attenuation in non-target directions, thereby extracting the initial speech signal from the multi-channel mixed signal.
[0040] Specifically, the process first determines whether the initial speech signal contains the target user's speech signal. If the initial speech signal contains the target user's speech signal, the initial speech signal is separated into target speech signal and non-target speech signal. The non-target speech signal is the speech signal from a user other than the target user.
[0041] The process of determining whether the initial voice signal contains the target user's voice signal includes: Based on the initial speech signal, determine the voiceprint confidence value C1, rhythm confidence value C2, and semantic confidence value C3; The target confidence value C is calculated using the target weighting formula: C = 0.4×C1 + 0.3×C2 + 0.3×C3; When the target confidence value C ≥ 0.8, it is determined that the initial voice signal contains the target user's voice signal.
[0042] One possible approach is to first determine the voiceprint features, speech rhythm features, and speech recognition results of the initial speech signal based on the initial speech signal, and then determine the voiceprint confidence value based on the matching degree between the voiceprint features and the target voiceprint.
[0043] Among them, voiceprint features can be represented as high-dimensional feature vectors, containing key information such as the frequency, amplitude, and spectral distribution of the user's pronunciation. They are the voice physiological features that can uniquely represent the target user.
[0044] In this step, a voiceprint feature extraction algorithm (such as Mel frequency cepstral coefficients (MFCC)) can be used to extract features from the preprocessed effective speech segments to obtain voiceprint features.
[0045] Among them, speaking rhythm characteristics refer to the behavioral characteristics of users when they speak, such as speaking speed, pause intervals, and sentence length distribution. There are significant differences in speaking rhythm among different users.
[0046] In this step, the speaking rhythm features can be extracted by analyzing the temporal distribution characteristics of effective speech segments. Specifically, this includes calculating the user's average speaking speed (number of syllables and words per unit time), the number of pauses during the speaking process, the average pause duration, the average sentence length (number of syllables and words per sentence), and the pause duration between sentences. These extracted parameters are then processed in a structured manner to obtain the speaking rhythm features.
[0047] Optionally, a speech recognition algorithm can be used to recognize the preprocessed valid speech segments, convert the speech signal into text information, and obtain the speech recognition result of the initial speech signal.
[0048] Furthermore, the voiceprint confidence value can be determined based on the matching degree between the voiceprint features and the target voiceprint.
[0049] Among them, the target voiceprint is the voiceprint feature of the target user that has been collected and stored in advance.
[0050] In this step, the matching degree between the voiceprint features and the target voiceprint is first calculated, for example, using the cosine similarity algorithm. The cosine similarity algorithm is suitable for matching high-dimensional feature vectors and is preferred as the matching degree calculation algorithm in this embodiment.
[0051] The matching degree can range from [0,1]. The closer the matching degree is to 1, the higher the similarity between the voiceprint feature and the target voiceprint, and the more likely it is to be the voice of the target user.
[0052] Among them, the voiceprint confidence value is a quantitative representation of the voiceprint matching degree, which is used to reflect the credibility of the voiceprint feature matching the target voiceprint.
[0053] In this step, the voiceprint matching degree can be directly used as the voiceprint confidence value, or the voiceprint confidence value can be obtained after correcting the matching degree through a preset mapping relationship.
[0054] Furthermore, the rhythm confidence value can be determined based on the degree of matching between the speech rhythm characteristics and the pre-recorded target rhythm.
[0055] Among them, the target rhythm is the speaking rhythm characteristics of the target user that have been collected and stored in advance.
[0056] In this step, the matching degree between the speech rhythm features extracted in step 201 and the target rhythm is first calculated. The matching degree can be calculated using a feature vector matching algorithm. The matching degree is calculated separately for each parameter of the speech rhythm features (speech rate, pause interval, sentence length, etc.), and then a weighted average is calculated to obtain the final speech rhythm matching degree. For example, the weights of the speech rate matching degree, pause interval matching degree, and sentence length matching degree are set to 0.4, 0.3, and 0.3, respectively, and then the weighted average is calculated as the final matching degree between the speech rhythm features and the target rhythm.
[0057] The rhythm confidence value ranges from [0,1]. The higher the matching degree, the higher the rhythm confidence value, reflecting the higher the credibility of the matching between the speaking rhythm and the target user.
[0058] Furthermore, the semantic confidence value can be determined based on the topic relevance between the speech recognition results and the historical dialogues.
[0059] Among them, historical dialogues refer to the past dialogue records between the target user and the robot. Historical dialogues need to be stored in advance and bound to the target user.
[0060] In this step, the first step is to extract topics from the speech recognition results and historical dialogue text. Topic extraction can employ topic models, keyword extraction algorithms, etc., to extract the core keywords and core topics from the speech recognition results, as well as the core topics and high-frequency keywords from the historical dialogue. Then, the topic relevance between the speech recognition results and the historical dialogue is calculated and normalized as the semantic confidence score.
[0061] Among them, semantic confidence is calculated based on the topic relevance between the current speech recognition result and the historical dialogue context, and is used to determine whether the current speech is consistent with the established dialogue main line.
[0062] Then, the target confidence value is calculated by weighting the voiceprint confidence value C1, the rhythm confidence value C2, and the semantic confidence value C3.
[0063] As an example, the target confidence score can be calculated using the following formula: C = 0.4 × C1 + 0.3 × C2 + 0.3 × C3 Where C is the target confidence value, C1 is the voiceprint confidence value, C2 is the rhythm confidence value, and C3 is the semantic confidence value.
[0064] Specifically, when the target confidence value C ≥ 0.8, the initial voice signal is determined to contain the target user's voice signal.
[0065] It should be noted that by simultaneously extracting the voiceprint features and speaking rhythm features of the initial speech signal, and combining them with the semantic relevance of the speech recognition results, verification is performed from three dimensions: user identity features (voiceprint), behavioral features (speaking rhythm), and content features (semantics). This avoids the limitations of single feature verification, effectively reduces the impact of environmental noise, pronunciation changes, emotional fluctuations, and other factors on the verification results, and improves the recognition accuracy.
[0066] Optionally, when multiple sound sources in the same direction are detected within the range of the target beam, the target speech signal is initially separated using a fast independent component analysis algorithm to obtain the initially separated sound signal. The initially separated sound signal is then input into the U-Net network for residual optimization to obtain the target speech signal and non-target speech signal. Finally, noise suppression processing is performed on the target speech signal to adaptively adjust the suppression depth according to the dynamically updated environmental noise baseline, resulting in the enhanced target speech signal.
[0067] Specifically, the Fast Independent Component Analysis (FastICA) algorithm is first used to separate the initial speech signal. The FastICA algorithm can separate individual sound source signals from a mixed speech signal even when the number of sound sources and propagation paths are unknown. It is suitable for the initial separation of multiple sound sources traveling in the same direction and can effectively distinguish the speech signals of different users, obtaining a preliminarily separated sound signal.
[0068] Optionally, the initially separated sound signal can be obtained using the following formula: Y = W·S, where W is the separation matrix, S is the mixed signal within the beam (initial speech signal), and Y is the initially separated sound signal.
[0069] Furthermore, the initially separated audio signal can be input into the U-Net network for residual optimization. The U-Net network has excellent feature extraction and semantic segmentation capabilities. It can extract features of the audio signal through the encoder, restore signal details through the decoder, and, combined with the residual connection structure, effectively compensate for problems such as signal distortion and incomplete separation during the initial separation process, accurately distinguishing target speech signals from non-target speech signals.
[0070] Furthermore, noise suppression processing can be applied to the target speech signal obtained after residual optimization to further improve the purity of the target speech signal. In this step, an adaptive noise suppression algorithm is adopted. First, the ambient noise acquisition module collects the current ambient noise signal in real time and dynamically updates the ambient noise benchmark. Then, based on the dynamically updated ambient noise benchmark, the noise suppression depth is adaptively adjusted.
[0071] Step 106: Control the robot based on the target voice signal.
[0072] One possible approach is to first perform speech recognition on the target speech signal to obtain the speech recognition result, and then generate a response command based on the speech recognition result. The robot is then controlled to execute the corresponding interactive action based on the response command, and the robot's servo motor is controlled to rotate following the beam direction of the target beam.
[0073] Specifically, speech recognition can be performed on the target speech signal to obtain accurate results. At this stage, the target speech signal has been filtered out for non-target user speech and environmental noise interference, resulting in a higher recognition accuracy than direct recognition of the initial speech signal. Then, based on the speech recognition results and pre-defined command mapping rules, executable response commands for the robot are generated. For example, if the speech recognition result is "forward," a response command of "control the robot drive module to start and move forward" is generated. This generated response command is then sent to the robot's control module, controlling the robot to execute the corresponding interactive action. Simultaneously, the robot's servo motors rotate to follow the beam direction of the target beam. The beam direction of the target beam reflects the target user's position in real time, as indicated by the servo motor rotation.
[0074] Optionally, it can also recognize hand gestures such as waving by detecting changes in airflow, without the need for visual assistance.
[0075] If the signal strength of the new sound source signal is ≥30% greater than the signal strength of the target sound source, and the duration exceeds 100ms, interruption detection is triggered to detect the pause time after the new sound source speaks. If the pause time T after the new sound source speaks is greater than 500ms, the current state is switched to the switching state, the locked target sound source is updated to the new sound source, and the robot is controlled to perform actions such as parameter storage, turning to the new sound source, and content recording.
[0076] If the pause time T after the new sound source speaks is less than or equal to 500ms and the original locked target sound source is restored, then the current state will be returned to the locked state, indicating that only other speakers are interrupting, and the robot will maintain the locked target sound source and beam parameters.
[0077] Optionally, when the current state is the response state, if the response execution ends and the duration of no new sound source input reaches the target time, the robot will switch the current state to the standby state and perform parameter reset and sound source detection preparation actions.
[0078] The target time can be 1 second. If the response execution ends and there is no new sound source input for a duration of 1 second, the robot will switch its current state to standby state and perform parameter reset and sound source detection preparation actions.
[0079] Optionally, when the current state is standby, if a valid sound source is detected and the sound source intensity is greater than a preset intensity threshold, the robot is controlled to transition to the locked state and perform sound source localization and verification actions to complete target locking.
[0080] The above solution can transform dialogue logic into acoustic physical events, enabling natural and smooth interaction in scenarios involving multiple people taking turns and interrupting.
[0081] In one test embodiment of the present invention, the test scenario involved 2-3 people speaking in the same direction near the robot (approximately 1 meter away, with about 40% overlap in their speech content) at an ambient noise level of approximately 60 dB. Using the solution of the present invention, the accuracy of separating the three sound sources in the same direction reached 89.2%, while the traditional MVDR beamforming solution achieved 51%. The state switching delay was reduced from 220 ms in the traditional solution to 85 ms, improving the smoothness of interaction by over 60%.
[0082] In this embodiment, firstly, multi-channel speech signals and robot state information are acquired. Based on the multi-channel speech signals, the distance information d and azimuth information of the target sound source are determined. Based on the distance information d and the objective function model, the beamwidth W is calculated. The objective function model is: W = 2° + 0.5° × (d / 1m). Based on the beamwidth W, the distance information, and the azimuth information, motion compensation is performed in conjunction with the robot state information to obtain the beam parameters of the target beam. Based on the target beam, an initial speech signal is acquired and processed to obtain the target speech signal. Based on the target speech signal, the robot is controlled. This solves the problems of existing robot voice control being susceptible to interference and inaccurate following, improving voice recognition accuracy and user experience. It is suitable for robot voice interaction scenarios in indoor, multi-person, high-noise, and high-reverberation environments, forming a fully acoustic closed-loop collaboration. The quality of the target voice output by the beam dynamic locking module directly affects the accuracy of the triple fusion verification; the verification results in turn adjust the sensitivity of beam switching. The separation results of the multi-source separation engine (such as whether there is an interruption) serve as direct input for state management, and the current state of the state management module (such as "listening" or "switching") determines whether the beam re-locks to a new target. There are non-obvious interdependencies and real-time feedback among the modules, which together solve the three-layer coupling dilemma that cannot be addressed by a single technology or simple combination.
[0083] Figure 2 This is a system architecture diagram of a multi-channel voice interaction system. The system is arranged in a hierarchical architecture of perception-processing-decision-execution, with five functional modules: The first module is the multi-channel acoustic perception layer, which collects audio signals through a 16-channel microphone array and completes noise benchmark updates after acoustic feature extraction; the second module is the beam dynamic locking module, which performs beam switching after triple verification of C1, C2, and C3 based on the sound source tracking algorithm W=f(d); the third module is the multi-source separation engine, which processes the signal using a FastICA and U-Net cascade method to achieve noise suppression; the fourth module is the interaction state management center, which has a built-in five-state model of lock-listen-response-switch-standby, and is configured with priority and semantic memory units; the fifth module is the voice interaction execution layer, which includes two units: speech recognition and response, and action collaborative feedback, to complete the final voice interaction execution.
[0084] Corresponding to the robot beam locking and voice separation method described in the above embodiments, Figure 3 This is a structural block diagram of the robot beam locking and voice separation device provided in the embodiments of this application.
[0085] Reference Figure 3 The robot beam-locking and voice separation device 300 includes: The acquisition module 310 is used to acquire multi-channel voice signals and robot status information; The determination module 320 is used to determine the distance information d and azimuth information of the target sound source based on the multi-channel speech signal; The calculation module 330 is used to calculate the beamwidth W based on the distance information d and the objective function model, wherein the objective function model is: W = 2° + 0.5° × (d / 1m); The motion compensation module 340 is used to perform motion compensation based on the beamwidth W, the distance information, and the orientation information, combined with the robot state information, to obtain the beam parameters of the target beam. The processing module 350 is used to acquire an initial speech signal based on the target beam and process the initial speech signal to obtain a target speech signal; The control module 360 is used to control the robot based on the target voice signal.
[0086] Optionally, the processing module includes: The judgment unit is used to determine whether the initial voice signal contains the voice signal of the target user; if it is determined that the initial voice signal contains the voice signal of the target user, the initial voice signal is subjected to voice separation to obtain the target voice signal and the non-target voice signal, wherein the non-target voice signal is the voice signal of a user other than the target user.
[0087] Optionally, the determination unit is specifically used for: Based on the initial speech signal, determine the voiceprint confidence value C1, the rhythm confidence value C2, and the semantic confidence value C3; The target confidence value C is calculated according to the target weighting formula, which is: C = 0.4×C1 + 0.3×C2 + 0.3×C3; When the target confidence value C ≥ 0.8, it is determined that the initial voice signal contains the voice signal of the target user.
[0088] Optionally, the determination unit is specifically used for: When multiple sound sources in the same direction are detected within the range of the target beam, the target speech signal is initially separated using a fast independent component analysis algorithm to obtain the initially separated sound signal. The preliminarily separated audio signal is input into the U-Net network, and residual optimization processing is performed to obtain the target speech signal and the non-target speech signal. The target speech signal is subjected to noise suppression processing to adaptively adjust the suppression depth according to the dynamically updated environmental noise benchmark, thereby obtaining an enhanced target speech signal.
[0089] Optionally, the control module is further configured to: The target speech signal is subjected to speech recognition to obtain the speech recognition result; Based on the speech recognition results, a response command is generated to control the robot to perform corresponding interactive actions based on the response command. The servo motor controlling the robot rotates in accordance with the beam direction of the target beam.
[0090] Optionally, if the signal strength of the new sound source signal is ≥30% greater than the signal strength of the target sound source, and the duration exceeds 100ms, interruption detection is triggered to detect the pause time after the new sound source speaks. If the pause time T after the new sound source speaks is greater than 500ms, then the current state is switched to the switching state, and the locked target sound source is updated to the new sound source; If the pause time T after the new sound source speaks is less than or equal to 500ms and the original locked target sound source is restored, then the current state will be returned to the locked state.
[0091] Optionally, the control module 350 is further configured to: When the current state is the response state, if the response execution ends and the duration of no new sound source input reaches the target time, the robot is controlled to switch the current state to the standby state and perform parameter reset and sound source detection preparation actions. When the current state is standby, if a valid sound source is detected and the sound source intensity is greater than a preset intensity threshold, the robot is controlled to switch to the locked state and perform sound source localization and verification actions to complete target locking.
[0092] In this embodiment, firstly, multi-channel speech signals and robot state information are acquired. Based on the multi-channel speech signals, the distance information d and azimuth information of the target sound source are determined. Based on the distance information d and the objective function model, the beamwidth W is calculated. The objective function model is: W = 2° + 0.5° × (d / 1m). Based on the beamwidth W, the distance information, and the azimuth information, motion compensation is performed in conjunction with the robot state information to obtain the beam parameters of the target beam. Based on the target beam, an initial speech signal is acquired and processed to obtain the target speech signal. Based on the target speech signal, the robot is controlled. This solves the problems of existing robot voice control being susceptible to interference and inaccurate following, improving voice recognition accuracy and user experience. It is suitable for robot voice interaction scenarios in indoor, multi-person, high-noise, and high-reverberation environments, forming a fully acoustic closed-loop collaboration. The quality of the target voice output by the beam dynamic locking module directly affects the accuracy of the triple fusion verification; the verification results in turn adjust the sensitivity of beam switching. The separation results of the multi-source separation engine (such as whether there is an interruption) serve as direct input for state management, and the current state of the state management module (such as "listening" or "switching") determines whether the beam re-locks to a new target. There are non-obvious interdependencies and real-time feedback among the modules, which together solve the three-layer coupling dilemma that cannot be addressed by a single technology or simple combination.
[0093] in addition, Figure 3 The robot beamlocking and voice separation device shown can be a software unit, a hardware unit, or a combination of software and hardware built into existing electronic devices. It can also be integrated into electronic devices as a separate accessory, or exist as a standalone electronic device.
[0094] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0095] Figure 4This is a schematic diagram of the structure of the electronic device provided in an embodiment of this application. For example... Figure 4 As shown, the electronic device 5 of this embodiment includes: at least one processor 50 ( Figure 4 (Only one is shown in the diagram) a processor, a memory 51, and a computer program 52 stored in the memory 51 and executable on the at least one processor 50, wherein the processor 50 executes the computer program 52 to implement the steps in any of the above embodiments of the robot beamlocking and voice separation method.
[0096] The electronic device may be a desktop computer, laptop, handheld computer, or cloud server, etc. This electronic device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that... Figure 4 This is merely an example of electronic device 5 and does not constitute a limitation on electronic device 5. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.
[0097] The processor 50 may be a central processing unit, or it may be other general-purpose processors, digital signal processors, application-specific integrated circuits, off-the-shelf programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0098] In some embodiments, the memory 51 may be an internal storage unit of the electronic device 5, such as a hard disk or memory of the electronic device 5. In other embodiments, the memory 51 may be an external storage device of the electronic device 5, such as a plug-in hard disk, smart memory card, secure digital card, flash memory card, etc., equipped on the electronic device 5. Further, the memory 51 may include both internal storage units and external storage devices of the electronic device 5. The memory 51 is used to store operating systems, applications, boot loaders, data, and other programs, such as the program code of the computer program. The memory 51 can also be used to temporarily store data that has been output or will be output.
[0099] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the steps in the above-described method embodiments.
[0100] This application provides a computer program product that, when run on an electronic device, enables the electronic device to implement the steps described in the various method embodiments above.
[0101] If the integrated unit is implemented as a software functional unit and used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a device / electronic device, a recording medium, a computer memory, a read-only memory, a random access memory, an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks.
[0102] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0103] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0104] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.
[0105] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0106] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for robot beam locking and voice separation, characterized in that, include: Acquire multi-channel voice signals and robot status information; Based on the multi-channel speech signal, the distance information d and azimuth information of the target sound source are determined; Based on the distance information d and the objective function model, the beamwidth W is calculated. The objective function model is: W = 2° + 0.5° × (d / 1m); Based on the beamwidth W, the distance information, and the orientation information, motion compensation is performed in conjunction with the robot state information to obtain the beam parameters of the target beam. The initial speech signal is acquired based on the target beam, and the initial speech signal is processed to obtain the target speech signal; The robot is controlled based on the target voice signal.
2. The method according to claim 1, characterized in that, The step of processing the initial speech signal to obtain the target speech signal includes: Determine whether the initial voice signal contains the target user's voice signal; If it is determined that the initial speech signal contains the speech signal of the target user, the initial speech signal is subjected to speech separation to obtain the target speech signal and the non-target speech signal, wherein the non-target speech signal is the speech signal of a user other than the target user.
3. The method according to claim 2, characterized in that, The step of determining whether the initial voice signal contains the target user's voice signal includes: Based on the initial speech signal, determine the voiceprint confidence value C1, the rhythm confidence value C2, and the semantic confidence value C3; The target confidence value C is calculated according to the target weighting formula, which is: C = 0.4×C1 + 0.3×C2 + 0.3×C3; When the target confidence value C ≥ 0.8, it is determined that the initial voice signal contains the voice signal of the target user.
4. The method according to claim 2, characterized in that, The step of performing speech separation on the initial speech signal to obtain the target speech signal and non-target speech signals includes: When multiple sound sources in the same direction are detected within the range of the target beam, the target speech signal is initially separated using a fast independent component analysis algorithm to obtain the initially separated sound signal. The preliminarily separated audio signal is input into the U-Net network, and residual optimization processing is performed to obtain the target speech signal and the non-target speech signal. The target speech signal is subjected to noise suppression processing to adaptively adjust the suppression depth according to the dynamically updated environmental noise benchmark, thereby obtaining an enhanced target speech signal.
5. The method according to claim 1, characterized in that, The control of the robot based on the target voice signal includes: The target speech signal is subjected to speech recognition to obtain the speech recognition result; Based on the speech recognition results, a response command is generated to control the robot to perform corresponding interactive actions based on the response command. The servo motor controlling the robot rotates in accordance with the beam direction of the target beam.
6. The method according to any one of claims 1-5, characterized in that, Also includes: If the signal strength of the new sound source signal is ≥30% greater than the signal strength of the target sound source, and the duration exceeds 100ms, interruption detection is triggered to detect the pause time after the new sound source speaks. If the pause time T after the new sound source speaks is greater than 500ms, then the current state is switched to the switching state, and the locked target sound source is updated to the new sound source; If the pause time T after the new sound source speaks is less than or equal to 500ms and the original locked target sound source is restored, then the current state will be returned to the locked state.
7. The method according to claim 6, characterized in that, Also includes: When the current state is the response state, if the response execution ends and the duration of no new sound source input reaches the target time, the robot is controlled to switch the current state to the standby state and perform parameter reset and sound source detection preparation actions. When the current state is standby, if a valid sound source is detected and the sound source intensity is greater than a preset intensity threshold, the robot is controlled to switch to the locked state and perform sound source localization and verification actions to complete target locking.
8. A robot beamlocking and voice separation device, characterized in that, include: The acquisition module is used to acquire multi-channel voice signals and robot status information; The determination module is used to determine the distance information d and azimuth information of the target sound source based on the multi-channel speech signal; The calculation module is used to calculate the beamwidth W based on the distance information d and the objective function model, wherein the objective function model is: W = 2° + 0.5° × (d / 1m); The motion compensation module is used to perform motion compensation based on the beamwidth W, the distance information, and the orientation information, combined with the robot state information, to obtain the beam parameters of the target beam. The processing module is used to acquire an initial speech signal based on the target beam and process the initial speech signal to obtain a target speech signal; The control module is used to control the robot based on the target voice signal.
9. An electronic device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the steps of the method as claimed in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, which is loaded by a processor to perform the steps of the method according to any one of claims 1 to 7.