Speech recognition method and device, electronic equipment, storage medium and program product

By combining image and audio information, using microphone arrays and computer vision algorithms, the sound source and object location are determined, solving the accuracy problem of voice interaction in a multi-device coexistence environment, and achieving accurate voice recognition and device wake-up.

CN120612932APending Publication Date: 2025-09-09BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510518738.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

In a smart home environment where multiple devices coexist, how to accurately wake up the target device for voice interaction? Existing technologies make it difficult to achieve accurate voice recognition and device wake-up in complex environments.

Method used

By collecting the image and audio information of the target object, using a microphone array combined with computer vision and direction of arrival estimation algorithm, the sound source position and object position are determined, thereby separating the target object's voice information from multiple sound source audio information for recognition.

Benefits of technology

It achieves precise tracking of target objects and precise collection of voice information in complex environments, improves the accuracy and robustness of voice interaction, and can accurately wake up smart devices for voice interaction in an environment where multiple devices coexist.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612932A_ABST
    Figure CN120612932A_ABST
Patent Text Reader

Abstract

The invention provides a voice recognition method and device, electronic equipment, a storage medium and a program product, and relates to the technical field of smart home. The method comprises the following steps: collecting image information including a target object; acquiring audio information through a microphone array in response to a preset posture of the target object included in the image information; the audio information comprises voice information of the target object; and performing voice recognition on the voice information of the target object according to the image information and the audio information. According to the invention, focusing of voice interaction between the target object and the intelligent device can be realized based on the preset posture, the target object is tracked through the microphone array, accurate awakening and accurate voice information collection can be realized in a multi-device coexistence environment, and the voice information of the target object is identified based on image information and audio information. Therefore, the intelligent device is awakened based on the preset posture to trigger acquisition of the audio information, and accurate and smooth voice interaction is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of smart home technology, and in particular to a speech recognition method and device, electronic equipment, storage medium, and program product. Background Art

[0002] With the rapid development of artificial intelligence, especially the maturity of speech recognition and natural language processing technologies, convenient and fast voice interaction has become embedded in a variety of smart home devices and has entered thousands of households. Currently, various smart home devices, including smart speakers, smart TVs, air conditioners, and lighting systems, generally integrate far-field voice wake-up and interaction functions. How to accurately wake up the target device in an environment where multiple devices coexist has become a pressing problem in this field.

[0003] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the Invention

[0004] To overcome the problems existing in the related art, the present disclosure provides a speech recognition method and device, an electronic device, a storage medium and a program product.

[0005] According to a first aspect of an embodiment of the present disclosure, a speech recognition method is provided, comprising: collecting image information including a target object; in response to a preset posture of the target object included in the image information, collecting audio information through a microphone array; the audio information includes speech information of the target object; and performing speech recognition on the speech information of the target object based on the image information and the audio information.

[0006] In some exemplary embodiments of the present disclosure, the performing speech recognition on the voice information of the target object based on the image information and the audio information includes: determining, based on the audio information, multiple sound source audio information corresponding to multiple sound sources respectively; determining the sound source positions corresponding to the multiple sound sources respectively; determining the object position of the target object based on the image information; determining the voice information of the target object from the multiple sound source audio information based on the sound source position and the object position; and performing speech recognition on the voice information of the target object.

[0007] In some exemplary embodiments of the present disclosure, determining the multiple sound source audio information corresponding to the multiple sound sources respectively based on the audio information includes: performing Fourier transform on the audio information to obtain audio frequency domain information; and performing blind source separation processing on the audio frequency domain information to determine the multiple sound source audio information corresponding to the multiple sound sources respectively.

[0008] In some exemplary embodiments of the present disclosure, determining the sound source positions corresponding to the multiple sound sources respectively includes: determining the probability of speech existence based on the audio frequency domain information and the multiple sound source audio information corresponding to the multiple sound sources respectively; obtaining the effective frequency point of the corresponding sound source in the audio frequency domain information; determining the components of the multiple sound sources in the audio frequency domain information based on the effective frequency point of the corresponding sound source in the audio frequency domain information and the speech existence probability; and determining the sound source positions corresponding to the multiple sound sources respectively based on the components of the multiple sound sources in the audio frequency domain information.

[0009] In some exemplary embodiments of the present disclosure, the sound source positions corresponding to the multiple sound sources are determined according to the components of the multiple sound sources in the audio frequency domain information, including: obtaining the coordinates of the microphone array and the viewing angle information of collecting the image information; dividing the space corresponding to the image information into multiple partitions according to the coordinates of the microphone array and the viewing angle information of collecting the image information; each of the multiple partitions has point coordinates; determining a steering vector for each of the point coordinates; determining a cross-correlation matrix according to the components of the multiple sound sources in the audio observation signal; determining a noise subspace according to the cross-correlation matrix; determining a spatial energy spectrum according to the noise subspace and the steering vector for each of the point coordinates; and determining the point coordinates corresponding to the maximum value in the spatial energy spectrum as the sound source position corresponding to the sound source in the scene.

[0010] In some exemplary embodiments of the present disclosure, the noise subspace is determined based on the mutual correlation matrix, including: performing eigenvalue decomposition on the mutual correlation matrix to determine a first matrix; the first matrix includes: an eigenvalue matrix and a first eigenvector; the first eigenvector corresponds to the direction of the signal space and the noise space; the eigenvalue represents the energy size in each direction; according to the size of the eigenvalue, the first eigenvector is divided into a signal subspace and a noise subspace; the signal subspace includes the eigenvector corresponding to the maximum eigenvalue; the noise subspace is the remaining eigenvector in the first eigenvector after removing the signal subspace.

[0011] In some exemplary embodiments of the present disclosure, determining the object position of the target object based on the image information includes: if the image information includes a preset posture of the target object, then determining the object position of the target object in the space corresponding to the image information based on the image information.

[0012] In some exemplary embodiments of the present disclosure, the voice information of the target object is determined from the multiple sound source audio information based on the sound source position and the object position, including: if the sound source position and the object position match, the sound source corresponding to the matched sound source position is determined as the target sound source; and based on the target sound source, the voice information of the target object is determined from the multiple sound source audio information respectively corresponding to the multiple sound sources.

[0013] In some exemplary embodiments of the present disclosure, the method further includes: determining a control instruction based on voice recognition of voice information of the target object; and controlling the target device according to the control instruction.

[0014] In some exemplary embodiments of the present disclosure, the method further includes: extracting a voiceprint feature of the target object based on voice information of the target object; and performing voice interaction with the target object based on the voiceprint feature.

[0015] According to a second aspect of an embodiment of the present disclosure, a speech recognition device is provided, comprising: an image information acquisition module for acquiring image information including a target object; an audio information acquisition module for acquiring audio information through a microphone array in response to a preset posture of the target object included in the image information; the audio information includes speech information of the target object; and a speech recognition module for performing speech recognition on the speech information of the target object based on the image information and the audio information.

[0016] In some exemplary embodiments of the present disclosure, the speech recognition module includes: a sound source audio information determination unit, used to determine multiple sound source audio information corresponding to multiple sound sources respectively based on the audio information; a sound source position determination unit, used to determine the sound source positions corresponding to the multiple sound sources respectively; an object position determination unit, used to determine the object position of the target object based on the image information; a target object speech information determination unit, used to determine the speech information of the target object from the multiple sound source audio information based on the sound source position and the object position; and a speech recognition unit, used to perform speech recognition on the speech information of the target object.

[0017] In some exemplary embodiments of the present disclosure, the sound source audio information determination unit is further used to: perform Fourier transform on the audio information to obtain audio frequency domain information; perform blind source separation processing on the audio frequency domain information to determine multiple sound source audio information corresponding to the multiple sound sources respectively.

[0018] In some exemplary embodiments of the present disclosure, the sound source position determination unit is further used to: determine the probability of speech existence based on the audio frequency domain information and the multiple sound source audio information corresponding to the multiple sound sources respectively; obtain the effective frequency point of the corresponding sound source in the audio frequency domain information; determine the components of the multiple sound sources in the audio frequency domain information based on the effective frequency point of the corresponding sound source in the audio frequency domain information and the speech existence probability; determine the sound source positions corresponding to the multiple sound sources respectively based on the components of the multiple sound sources in the audio frequency domain information.

[0019] In some exemplary embodiments of the present disclosure, the sound source position determination unit is further used to: obtain the coordinates of the microphone array and the viewing angle information of collecting the image information; divide the space corresponding to the image information into multiple partitions according to the coordinates of the microphone array and the viewing angle information of collecting the image information; each of the multiple partitions has point coordinates; determine the steering vector of each of the point coordinates; determine the cross-correlation matrix according to the components of the multiple sound sources in the audio observation signal; determine the noise subspace according to the cross-correlation matrix; determine the spatial energy spectrum according to the noise subspace and the steering vector of each of the point coordinates; determine the point coordinates corresponding to the maximum value in the spatial energy spectrum as the sound source position corresponding to the sound source in the scene.

[0020] In some exemplary embodiments of the present disclosure, the sound source position determination unit is further used to: perform eigenvalue decomposition on the mutual correlation matrix to determine a first matrix; the first matrix includes: an eigenvalue matrix and a first eigenvector; the first eigenvector corresponds to the direction of the signal space and the noise space; the eigenvalue represents the energy size in each direction; according to the size of the eigenvalue, the first eigenvector is divided into a signal subspace and a noise subspace; the signal subspace includes the eigenvector corresponding to the maximum eigenvalue; the noise subspace is the remaining eigenvector in the first eigenvector after removing the signal subspace.

[0021] In some exemplary embodiments of the present disclosure, the target object voice information determination unit is further used to: if the sound source position and the object position match, determine the sound source corresponding to the matched sound source position as the target sound source; and determine the voice information of the target object from the multiple sound source audio information corresponding to the multiple sound sources according to the target sound source.

[0022] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to: implement any one of the speech recognition methods described.

[0023] According to a fourth aspect of an embodiment of the present disclosure, a non-transitory computer-readable storage medium is provided, which, when instructions in the storage medium are executed by a processor of a mobile terminal, enables the mobile terminal to execute any one of the speech recognition methods described above.

[0024] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any one of the speech recognition methods described above.

[0025] The technical solutions provided by the embodiments of the present disclosure may have the following beneficial effects:

[0026] The disclosed embodiment collects image information including the target object, and in response to the preset posture of the target object included in the collected image information, collects audio information including the voice information of the target object through a microphone array, and can achieve focusing of the target object for voice interaction with the smart device based on the preset posture. The target object can be accurately tracked through the microphone array, so that it can be accurately awakened in an environment where multiple devices coexist, and accurate voice information can be collected. The voice information of the target object is recognized based on the above-collected image information and audio information, so as to realize waking up the smart device based on the preset posture to trigger the collection of audio information, and can accurately track the target object to achieve accurate and smooth voice interaction.

[0027] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0029] Figure 1 This is a flow chart of a speech recognition method according to an exemplary embodiment of the present disclosure. Figure 1 .

[0030] Figure 2 This is a flow chart of a speech recognition method according to an exemplary embodiment of the present disclosure. Figure 2 .

[0031] Figure 3 This is a flow chart of a speech recognition method according to an exemplary embodiment of the present disclosure. Figure 3 .

[0032] Figure 4 This is a flow chart of a speech recognition method according to an exemplary embodiment of the present disclosure. Figure 4 .

[0033] Figure 5 This is a first scene diagram of a speech recognition method according to an exemplary embodiment of the present disclosure.

[0034] Figure 6 4 is a second scene diagram of a speech recognition method according to an exemplary embodiment of the present disclosure.

[0035] Figure 7 The block diagram of a speech recognition device according to an exemplary embodiment of the present disclosure is shown.

[0036] Figure 8 is a block diagram of an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0037] Some exemplary embodiments of the present disclosure will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. Various changes, modifications and equivalents of the methods, devices and / or systems described herein will become apparent after understanding the present disclosure. For example, the order of operations described herein is merely an example and is not limited to those orders set forth herein, but may be changed as becomes apparent after understanding the present disclosure, except for operations that must be performed in a specific order. In addition, descriptions of features known in the art may be omitted for clarity and brevity.

[0038] The following exemplary embodiments of the present disclosure do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0039] With the rapid development of artificial intelligence, especially the maturity of speech recognition and natural language processing technologies, convenient and fast voice interaction has been embedded in various smart home devices and has entered thousands of households. However, how to focus on the target device among so many devices is also a challenge.

[0040] Figure 1 This is a flow chart of a speech recognition method according to an exemplary embodiment of the present disclosure. Figure 1 The method of this embodiment can be applied to electronic devices that can collect image information and audio information, including smart speakers, smart phones, and smart tablet terminal devices, and can also include server terminals such as local servers and cloud servers. The server terminal can be deployed in a computer cluster consisting of one computer or multiple computers.

[0041] like Figure 1As shown, in some embodiments, the speech recognition method of the present disclosure example includes:

[0042] In step S110, image information including the target object is collected;

[0043] The speech recognition method of the embodiment of the present disclosure is applied to an electronic device, which has an image acquisition device and an audio acquisition device, and collects image information through the image acquisition device. The image information collected by the image acquisition device includes the target object. Taking a smart TV as an example, the smart TV has a camera, which serves as an image acquisition device and can collect image information containing the target object. The camera can include: a combination of one or more types of cameras such as ordinary image cameras, depth cameras, infrared cameras, wide-angle cameras, etc. In a smart interaction scenario, the smart TV is installed in a home scene, and the target object can be a user. In one method, the camera continuously collects image information, including image information collected when the user is not within the camera's image acquisition range and image information collected when the user is within the camera's image acquisition range. In another method, when the user enters the camera's image acquisition range, image information is collected, and the image information includes the user.

[0044] In step S120, in response to the image information including the preset posture of the target object, audio information is collected through the microphone array; the audio information includes the voice information of the target object;

[0045] In a home scene, there are multiple electronic devices, each of which can collect the posture of a target object. When it is necessary to interact with a target device among multiple electronic devices, it is necessary to locate and focus on the target device through a preset posture. For example, the preset posture is that the target object raises his right thumb facing the electronic device. The electronic device detects the collected image information. If it is detected in the image information that the target object raises his right thumb facing the electronic device, it is determined that the target device that the target object needs to interact with is the current electronic device, that is, the electronic device determines that the image information includes the preset posture of the target object, and collects audio information through the microphone array. In this electronic device, the microphone array collects audio information as an audio collection device, and the microphone array includes at least two microphones. In order to enhance the accuracy of speech recognition, the number of microphones in the microphone array can be appropriately increased. In the scene where the above electronic device is located, there may be multiple objects. The voice information emitted by each object as a sound source will be mixed. The audio information includes mixed voice information of multiple objects. When performing speech interaction recognition, it is necessary to pay attention to the voice information of the target object in the audio information.

[0046] In step S130 , speech recognition is performed on the speech information of the target object based on the image information and the audio information.

[0047] After capturing image information through an image acquisition device and audio information containing the target object's voice information through a microphone array, speech recognition is performed on the target object's voice information. In complex multimodal interaction scenarios in smart homes, audio information alone may not be able to accurately identify the target object's voice content. By combining image and audio information, problems such as multi-source aliasing and noise interference can be effectively addressed, thereby improving the accuracy and robustness of speech recognition. Image information can be used to achieve visual object localization. Visual object localization can be achieved by using computer vision algorithms to detect the target object in the image, extract the target object's bounding box or key points, and calculate its position in the image plane or three-dimensional space based on the target object's bounding box center point or key point distribution, thereby achieving visual location of the target object. Common computer vision algorithms include YOLO (You Only Look Once) and Faster R-CNN (Faster Region-based Convolutional Neural Networks). Audio information can be used to achieve auditory object localization. A microphone array can be used to estimate the direction of arrival of the recorded sound signal to determine the direction of each sound source. By combining the spatial layout of the microphone array and the direction of the sound source, the specific position of each sound source in three-dimensional space is calculated, thereby achieving auditory localization of the target object. Common DOA estimation algorithms include MUSIC (Multiple Signal Classification), ESPRIT (Estimation of Signal Parameters via Rotational Invariance Techniques), and Beamforming. A dual-modal localization approach based on auditory and visual localization can distinguish multiple sound sources in a mixed speech, thereby isolating the target object's voice information from the mixed speech. The speech recognition engine then performs speech recognition on the target object's voice information, enabling voice interaction. By combining image and audio information, the accuracy of speech recognition can be significantly improved, especially in complex environments. The fusion of multimodal information enhances the system's anti-interference capabilities, enabling it to maintain high recognition accuracy even in noisy environments. By combining visual and audio information, the system can accurately locate the target object and extract its voice information, avoiding misidentification of other sound sources. The system is suitable for a variety of application scenarios, such as smart conferencing, security monitoring, smart home, etc., and has wide adaptability.

[0048] The disclosed embodiment collects image information including the target object, and in response to the preset posture of the target object included in the collected image information, collects audio information including the voice information of the target object through a microphone array, and can achieve focusing of the target object for voice interaction with the smart device based on the preset posture. The target object can be accurately tracked through the microphone array, so that it can be accurately awakened in an environment where multiple devices coexist, and accurate voice information can be collected. The voice information of the target object is recognized based on the above-collected image information and audio information, so as to realize waking up the smart device based on the preset posture to trigger the collection of audio information, and can accurately track the target object to achieve accurate and smooth voice interaction.

[0049] In the embodiment of the present disclosure, speech recognition is performed on the speech information of the target object based on the image information and the audio information, including:

[0050] Determining, based on the audio information, a plurality of sound source audio information corresponding to the plurality of sound sources respectively;

[0051] Determine the sound source positions corresponding to the multiple sound sources;

[0052] determining an object position of a target object based on the image information;

[0053] Determining the voice information of the target object from the audio information of the multiple sound sources according to the sound source position and the object position;

[0054] Perform speech recognition on the target object's speech information.

[0055] The audio information obtained above through the microphone array is obtained by mixing the voice information emitted by multiple sound sources in the scene where the electronic device is located. By separating the audio sources of the audio information, multiple sound source audio information corresponding to the multiple sound sources can be obtained, wherein each sound source corresponds to sound source audio information.

[0056] A microphone array is an array consisting of two or more microphones. Since the physical dimensions of the electronic device are known, and the specific location of the microphone array on the electronic device is also known, the location coordinates of each microphone in the microphone array can be obtained by establishing a coordinate system on the electronic device. When there are multiple electronic devices in a smart home scenario, the microphone array of each electronic device can establish a coordinate system on the corresponding electronic device to obtain the corresponding location coordinates. The location coordinates of the microphones can provide a basic positional relationship for microphone array direction finding, which is conducive to the subsequent calculation of the steering vector. In a smart home scenario, multiple objects may emit sounds at the same time. In this case, the multiple microphones in the microphone array of each electronic device will collect the mixed speech of each object, that is, the audio information mentioned above. In order to achieve accurate tracking of the target object, it is necessary to perform sound source separation on the audio information containing the mixed speech, so that the sound source audio information corresponding to each sound source can be obtained. Based on the multiple sound source audio information corresponding to the multiple sound sources obtained by the above sound source separation, the sound source positions corresponding to the multiple sound sources can be determined.

[0057] Through the above steps, the sound source can be separated and the location of the sound source can be achieved based on the audio information; further, the image information is analyzed and calculated to determine the object position of the target object in the scene where the electronic device is located; when the image acquisition device acquires image information, the image acquisition device usually has a fixed acquisition range. For example, if a monocular camera is installed on the electronic device, the acquisition range of the monocular camera is a sector-shaped area corresponding to a maximum angle of 180°. Since the indoor environment in a home scene is a three-dimensional space formed by multiple walls, the actual acquisition range will change based on the shape of the space on the basis of the above sector-shaped area. If a binocular camera is installed on the electronic device, the corresponding acquisition range can be expanded. Based on the acquisition range of the camera when acquiring image information, a certain context can be provided for image positioning, and the object position of the target object can be determined through the image information.

[0058] Based on the sound source position and object position obtained above, the voice information of the target object can be separated and determined from the audio information of multiple sound sources, thereby realizing the separate extraction of the voice information corresponding to the target object in complex scenes.

[0059] A microphone array is used to collect audio signals, and the direction of each sound source is calculated using a direction-of-arrival estimation algorithm. Combined with the spatial layout of the microphone array, this direction information is converted into the specific sound source location in three-dimensional space. A camera is used to capture video frames containing the target object, and the target object's location is detected using a computer vision algorithm. The coordinate systems of the visual system and the audio system are aligned to the same reference frame to ensure that their position information can be directly compared, thereby determining the specific location of the target object. Based on the specific location of the target object, the target object's voice information is separated from the audio information of multiple sound sources and determined. By extracting the above target object's voice information, voice recognition can be performed to achieve voice interaction.

[0060] The disclosed embodiments can effectively combine image information and audio information, accurately identify the voice information of the target object from complex environmental noise, perform voice recognition, and convert it into meaningful text content, thereby achieving efficient human-computer voice interaction.

[0061] In the embodiment of the present disclosure, determining, based on the audio information, multiple sound source audio information corresponding to the multiple sound sources respectively includes:

[0062] Perform Fourier transform on the audio information to obtain audio frequency domain information;

[0063] Blind source separation is performed on the audio frequency domain information to determine multiple sound source audio information corresponding to the multiple sound sources.

[0064] The acoustic front-end for voice interaction in smart home scenarios can use beamforming or blind source separation algorithms for far-field sound pickup. Blind source separation algorithms are widely used in various smart devices due to their low requirements for microphone array consistency and number of microphones, as well as their significant improvement in signal-to-noise ratio. However, the order in which sound sources are output by the separation algorithm is random. Therefore, a speaker tracking algorithm is required to determine the order of valid sound sources to ensure that only valid sound sources are fed into the ASR (Automatic Speech Recognition) engine, thus ensuring smooth voice interaction. Speaker tracking algorithms based on blind source separation typically ensure the output order of sound sources in two ways: by controlling the update of the covariance matrix in blind source separation and by controlling the update of the covariance matrix of the corresponding sound sources based on the speaker's frequency domain energy and utterance time, thereby ensuring that the speaker's sound sources are output in a predetermined order. Post-blind source separation processing typically uses the similarity between speaker characteristics such as the speaker's pitch frequency and the corresponding characteristics of the separated sound sources to distinguish speakers.

[0065] Collect audio information through microphone array x m(t), where m represents the audio information collected by the mth microphone; audio information is usually represented in the time domain, including the sound intensity that changes over time; Fourier transform of audio information can obtain audio frequency domain information, which represents the energy distribution at different frequencies. Specifically, the audio information x m (t) After the window is transformed by the Short Time Fourier Transform (STFT) of N points, the audio frequency domain information X is obtained. m (t,f); where the frame is marked as t∈{1,...,T} and the frequency is marked as

[0066] In the embodiment, the audio frequency domain information is determined as follows:

[0067] X m (t,f)=STFT(x m (t))

[0068] Among them, X m (t,f) is the audio frequency domain information; x m (t) is the audio information; STFT() is the short-time Fourier transform.

[0069] The audio frequency domain information obtained above is subjected to blind source separation processing by using a blind source separation algorithm W(f) to determine a plurality of sound source audio information corresponding to the plurality of sound sources.

[0070] The blind source separation algorithm separates the audio frequency domain information to obtain the sound source audio information, which is a posteriori frequency domain estimate. Taking a microphone array consisting of two microphones as an example, the blind source separation process includes:

[0071] Y(t,f)=[Y1(t,f),Y2(t,f)] T =W(f)X(t,f)

[0072] X(t,f)=[X1(t,f),X2(t,f)] T

[0073] Among them, Y1(t,f) and Y2(t,f) represent the frequency domain estimates of sound source s1 and sound source s2; Y(t,f) is the sound source audio information; X1(t,f) is the mixed audio of sound source s1 and sound source s2 collected by microphone 1; X2(t,f) is the mixed audio of sound source s1 and sound source s2 collected by microphone 2; T is the transpose.

[0074] The audio frequency domain information obtained through Fourier transform can better display the frequency components of the signal, helping to identify and distinguish different sound sources. By selecting appropriate window functions and parameter settings, background noise can be effectively removed or suppressed in the frequency domain. It can accurately separate the audio information of each sound source from a mixed audio signal, improving user experience and system performance in various application scenarios such as speech recognition, meeting recording, and intelligent assistants. Blind source separation technology significantly improves the intelligibility and usability of audio information, especially in complex scenarios such as multi-person conversations and noisy environments.

[0075] Figure 2 This is a flow chart of a speech recognition method according to an exemplary embodiment of the present disclosure. Figure 2 .like Figure 2 As shown, in the embodiment, determining the sound source positions corresponding to the multiple sound sources respectively includes:

[0076] Step S210: determining a probability of speech presence based on the audio frequency domain information and the multiple sound source audio information corresponding to the multiple sound sources;

[0077] Step S220: Obtaining effective frequency points corresponding to the sound source in the audio frequency domain information;

[0078] Step S230: determining the components of the multiple sound sources in the audio frequency domain information according to the effective frequency points and speech existence probabilities corresponding to the sound sources in the audio frequency domain information;

[0079] Step S240: determining the sound source positions corresponding to the multiple sound sources respectively according to the components of the multiple sound sources in the audio frequency domain information.

[0080] Due to the characteristics of the blind source separation algorithm, the multiple sound source audio information corresponding to the multiple sound sources obtained above are randomly output, so it is impossible to know which of the sound sources s1 and s2 corresponds to the target object, so tracking and separation of the target object is required.

[0081] First, the probability of speech existence is determined using audio frequency domain information and multiple sound source audio information corresponding to multiple sound sources.

[0082] Taking the microphone array composed of two microphones as an example, when the sound sources include sound sources s1 and s2, sound source s1 represents one path and sound source s2 represents another path. The calculation based on one of the paths specifically includes:

[0083] Using audio frequency domain information X m (t, f) and the multiple sound source audio information Y(t, f) corresponding to the multiple sound sources respectively calculate the probability of speech existence M(t, f); wherein the value of the speech existence probability M(t, f) is between 0 and 1, 0 represents that there is no speaker's speech at this frequency point, and 1 represents that there is a speaker's speech signal at this frequency point.

[0084] In the embodiment, the probability of speech presence is determined in the following manner:

[0085] M(t,f)=[abs(Y1(t,f)) / abs(X1(t,f)),abs(Y1(t,f)) / abs(X2(t,f))] T

[0086] Where M(t, f) is the probability of speech presence; Y1(t, f) represents the frequency domain estimation of sound source s1; X1(t, f) is the mixed audio of sound sources s1 and s2 collected by microphone 1; X2(t, f) is the mixed audio of sound sources s1 and s2 collected by microphone 2; T is the transpose; abs is the absolute value function, which represents the amplitude of the signal.

[0087] In an embodiment, effective frequency points corresponding to the sound source in the audio frequency domain information are obtained. In audio signal processing tasks, effective frequency points refer to frequency points in the audio frequency domain information that are related to a specific sound source and have significant energy distribution. These frequency points can reflect the characteristic frequency range of the sound source and its energy concentration area, and are an important basis for distinguishing different sound sources. By analyzing the audio frequency domain information and extracting effective frequency points, the performance of tasks such as sound source separation, speech recognition, and noise suppression can be significantly improved.

[0088] Based on the effective frequency points corresponding to the sound sources in the audio frequency domain information obtained above and the speech existence probability obtained above, the components of the multiple sound sources in the audio frequency domain information are calculated.

[0089] In the embodiment, the component of the sound source in the audio frequency domain information is determined as follows:

[0090] Z1(t,f)=M(t,f)X(t,f)

[0091] Among them, Z1(t,f) is the component of the sound source s1 in the audio frequency domain information; M(t,f) is the probability of speech existence.

[0092] Based on the same calculation method, the component of the sound source s2 in the audio frequency domain information can be calculated.

[0093] Based on the components of the sound source s1 and the sound source s2 in the audio frequency domain information, DOA (Direction of Arrival) direction finding is performed to determine the sound source positions corresponding to the multiple sound sources.

[0094] The probability of speech existence is calculated using the multiple sound source audio information and audio frequency domain information corresponding to the multiple sound sources obtained by blind source separation. The phase of the effective frequency point of each channel is obtained according to the speech existence probability. Only the effective phase is used for DOA direction finding, and the positions of the two sound sources can be obtained, realizing the accurate positioning of the sound source positions. Finally, combined with the position of the speaker given by visual information, one of the two sound sources is selected and sent to the speech recognition engine to ensure the correctness of voice interaction.

[0095] Figure 3 This is a flow chart of a speech recognition method according to an exemplary embodiment of the present disclosure. Figure 3 .like Figure 3 As shown, in the embodiment of the present disclosure, determining the sound source positions corresponding to the multiple sound sources respectively according to the components of the multiple sound sources in the audio frequency domain information includes:

[0096] Step S310: obtaining coordinates of a microphone array and viewing angle information of collected image information; dividing the space corresponding to the image information into a plurality of partitions according to the coordinates of the microphone array and the viewing angle information of collected image information; each of the plurality of partitions has point coordinates;

[0097] Step S320: Determine the steering vector of each point coordinate;

[0098] Step S330: determining a cross-correlation matrix based on the components of the multiple sound sources in the audio observation signal;

[0099] Step S340: determining the noise subspace according to the cross-correlation matrix;

[0100] Step S350: determining a spatial energy spectrum based on the noise subspace and the steering vector of each point coordinate;

[0101] Step S360: Determine the point coordinates corresponding to the maximum value in the spatial energy spectrum as the sound source position corresponding to the sound source in the scene.

[0102] There are many ways to locate the sound source. This embodiment of the present disclosure provides a method for locating the sound source by combining image information and audio information. The specific steps include:

[0103] Get the coordinates (y) of the microphone array k ,k=1,2,...,M) and the viewing angle information of the collected image information, according to the coordinate y of the microphone array k , k=1,2,...,M and the viewing angle information of the collected image information, the space G corresponding to the image information is divided into multiple partitions, and the point coordinates of each partition are represented by x i ,i=1,2,...,G, calculate the steering vector for each coordinate position

[0104] In the embodiment, the steering vector of each point coordinate is determined as follows:

[0105] a i (f)=[a i1 (f),a i2 (f),...,a iM (f)] T

[0106]

[0107] Among them, a i (f) is the steering vector of each point coordinate; f=1,2,...,F is the frequency point; f s is the sampling rate; c is the speed of sound, which is generally 340m / s.

[0108] According to the components Z1(t,f) of multiple sound sources in the audio observation signal, the cross-correlation matrix is ​​calculated:

[0109]

[0110] Where R(t, f) is the cross-correlation matrix; (·) H is the conjugate transpose.

[0111] Based on the cross-correlation matrix obtained by the above calculation, the noise subspace is calculated; and then the spatial energy spectrum is determined according to the noise subspace and the steering vector of each point coordinate.

[0112] In the embodiment, the spatial energy spectrum is determined in the following manner:

[0113]

[0114] Among them, p i is the spatial energy spectrum; a i (f) is the steering vector; U N is the noise subspace.

[0115] The spatial energy spectrum is numerically counted, and the point coordinates corresponding to the maximum value in the spatial energy spectrum are determined as the corresponding sound source position of the sound source in the scene.

[0116] In the embodiment, the following method is used to determine the coordinates of the point corresponding to the maximum value in the spatial energy spectrum:

[0117] θ1=argmax(p i )

[0118] Where θ1 is the sound source position of the sound source s1.

[0119] Using the above calculation process, the sound source position θ2 of the sound source s2 is also calculated.

[0120] In the embodiment, determining the noise subspace according to the cross-correlation matrix includes:

[0121] Performing eigenvalue decomposition on the cross-correlation matrix to determine a first matrix; the first matrix includes: an eigenvalue matrix and a first eigenvector; the first eigenvector corresponds to the direction of the signal space and the noise space; the eigenvalue represents the energy in each direction;

[0122] According to the size of the eigenvalue, the first eigenvector is divided into a signal subspace and a noise subspace; the signal subspace includes the eigenvector corresponding to the maximum eigenvalue; the noise subspace is the remaining eigenvector in the first eigenvector after removing the signal subspace.

[0123] The cross-correlation matrix R(t, f) is decomposed by eigenvalue to obtain the first matrix U consisting of the eigenvalue matrix Σ and the eigenvector. The eigenvector corresponds to the direction of the signal space and the noise space, and the eigenvalue reflects the energy in each direction. Therefore, according to the size of the eigenvalue, the eigenvector is divided into the signal subspace U s and noise subspace U N The signal subspace is composed of the subspace corresponding to the largest eigenvalue Σ s The noise subspace is composed of the eigenvectors of N constitute.

[0124] In the embodiment, the noise subspace is calculated as follows:

[0125]

[0126] Among them, U N is the noise subspace; Σ is the eigenvalue matrix; U s is the signal subspace; R(f) is the cross-correlation matrix. Since the above calculation formulas are specific to each frame, t is omitted.

[0127] In the embodiment of the present disclosure, determining the object position of the target object based on image information includes:

[0128] If the preset posture of the target object is included in the image information, the object position of the target object in the space corresponding to the image information is determined according to the image information.

[0129] Taking the above-mentioned preset posture of the target object raising the right thumb facing the electronic device as an example, when it is detected that the image information includes the target object raising the right thumb facing the electronic device, the object position of the target object in the space corresponding to the image information is calculated based on the image where the target user currently raises the right thumb to the electronic device; if the camera uses a monocular camera, the two-dimensional coordinates of the target object on the image plane can be determined by the center point of the bounding box or the average position of the key points of the target object, and then the object position is obtained. If a binocular camera or a multi-view sensor is used, combined with triangulation or depth estimation technology (such as the deep learning model DepthNet), the position of the target object in three-dimensional space is calculated. If the image acquisition device uses a depth camera, the object position of the target object in space can be directly calculated by the depth camera.

[0130] The preset posture information is used to provide additional contextual information for the spatial position estimation of the target object, thereby improving the accuracy and reliability of positioning.

[0131] In an embodiment of the present disclosure, determining the voice information of a target object from multiple sound source audio information based on the sound source position and the object position includes:

[0132] If the sound source position matches the object position, the sound source corresponding to the matched sound source position is determined as the target sound source;

[0133] According to the target sound source, the speech information of the target object is determined from the plurality of sound source audio information respectively corresponding to the plurality of sound sources.

[0134] The coordinate systems of the sound source position and the object position are unified, converted to the unified coordinate system for position matching, and the distance between each sound source position and the target object position is calculated. If the distance between a sound source position and the target object position is less than a set threshold, it means that the sound source position and the object position match, and the sound source corresponding to the matched sound source position is determined as the target sound source; multiple sound sources correspond to multiple sound source audio information, and the corresponding target sound source is found from the multiple sound sources to determine the voice information of the target object of the target sound source in the multiple sound source audio information.

[0135] Combining visual and audio information can significantly improve the accuracy of target sound source localization. In multi-person conversation scenarios, by precisely matching the speaker's object location and sound source location, each person's voice information can be more accurately separated.

[0136] In the embodiment, the object position θ obtained by the above calculation v (t), select the target object's voice information s(t) from the two sound sources s1 and s2 and input it into the recognition engine for voice recognition, thus completing voice interaction.

[0137] In the embodiment, the voice information of the target object is calculated as follows:

[0138]

[0139] dist(θ1,θ2)=|θ1-θ2|

[0140] s(t)=ISTFT(S(t,f))

[0141] Among them, s(t) is the speech information of the target object; θ1 is the sound source position of sound source s1; θ2 is the sound source position of sound source s2; ISTFT() is the inverse short-time Fourier transform.

[0142] In the embodiment of the present disclosure, the voice recognition method further includes: performing voice recognition on the voice information of the target object to determine a control instruction; and controlling the target device according to the control instruction.

[0143] By performing speech recognition on the target user's voice information and analyzing the recognized text using a Natural Language Understanding (NLU) model, the system extracts the user's control intent and key information from the text (such as device name, action type, parameter value, etc.) to form structured control instructions. For example, the voice message "Turn up the TV volume" can be parsed into a control instruction containing [Device Type: TV, Action: Increase Volume, Location: Living Room]. Based on this control instruction, the TV volume can be increased by a certain amount, thereby achieving control of the target device (such as a TV).

[0144] Combining speech recognition, semantic parsing, and device control technologies enables an efficient and intelligent human-computer interaction system. This approach not only improves the accuracy and robustness of voice control but also offers broad applicability and flexible application scenarios. Structured control instructions clarify device control logic, improving system responsiveness and stability. Direct interaction with devices through natural language eliminates the need to memorize complex operating instructions, lowering the barrier to entry.

[0145] In the embodiment of the present disclosure, the speech recognition method further includes: extracting voiceprint features of the target object based on the speech information of the target object; and performing speech interaction with the target object based on the voiceprint features.

[0146] In intelligent voice interaction systems, speaker recognition is a key technology used to verify or identify the speaker's identity. By extracting the target subject's voiceprint characteristics, personalized services and security authentication can be achieved. For example, in scenarios such as smart homes, bank customer service, and medical assistance, voiceprint characteristics can help confirm the user's identity and provide a customized interactive experience based on their identity.

[0147] Unique acoustic features are extracted from the target subject's voice information for identification. The extracted voiceprint features are compared with known voiceprints in the database to confirm the target subject's identity. Based on the target subject's identity information, personalized voice services or specific operations are provided. Voiceprint features are typically based on the spectral characteristics of the speech signal. Common features include Mel-Frequency Cepstral Coefficients (MFCC), Linear Predictive Coding (LPC), i-vectors, or x-vectors. The target subject's voiceprint features are compared with known voiceprint templates in the database to calculate a similarity score. If the score exceeds the preset threshold, the target subject's identity is confirmed; otherwise, the authentication fails.

[0148] With the help of advanced voiceprint feature extraction and matching algorithms, the identity of the target object can be accurately identified in complex environments. With the help of advanced voiceprint feature extraction and matching algorithms, the identity of the target object can be accurately identified in complex environments.

[0149] Figure 4 This is a flow chart of a speech recognition method according to an exemplary embodiment of the present disclosure. Figure 4 ,like Figure 4 As shown, the embodiment of the present disclosure provides a process of a speech recognition method, including:

[0150] The method includes collecting image information, wherein the image information includes a target object; in response to a preset posture of the target object included in the image information, collecting audio information through a microphone array; the audio information includes voice information of the target object; performing Fourier transform on the audio information to obtain audio frequency domain information X1(t, f) and X2(t, f); performing blind source separation processing on the audio frequency domain information to determine multiple sound source audio information Y1(t, f) and Y2(t, f) corresponding to multiple sound sources; determining a speech existence probability M(t, f) based on the audio frequency domain information and the multiple sound source audio information corresponding to the multiple sound sources; determining components of the multiple sound sources in the audio frequency domain information based on effective frequency points of the corresponding sound sources in the audio frequency domain information and the speech existence probability; determining sound source positions θ1(t) and θ2(t) corresponding to the multiple sound sources based on the components of the multiple sound sources in the audio frequency domain information; and determining an object position θ of the target object based on the image information. v (t); determining the voice information of the target object from the audio information of multiple sound sources according to the sound source position and the object position; and performing speech recognition on the voice information of the target object.

[0151] Figure 5 is a first scene diagram of a speech recognition method according to an exemplary embodiment of the present disclosure, such as Figure 3As shown, when the speech recognition method provided by the embodiment of the present disclosure is applied to a smart home scenario, the speech recognition method is deployed in a smart TV. The smart TV has a camera and a microphone array. The microphone array includes microphone 1 and microphone 2. The camera has a collection angle of view. Within the collection angle of view, image information including the target object is collected, wherein the collection angle of view includes object 1 and object 2, taking object 1 as the target object as an example; when the image information includes a preset posture of the target object, audio information is collected by microphone 1 and microphone 2 of the microphone array, and the audio information is a mixture of voice information emitted by object 1 and voice information emitted by object 2. Perform Fourier transform on audio information to obtain audio frequency domain information; perform blind source separation on audio frequency domain information to determine multiple sound source audio information corresponding to multiple sound sources; determine the probability of speech existence based on the audio frequency domain information and the multiple sound source audio information corresponding to the multiple sound sources; obtain the effective frequency point of the corresponding sound source in the audio frequency domain information; determine the components of the multiple sound sources in the audio frequency domain information based on the effective frequency point of the corresponding sound source in the audio frequency domain information and the probability of speech existence; obtain the coordinates of the microphone array and the viewing angle information of the collected image information; divide the space corresponding to the image information into multiple partitions based on the coordinates of the microphone array and the viewing angle information of the collected image information; each partition in the multiple partitions has point coordinates; determine the steering vector of each point coordinate; determine the cross-correlation matrix based on the components of the multiple sound sources in the audio observation signal; perform eigenvalue decomposition on the cross-correlation matrix to determine the first matrix; the first matrix includes: an eigenvalue matrix and a first eigenvector; the first eigenvector The quantity corresponds to the direction of the signal space and the noise space; the eigenvalue represents the energy size in each direction; according to the size of the eigenvalue, the first eigenvector is divided into a signal subspace and a noise subspace; the signal subspace includes the eigenvector corresponding to the maximum eigenvalue; the noise subspace is the remaining eigenvector in the first eigenvector after removing the signal subspace; according to the noise subspace and the steering vector of each point coordinate, the spatial energy spectrum is determined; the point coordinates corresponding to the maximum value in the spatial energy spectrum are determined as the sound source position corresponding to the sound source in the scene. If the image information includes a preset posture of the target object, the object position of the target object in the space corresponding to the image information is determined according to the image information; according to the sound source position and the object position, the voice information of the target object, that is, the voice information of object 1, is determined from the audio information of multiple sound sources; voice recognition is performed on the voice information of object 1, and a control instruction is determined based on the voice recognition of the voice information of object 1; the target device is controlled according to the control instruction.

[0152] Figure 6 is a second scene diagram of a speech recognition method according to an exemplary embodiment of the present disclosure, such as Figure 6As shown, in a smart home scenario, when there are multiple electronic devices, such as a smart TV, a smart refrigerator, and a smart speaker, all of which include cameras and microphone arrays, when the target object only needs to interact with the smart TV by voice, it faces the smart TV, and the preset gesture is that the target object raises its right thumb facing the electronic device. However, since the smart refrigerator and smart speaker cannot capture the side of the target object facing the smart TV, they cannot capture the corresponding preset gesture. Then, when the target object raises its right thumb facing the smart TV, the image information captured by the smart TV includes the preset gesture of the target object, and the smart TV will be awakened. The smart TV's microphone array will capture audio information, and then voice recognition of the target object's voice information will be performed based on the image information and audio information. However, since the smart refrigerator and smart speaker do not capture the preset gesture, they will not be awakened. Using the above preset gesture to wake up the activated device and then perform voice interaction can not only activate the target device for interaction and control, but also avoid the limitation of single gesture recognition on the complexity of instructions. At the same time, the disadvantage of the object position of the target object can be realized based on image information. Combining the sound source position and the object position, the voice information of the target object can be determined from multiple sound source audio information, and then voice recognition can be performed on the voice information of the target object.

[0153] The disclosed embodiment can focus on the target interactive device among many smart devices, and the instructions will not be executed by other unrelated devices; secondly, there is no need to use complex gestures for interaction, and voice interaction is simple, accurate and fast; thirdly, by combining gesture and visual information, the target object can be tracked very accurately, and even if the device can use a minimum of 2 microphones, it can still have an accurate and smooth voice interaction experience, greatly reducing the hardware cost of the device.

[0154] It should be noted that the acquisition, storage, use, and processing of data in the technical solution disclosed herein are in compliance with the relevant provisions of national laws and regulations. Various types of data such as personal identity data, operation data, behavioral data, etc. related to individuals, customers, and groups obtained in the embodiments of the present disclosure have been authorized.

[0155] The following are embodiments of the apparatus disclosed herein, which can be used to implement the method embodiments disclosed herein. For details not disclosed in the apparatus embodiments disclosed herein, please refer to the method embodiments disclosed herein.

[0156] Figure 7 This is a block diagram of a speech recognition device according to an exemplary embodiment of the present disclosure. The device of this embodiment can be applied to electronic devices, including smart speakers, smartphones, and smart tablets. It can also include a server, such as a local server or a cloud server. The server can be deployed on a single computer or in a computer cluster consisting of multiple computers.

[0157] like Figure 7 As shown, the speech recognition device 700 may include:

[0158] An image information acquisition module 710 is used to acquire image information including a target object;

[0159] The audio information collection module 720 is configured to collect audio information through a microphone array in response to the image information including a preset posture of the target object; the audio information includes voice information of the target object;

[0160] The speech recognition module 730 is configured to perform speech recognition on the speech information of the target object based on the image information and audio information.

[0161] In some exemplary embodiments of the present disclosure, the speech recognition module 730 includes:

[0162] A sound source audio information determining unit, configured to determine a plurality of sound source audio information corresponding to the plurality of sound sources respectively according to the audio information;

[0163] a sound source position determination unit, configured to determine the sound source positions corresponding to the plurality of sound sources;

[0164] an object position determining unit, configured to determine an object position of a target object based on image information;

[0165] a target object voice information determination unit, configured to determine the voice information of the target object from the plurality of sound source audio information according to the sound source position and the object position;

[0166] The speech recognition unit is used to perform speech recognition on the speech information of the target object.

[0167] In some exemplary embodiments of the present disclosure, the sound source audio information determination unit is further used to: perform Fourier transform on the audio information to obtain audio frequency domain information; perform blind source separation processing on the audio frequency domain information to determine multiple sound source audio information corresponding to multiple sound sources respectively.

[0168] In some exemplary embodiments of the present disclosure, the sound source position determination unit is further used to: determine the probability of speech existence based on the audio frequency domain information and the multiple sound source audio information corresponding to the multiple sound sources respectively; obtain the effective frequency point of the corresponding sound source in the audio frequency domain information; determine the components of the multiple sound sources in the audio frequency domain information based on the effective frequency point of the corresponding sound source in the audio frequency domain information and the probability of speech existence; determine the sound source positions corresponding to the multiple sound sources respectively based on the components of the multiple sound sources in the audio frequency domain information.

[0169] In some exemplary embodiments of the present disclosure, the sound source position determination unit is further used to: obtain the coordinates of the microphone array and the viewing angle information of the collected image information; divide the space corresponding to the image information into multiple partitions according to the coordinates of the microphone array and the viewing angle information of the collected image information; each of the multiple partitions has point coordinates; determine the steering vector of each point coordinate; determine the cross-correlation matrix according to the components of multiple sound sources in the audio observation signal; determine the noise subspace according to the cross-correlation matrix; determine the spatial energy spectrum according to the noise subspace and the steering vector of each point coordinate; determine the point coordinate corresponding to the maximum value in the spatial energy spectrum as the sound source position corresponding to the sound source in the scene.

[0170] In some exemplary embodiments of the present disclosure, the sound source position determination unit is further used to: perform eigenvalue decomposition on the mutual correlation matrix to determine a first matrix; the first matrix includes: an eigenvalue matrix and a first eigenvector; the first eigenvector corresponds to the direction of the signal space and the noise space; the eigenvalue represents the energy size in each direction; according to the size of the eigenvalue, the first eigenvector is divided into a signal subspace and a noise subspace; the signal subspace includes the eigenvector corresponding to the maximum eigenvalue; the noise subspace is the remaining eigenvector in the first eigenvector after removing the signal subspace.

[0171] In some exemplary embodiments of the present disclosure, the object position determining unit is further configured to: if the image information includes a preset posture of the target object, determine the object position of the target object in the space corresponding to the image information based on the image information.

[0172] In some exemplary embodiments of the present disclosure, the target object voice information determination unit is also used to: if the sound source position and the object position match, determine the sound source corresponding to the matched sound source position as the target sound source; and according to the target sound source, determine the voice information of the target object from multiple sound source audio information corresponding to the multiple sound sources.

[0173] In some exemplary embodiments of the present disclosure, the speech recognition apparatus further includes: a target device control module, configured to: determine a control instruction based on speech recognition of speech information of a target object; and control the target device according to the control instruction.

[0174] In some exemplary embodiments of the present disclosure, the speech recognition device further includes: a speech interaction module, configured to: extract voiceprint features of the target object based on the speech information of the target object; and perform speech interaction with the target object based on the voiceprint features.

[0175] Figure 81 is a block diagram of an electronic device according to an exemplary embodiment of the present disclosure. For example, the device 1100 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0176] Reference Figure 8 , the device 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output (I / O) interface 812 , a sensor component 814 , and a communication component 816 .

[0177] The processing component 802 generally controls the overall operation of the device 800, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 802 may include one or more modules to facilitate interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate interaction between the multimedia component 808 and the processing component 802.

[0178] The memory 804 is configured to store various types of data to support operations on the device 800. Examples of such data include instructions for any application or method operating on the device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0179] The power supply component 806 provides power to the various components of the device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device 800.

[0180] The multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor can not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have a focal length and optical zoom capability.

[0181] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), which is configured to receive external audio signals when the device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.

[0182] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include but are not limited to: a home button, volume buttons, a start button, and a lock button.

[0183] The sensor assembly 814 includes one or more sensors for providing various aspects of the status assessment of the device 800. For example, the sensor assembly 814 can detect the open / closed state of the device 800, the relative positioning of components, such as the display and keypad of the device 800. The sensor assembly 814 can also detect changes in the position of the device 800 or a component of the device 800, the presence or absence of user contact with the device 800, the orientation or acceleration / deceleration of the device 800, and temperature changes of the device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0184] The communication component 816 is configured to facilitate wired or wireless communication between the device 800 and other devices. The device 800 can access a wireless network based on a communication standard, such as WiFi, 3G, 4G, 5G, other communication standards, or a combination thereof. In some embodiments of the present disclosure, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In some embodiments of the present disclosure, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0185] In some embodiments of the present disclosure, the apparatus 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described methods.

[0186] In some embodiments of the present disclosure, a non-transitory computer-readable storage medium including instructions is further provided, such as a memory 804 including instructions, and the instructions can be executed by the processor 820 of the apparatus 800 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0187] In some embodiments of the present disclosure, a non-transitory computer-readable storage medium, when instructions in the storage medium are executed by a processor of a mobile terminal, enables the mobile terminal to perform a speech recognition method, the method comprising:

[0188] Collecting image information including the target object;

[0189] In response to the image information including a preset posture of the target object, collecting audio information through a microphone array; the audio information includes voice information of the target object;

[0190] Perform speech recognition on the speech information of the target object based on the image information and audio information.

[0191] In some embodiments of the present disclosure, a computer program product is further provided, including a computer program / instruction. When the computer program / instruction is executed by a processor, a speech recognition method is implemented. The method includes:

[0192] Collecting image information including the target object;

[0193] In response to the image information including a preset posture of the target object, collecting audio information through a microphone array; the audio information includes voice information of the target object;

[0194] Based on the image information and audio information, speech recognition is performed on the speech information of the target object. After considering the specification and practicing the invention disclosed herein, those skilled in the art will easily come up with other embodiments of the present disclosure. This application is intended to cover any variations, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common knowledge or customary technical means in the art not disclosed in the present disclosure. The description and examples are to be regarded as exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0195] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A speech recognition method, characterized in that: include: Collecting image information including the target object; In response to the image information including a preset posture of the target object, collecting audio information through a microphone array; The audio information includes the voice information of the target object; Perform speech recognition on the speech information of the target object based on the image information and audio information.

2. The method according to claim 1, characterized in that The performing speech recognition on the speech information of the target object based on the image information and the audio information includes: Determining, based on the audio information, a plurality of sound source audio information corresponding to the plurality of sound sources respectively; Determining sound source positions corresponding to the multiple sound sources respectively; determining an object position of the target object based on the image information; Determining the voice information of the target object from the plurality of sound source audio information according to the sound source position and the object position; Perform speech recognition on the speech information of the target object.

3. The method according to claim 2, characterized in that The step of determining, based on the audio information, a plurality of sound source audio information corresponding to the plurality of sound sources, respectively, includes: Performing Fourier transform on the audio information to obtain audio frequency domain information; Blind source separation is performed on the audio frequency domain information to determine a plurality of sound source audio information corresponding to the plurality of sound sources respectively.

4. The method according to claim 2, characterized in that The determining of the sound source positions corresponding to the plurality of sound sources respectively includes: Determining a probability of speech existence based on the audio frequency domain information and the multiple sound source audio information corresponding to the multiple sound sources respectively; Obtaining the effective frequency point of the corresponding sound source in the audio frequency domain information; Determining components of multiple sound sources in the audio frequency domain information according to effective frequency points corresponding to the sound sources in the audio frequency domain information and the probability of existence of the speech; The sound source positions corresponding to the multiple sound sources are determined according to the components of the multiple sound sources in the audio frequency domain information.

5. The method according to claim 4, characterized in that Determining, according to the components of the multiple sound sources in the audio frequency domain information, sound source positions corresponding to the multiple sound sources respectively, includes: Obtaining coordinates of a microphone array and information of a viewing angle for collecting the image information; dividing a space corresponding to the image information into a plurality of partitions according to the coordinates of the microphone array and the information of the viewing angle for collecting the image information; each of the plurality of partitions having point coordinates; Determining a steering vector for each of the point coordinates; determining a cross-correlation matrix according to components of the multiple sound sources in the audio observation signal; determining a noise subspace according to the cross-correlation matrix; determining a spatial energy spectrum according to the noise subspace and the steering vector of each of the point coordinates; The point coordinates corresponding to the maximum value in the spatial energy spectrum are determined as the sound source position corresponding to the sound source in the scene.

6. The method according to claim 5, characterized in that Determining a noise subspace according to the cross-correlation matrix includes: Performing eigenvalue decomposition on the cross-correlation matrix to determine a first matrix; the first matrix includes: an eigenvalue matrix and a first eigenvector; the first eigenvector corresponds to a direction in a signal space and a noise space; the eigenvalue represents the energy in each direction; According to the size of the eigenvalue, the first eigenvector is divided into a signal subspace and a noise subspace; the signal subspace includes the eigenvector corresponding to the maximum eigenvalue; the noise subspace is the remaining eigenvector in the first eigenvector after removing the signal subspace.

7. The method according to claim 2, characterized in that Determining the object position of the target object according to the image information includes: If the preset posture of the target object is included in the image information, the object position of the target object in the space corresponding to the image information is determined according to the image information.

8. The method according to claim 2, characterized in that Determining the voice information of the target object from the plurality of sound source audio information according to the sound source position and the object position includes: If the sound source position matches the object position, determining the sound source corresponding to the matched sound source position as the target sound source; According to the target sound source, the voice information of the target object is determined from the multiple sound source audio information respectively corresponding to the multiple sound sources.

9. The method according to claim 1, characterized in that The method further comprises: Determining a control instruction based on voice recognition of the target object's voice information; The target device is controlled according to the control instruction.

10. The method according to claim 1, characterized in that The method further comprises: Extracting the voiceprint features of the target object according to the voice information of the target object; Perform voice interaction with the target object based on the voiceprint feature.

11. A speech recognition device, characterized in that: include: An image information acquisition module, used to acquire image information including the target object; an audio information collection module, configured to collect audio information through a microphone array in response to the image information including a preset posture of the target object; the audio information includes voice information of the target object; The speech recognition module is used to perform speech recognition on the speech information of the target object based on the image information and audio information.

12. The device according to claim 11, characterized in that The speech recognition module comprises: a sound source audio information determining unit, configured to determine, based on the audio information, a plurality of sound source audio information corresponding to the plurality of sound sources; a sound source position determining unit, configured to determine the sound source positions corresponding to the plurality of sound sources; an object position determining unit, configured to determine an object position of the target object based on the image information; a target object voice information determining unit, configured to determine the voice information of the target object from the plurality of sound source audio information according to the sound source position and the object position; The speech recognition unit is used to perform speech recognition on the speech information of the target object.

13. The device according to claim 12, characterized in that The sound source audio information determination unit is further configured to: perform Fourier transform on the audio information to obtain audio frequency domain information; and perform blind source separation processing on the audio frequency domain information to determine multiple sound source audio information corresponding to the multiple sound sources.

14. The device according to claim 12, characterized in that The sound source location determination unit is further configured to: determine a probability of speech presence based on the audio frequency domain information and multiple sound source audio information corresponding to the multiple sound sources; obtain an effective frequency point corresponding to the sound source in the audio frequency domain information; and determine a component of the multiple sound sources in the audio frequency domain information based on the effective frequency point corresponding to the sound source in the audio frequency domain information and the speech presence probability; The sound source positions corresponding to the multiple sound sources are determined according to the components of the multiple sound sources in the audio frequency domain information.

15. The device according to claim 14, characterized in that The sound source position determination unit is further configured to: obtain the coordinates of the microphone array and the viewing angle information of the image information; divide the space corresponding to the image information into a plurality of partitions according to the coordinates of the microphone array and the viewing angle information of the image information; each of the plurality of partitions has a point coordinate; determine a steering vector for each of the point coordinates; determine a cross-correlation matrix according to the components of the plurality of sound sources in the audio observation signal; determine a noise subspace according to the cross-correlation matrix; and determine a spatial energy spectrum according to the noise subspace and the steering vector for each of the point coordinates; The point coordinates corresponding to the maximum value in the spatial energy spectrum are determined as the sound source position corresponding to the sound source in the scene.

16. The device according to claim 15, characterized in that The sound source position determination unit is further used to: perform eigenvalue decomposition on the mutual correlation matrix to determine a first matrix; the first matrix includes: an eigenvalue matrix and a first eigenvector; the first eigenvector corresponds to the direction of the signal space and the noise space; the eigenvalue represents the energy size in each direction; according to the size of the eigenvalue, the first eigenvector is divided into a signal subspace and a noise subspace; the signal subspace includes the eigenvector corresponding to the maximum eigenvalue; the noise subspace is the remaining eigenvector in the first eigenvector after removing the signal subspace.

17. The device according to claim 12, characterized in that The target object voice information determination unit is further used to: if the sound source position and the object position match, determine the sound source corresponding to the matched sound source position as the target sound source; and determine the voice information of the target object from the multiple sound source audio information corresponding to the multiple sound sources according to the target sound source.

18. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to implement the speech recognition method according to any one of claims 1 to 10.

19. A non-transitory computer-readable storage medium, which, when instructions in the storage medium are executed by a processor of a mobile terminal, enables the mobile terminal to execute the speech recognition method according to any one of claims 1 to 10.

20. A computer program product, characterized in that The invention comprises a computer program, which implements the speech recognition method according to any one of claims 1 to 10 when the computer program is executed by a processor.

Citation Information

Patent Citations

  • Multi-sound-source locating method based on spherical microphone array

    CN102866385A

  • Household appliance control method, device and system and intelligent air conditioner

    CN106440192A

  • Wake-up method for intelligent product, intelligent product, and computer-readable storage medium

    CN107679506A

  • Method for improving voice wake-up rate and correcting DOA

    CN108122563A

  • Sound source positioning method and device based on image recognition and voice recognition

    CN109506568A