Speech signal processing method and device, readable storage medium and electronic equipment
By combining speech signals and image sequences, and utilizing single-mode and multi-mode processing methods, the speech segment signal output is selected based on image quality, thus solving the problem of decreased speech recognition rate in complex environments and improving recognition accuracy.
Patent Information
- Application Number
- CN202310035251.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-10
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2043-01-10
AI Technical Summary
In complex environments such as high noise or when the face is obscured, the recognition rate of traditional speech recognition technology decreases, and the performance of multimodal speech recognition methods drops significantly when visual features are invalid, failing to effectively remove the influence of invalid visual features.
By acquiring speech signals and image sequences, speech segment signals are extracted using single-mode and multi-mode speech processing methods, and the target speech segment signal is selected for output based on image quality information, thus achieving real-time switching.
It improves the accuracy of speech recognition, especially when the image quality is not up to standard, and reduces the impact of invalid visual features on recognition accuracy.
Smart Images

Figure CN116013262B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, and particularly relates to a speech signal processing method and device, a computer readable storage medium and an electronic device. BACKGROUND
[0002] Traditional speech recognition technology only processes speech signals to obtain a recognition result. Such a speech recognition method has good recognition effect in a clear speech environment. However, in some complex environments such as high noise, the recognition rate of the traditional speech recognition technology will rapidly decrease. In order to improve the speech recognition rate, there is currently a multi-modal speech recognition method that assists speech recognition by means of a lip movement video, which improves the speech recognition rate in a high noise environment to a certain extent.
[0003] However, in a real-time speech interaction system, in the case that a user's face is blocked or a face image is unclear, the visual features obtained based on image recognition become invalid interference input, and the performance of the multi-modal speech recognition method will significantly decrease. Therefore, in the case that the visual features are invalid, how to remove the invalid features and only recognize valid speech signals is a problem to be solved. SUMMARY
[0004] In order to solve the above technical problems, the present disclosure is proposed. Embodiments of the present disclosure provide a speech signal processing method and device, a computer readable storage medium and an electronic device.
[0005] An embodiment of the present disclosure provides a speech signal processing method, which comprises: acquiring a speech signal and an image sequence in a target space; based on the speech signal, extracting a first speech segment signal from the speech signal by a first speech processing manner; based on the speech signal and the image sequence, extracting a second speech segment signal from the speech signal by a second speech processing manner; determining whether a current speech signal processing state meets a speech signal output condition; in response to the speech signal processing state meeting the speech signal output condition, determining image quality information of the image sequence; based on the image quality information of the image sequence, determining a target speech segment signal from the first speech segment signal and the second speech segment signal, and outputting the target speech segment signal.
[0006] According to another aspect of the embodiments of the present disclosure, a voice signal processing apparatus is provided, which comprises: an acquisition module configured to acquire a voice signal and an image sequence in a target space; a first extraction module configured to extract, based on the voice signal, a first voice segment signal from the voice signal by using a first voice processing manner; a second extraction module configured to extract, based on the voice signal and the image sequence, a second voice segment signal from the voice signal by using a second voice processing manner; a first determination module configured to determine whether a current voice signal processing state meets a voice signal output condition; a second determination module configured to determine image quality information of the image sequence in response to the voice signal processing state meeting the voice signal output condition; and an output module configured to determine a target voice segment signal from the first voice segment signal and the second voice segment signal based on the image quality information of the image sequence, and output the target voice segment signal.
[0007] According to another aspect of the embodiments of the present disclosure, a computer readable storage medium is provided, which stores a computer program for executing the voice signal processing method.
[0008] According to another aspect of the embodiments of the present disclosure, an electronic device is provided, which comprises: a processor; a memory configured to store processor-executable instructions; and the processor configured to read the executable instructions from the memory and execute the instructions to implement the voice signal processing method.
[0009] Based on the voice signal processing method, apparatus, computer readable storage medium and electronic device provided by the above embodiments of the present disclosure, the first voice segment signal is extracted from the voice signal by using the first voice processing manner based on the voice signal, the second voice segment signal is extracted from the voice signal by using the second voice processing manner based on the voice signal and the image sequence, then the image quality information of the image sequence is determined in response to the current voice signal processing state meeting the voice signal output condition, and finally the target voice segment signal is determined from the first voice segment signal and the second voice segment signal based on the image quality information of the image sequence, and the target voice segment signal is output. The embodiments of the present disclosure realize automatic selection of the first voice segment signal extracted by the single-mode voice processing manner or the second voice segment signal extracted by the multi-mode voice processing manner as the target voice segment signal required for voice recognition according to the image quality of the photographed image sequence, so that the source of the output voice segment signal is selected according to the image quality when the output target voice segment signal is recognized, and thus the accuracy of voice recognition is improved.
[0010] The technical solutions of the present disclosure will be further described in detail below with reference to the drawings and embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0011] The above and other objects, features and advantages of the present disclosure will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings in which:
[0012] Figure 1 is a system diagram to which the present disclosure is applicable.
[0013] Figure 2 is a flowchart of a voice signal processing method according to an exemplary embodiment of the present disclosure.
[0014] Figure 3 is a flowchart of a voice signal processing method according to another exemplary embodiment of the present disclosure.
[0015] Figure 4 is a flowchart of a voice signal processing method according to another exemplary embodiment of the present disclosure.
[0016] Figure 5 is a flowchart of a voice signal processing method according to another exemplary embodiment of the present disclosure.
[0017] Figure 6 is a flowchart of a voice signal processing method according to another exemplary embodiment of the present disclosure.
[0018] Figure 7 is a structural diagram of a voice signal processing apparatus according to an exemplary embodiment of the present disclosure.
[0019] Figure 8 is a structural diagram of a voice signal processing apparatus according to another exemplary embodiment of the present disclosure.
[0020] Figure 9 is a structural diagram of an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0021] Hereinafter, example embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. It should be understood that the exemplary embodiments described herein are merely a part of the present disclosure, and the present disclosure is not limited by the exemplary embodiments described herein.
[0022] It should be noted that the relative arrangement of the components and steps illustrated in the embodiments set forth herein is not limiting of the scope of the present disclosure. The numerical expressions and values set forth in the detailed description are also not limiting of the scope of the present disclosure.
[0023] Those skilled in the art can understand that the terms "first", "second" and the like in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, and do not represent any specific technical meaning, nor represent the inherent logical order between them.
[0024] It should also be understood that in the embodiments of the present disclosure, "multiple" can mean two or more, and "at least one" can mean one, two or more.
[0025] It should also be understood that for any component, data or structure mentioned in the embodiments of the present disclosure, it can be understood as one or more in general, without explicit limitation or in the context of the opposite indication.
[0026] In addition, the term "and / or" in the present disclosure is only a description of the association relationship between the associated objects, which means that there can be three relationships, for example, A and / or B can represent: A exists alone, A and B exist together, and B exists alone. In addition, the character " / " in the present disclosure generally represents that the front and rear associated objects are in an "or" relationship.
[0027] It should also be understood that the description of the embodiments of the present disclosure emphasizes the differences between the embodiments, and the same or similar parts can be referred to each other, and for the sake of brevity, will not be repeated.
[0028] At the same time, it should be understood that in order to facilitate the description, the size of each part shown in the drawings is not drawn according to the actual proportional relationship.
[0029] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application or uses.
[0030] The techniques, methods and devices known to those skilled in the relevant art can not be discussed in detail, but in appropriate cases, the techniques, methods and devices should be considered as part of the specification.
[0031] It should be noted that: similar numbers and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further discussed in the subsequent drawings.
[0032] The embodiments of the present disclosure can be applied to terminal devices, computer systems, servers, and other electronic devices, which can operate with many other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with terminal devices, computer systems, servers, and other electronic devices include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems, and the like.
[0033] Terminal devices, computer systems, servers, and other electronic devices can be described in the general context of computer system executable instructions, such as program modules, executed by the computer system. Generally, program modules can include routines, programs, objects, components, logic, data structures, and the like, which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in a distributed cloud computing environment, in which tasks are performed by remote processing devices that are linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computer system storage media, including storage devices.
[0034] SUMMARY
[0035] In a real-time voice interaction system, in the case of user face being blocked, face image being unclear, and the like, the visual features obtained based on image recognition become invalid interference input, and the performance of the multi-modal voice recognition method will decrease significantly, so the system needs to realize real-time switching of single and multi-modal recognition according to the quality of the photographed image, and solve the voice recognition problem in the case of visual blocking. The current multi-modal voice recognition method does not judge the quality of the user's face image, and cannot avoid the influence of invalid visual features on the voice recognition effect.
[0036] To solve this problem, the embodiments of the present disclosure provide a voice signal processing method, which can realize real-time image quality judgment on a photographed image sequence, and output voice signals extracted by a single-modal processing mode or voice signals extracted by a multi-modal processing mode according to the judgment result, thereby reducing the influence of visual features on voice recognition accuracy when the image quality does not meet the conditions.
[0037] Exemplary System
[0038] Figure 1 An exemplary system architecture 100 of a voice signal processing method or a voice signal processing apparatus to which embodiments of the present disclosure can be applied is shown.
[0039] like Figure 1 As shown, the system architecture 100 may include a terminal device 101, a network 102, a server 103, an image acquisition device 104, and a voice acquisition device 105. The network 102 serves as the medium for providing a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0040] Image acquisition device 104 and voice acquisition device 105 can be installed within a target space, which can be of various types, such as the interior of a vehicle or a house. Image acquisition device 104 and voice acquisition device 105 are used to acquire image sequences and voice signals for a target user. The acquired image sequences and voice signals can be saved to terminal device 101 or sent by terminal device 101 to server 103.
[0041] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as search applications, browser applications, instant messaging tools, etc.
[0042] Terminal device 101 can be various electronic devices, including but not limited to mobile terminals such as vehicle terminals, mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), etc., as well as fixed terminals such as digital TVs, desktop computers, etc.
[0043] Server 103 can be a server that provides various services, such as a backend server that processes image sequences and audio signals uploaded by terminal device 101. The backend server can process the received image sequences and audio signals according to a first audio processing method and a second audio processing method, and output the target audio segment signal.
[0044] It should be noted that the voice signal processing method provided in the embodiments of this disclosure can be executed by the server 103 or by the terminal device 101. Accordingly, the voice signal processing device can be set in the server 103 or in the terminal device 101.
[0045] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included. If image sequences and voice signals do not need to be acquired remotely, the above system architecture may exclude the network and include only servers or terminal devices.
[0046] Exemplary Method
[0047] Figure 2 is a flowchart of a voice signal processing method provided by an example embodiment of the present disclosure. The embodiment can be applied on an electronic device (such as a terminal device 101 or a server 103) as shown in Figure 1 , the method comprises the following steps: Figure 2
[0048] Step 201: Obtain a voice signal and an image sequence in a target space.
[0049] In the embodiment, the electronic device can obtain a voice signal and an image sequence in a target space. The target space can be any type of space, for example, a vehicle interior space, a room interior space, etc. The voice signal and the image sequence can be collected for a user in the target space.
[0050] The image sequence can be obtained by an image collection device 104 as shown in Figure 1 , which can take pictures for a certain position to obtain the image sequence. The voice signal can be collected by a voice collection device 105 as shown in Figure 1 .
[0051] It should be noted that the voice signal and the image sequence are corresponding in the time scale, that is, the starting time and the ending time of the voice signal are the same as the starting time and the ending time of the image sequence.
[0052] Step 202: Based on the voice signal, a first voice segment signal is extracted from the voice signal by a first voice processing mode.
[0053] In the embodiment, the electronic device can extract a first voice segment signal from the voice signal by a first voice processing mode based on the voice signal.
[0054] The first voice processing mode is a mode of processing the voice signal alone, which can also be referred to as a single-mode voice processing mode. The first voice processing mode can intercept the input single-mode voice signal, cut the continuous audio stream form of the voice signal into multiple voice segment signals, and set a start mark and an end mark at the start point and the end point of each voice segment signal.
[0055] Optionally, the first voice processing manner can be implemented by using a VAD (Voice Activity Detection) method. The purpose of the VAD is to identify and eliminate long silent signal segments from the audio stream, so as to extract the valid voice segment signal (query). For example, if the user utters the voice "today how is the weather", the VAD method can extract the voice segment signal corresponding to "today how is the weather" from the collected voice signal. Usually, a start mark (vad_start) is set before the audio frame corresponding to the word "today", and an end mark (vad_end) is set after the audio frame corresponding to the word "style". The electronic device can determine the voice segment signal between the start mark and the end mark as the first voice segment signal and extract it.
[0056] In step 203, the second voice segment signal is extracted from the voice signal based on the voice signal and the image sequence by using a second voice processing manner.
[0057] In this embodiment, the electronic device can extract the second voice segment signal from the voice signal based on the voice signal and the image sequence by using the second voice processing manner.
[0058] The second voice processing manner is a manner of combining the image sequence and the voice signal for voice recognition, and is usually referred to as a multi-modal voice processing manner. Usually, the execution process of the second voice processing manner is as follows: the voice signal is input into a voice feature extraction network to obtain voice feature data, the images included in the image sequence are input into an image feature extraction network to obtain image sequence feature data, the image sequence feature data represents the motion features of the target part (such as the lip, eyeball, etc.) of the user; then the voice feature data and the image sequence feature data are spliced and input into a feature fusion network to obtain a mask value; finally, the above voice signal is multiplied by the mask value, so that the second voice segment signal can be obtained.
[0059] In step 204, it is determined whether the current voice signal processing state meets the voice signal output condition.
[0060] In this embodiment, the electronic device can determine whether the current voice signal processing state meets the voice signal output condition. The voice signal processing state can include a busy state and an idle state. Usually, if the current voice signal processing state is the idle state, it can be determined that the current voice signal processing state meets the voice signal output condition.
[0061] The busy state indicates that the electronic device is currently processing some information and cannot output the voice segment signal. For example, if the electronic device is currently processing a voice signal, i.e., an operation of extracting a first voice segment signal or a second voice segment signal from the voice signal has not ended, it is determined that the voice signal processing state is the busy state. If the voice signal is not currently being processed, i.e., the first voice segment signal or the second voice segment signal has been obtained completely, it is determined that the voice signal processing state is the idle state. Alternatively, if the electronic device is currently outputting a voice segment signal, it is determined that the voice signal processing state is the busy state. If no voice segment signal is currently being output, it is determined that the voice signal processing state is the idle state. Alternatively, if a current target resource occupation rate (e.g., CPU occupation rate) of the electronic device for performing voice recognition exceeds a preset occupation rate, it is determined that the voice signal processing state is the busy state. Otherwise, it is determined that the voice signal processing state is the idle state.
[0062] In step 205, in response to the voice signal processing state meeting the voice signal output condition, image quality information of the image sequence is determined.
[0063] In this embodiment, the electronic device can determine the image quality information of the image sequence in response to the voice signal processing state meeting the voice signal output condition.
[0064] The image quality information can be information indicating whether the overall quality of the images included in the image sequence is qualified, e.g., a number 1 indicating qualification and a number 0 indicating disqualification. Generally, the electronic device can determine an image quality score of each image in the image sequence, and then determine the image quality information of the image sequence according to the image quality scores of each image according to a preset statistical method. For example, an average value (or a median value, etc.) of the image quality scores can be determined. If the average value is greater than or equal to a preset value, the image quality information indicating that the image quality is qualified is generated. If the average value is less than the preset value, the image quality information indicating that the image quality is disqualified is generated.
[0065] Optionally, the electronic device can determine the sharpness of each image in the image sequence according to a method of determining image sharpness, and determine the sharpness as the image quality score. The electronic device can also determine the position of a target part of the user from each image according to a target detection method, and then determine the proportion of the area of the target part to the total area of the image, and determine the proportion as the image quality score. The electronic device can also determine the completeness value of the display of the target part as the image quality score. As an example, the completeness value of the display of the target part can be obtained by determining the number of key points on the target part, i.e., taking the ratio of the number of key points included in the currently displayed target part to the number of key points included in the target part when it is displayed completely as the completeness value.
[0066] Optionally, a plurality of image quality scores can be obtained according to a plurality of methods for determining image quality scores, and then the image quality scores are fused (for example, average, weighted sum, etc.) to obtain the image quality information.
[0067] It should be noted that, since the images included in the image sequence are acquired in sequence in real time, the image quality score of each image can be determined in real time after each image is acquired, thereby reducing the delay of determining the image quality information of the image sequence.
[0068] In step 206, the target speech segment signal is determined from the first speech segment signal and the second speech segment signal based on the image quality information of the image sequence, and the target speech segment signal is output.
[0069] In this embodiment, the electronic device can determine the target speech segment signal from the first speech segment signal and the second speech segment signal based on the image quality information of the image sequence, and output the target speech segment signal. The output target speech segment signal can be further used for speech recognition to obtain a recognition result, and the recognition result can be further used for human-computer interaction.
[0070] Optionally, the electronic device can determine the second speech segment signal as the target speech segment signal in response to the image quality information of the image sequence indicating that the overall quality of the image sequence is qualified, and determine the first speech segment signal as the target speech segment signal in response to the image quality information of the image sequence indicating that the overall quality of the image sequence is unqualified. By outputting the first speech segment signal as the target speech segment signal when the image quality information indicates that the overall quality of the image sequence is unqualified, the accuracy of the output speech segment signal can be avoided to be reduced when the overall image quality of the image sequence is unqualified, thereby improving the accuracy of subsequent speech recognition.
[0071] The method provided by the above embodiments of the present disclosure extracts a first voice segment signal from the voice signal based on the voice signal by a first voice processing manner, extracts a second voice segment signal from the voice signal based on the voice signal and the image sequence by a second voice processing manner, determines image quality information of the image sequence in response to the current voice signal processing state meeting a voice signal output condition, and finally determines a target voice segment signal from the first voice segment signal and the second voice segment signal based on the image quality information of the image sequence, and outputs the target voice segment signal. The embodiments of the present disclosure automatically select the first voice segment signal extracted by the single-mode voice processing manner or the second voice segment signal extracted by the multi-mode voice processing manner as the target voice segment signal required for voice recognition according to the image quality of the photographed image sequence, so that when the output target voice segment signal is recognized, the source of the output voice segment signal is selected according to the image quality, thereby helping to improve the accuracy of voice recognition.
[0072] In some optional implementations, as shown in Figure 3 Step 203 includes:
[0073] Step 2031, determining audio feature data of the voice signal based on a preset audio feature extraction network.
[0074] Optionally, the structure of the audio feature extraction network can include but is not limited to RNN (Recurrent Neural Network), LSTM (Long Short-Term Memory), UNet (U-shaped network), Complex UNet, and Transformer architecture based on self-attention mechanism and cross-domain attention mechanism.
[0075] Step 2032, determining image sequence feature data of the image sequence based on a preset image sequence feature extraction network.
[0076] The image sequence feature data represents the action state of the target part (such as the lip, the eye, etc.) of the user in the image. For example, the image sequence feature extraction network can determine the lip region from each frame of image, and then determine the lip contour feature data (such as including the distance between the corners of the mouth, the distance between the upper and lower lips, etc.) of each lip region, and combine the lip contour feature data of each lip region into image sequence feature data representing the lip action state change feature. For another example, the image sequence feature extraction network can determine the eye region from each frame of image, and then determine the eye feature data (such as including the line of sight angle, the opening degree of the eye, etc.) according to the eye region, and combine the eye feature data into image sequence feature data representing the eye action state change feature.
[0077] Step 2033, merge the audio feature data and the image sequence feature data, and input the merged data into a pre-trained feature fusion network to obtain mask data.
[0078] The above-mentioned merging refers to directly merging the channels included in the audio feature data and the image sequence feature data. Generally, the feature fusion network can include a fusion sub-network and a decoding network. The fusion sub-network can perform a feature fusion method to fuse the audio feature data and the image sequence feature data to obtain fused feature data. As an example, the feature fusion method can include any one of the following: a concat feature fusion method, an elemwise_add feature fusion method, an attention feature fusion method, etc.
[0079] The decoding network can decode the fused feature data to obtain the mask data. Generally, the decoding network can have an up-sampling function. The fused feature data is usually small-scale feature data. Through the decoding network, the small-scale fused feature data can be up-sampled to obtain mask data with the same scale as the frequency domain data of the speech signal.
[0080] Step 2034, based on the mask data, extracting a second speech segment signal from the speech signal.
[0081] Specifically, the mask data can be multiplied by the frequency domain data (e.g., obtained by performing short-time Fourier transform on the speech signal) of the speech signal to obtain a frequency domain second speech segment signal. Alternatively, the frequency domain second speech segment signal can be processed, such as inverse Fourier transform, to obtain a time domain second speech segment signal.
[0082] The embodiment extracts the second speech segment signal from the speech signal by using the audio feature extraction network, the image sequence feature extraction network, and the feature fusion network, fully utilizes the high prediction accuracy of the neural network, and improves the accuracy of extracting the second speech segment signal from the speech signal according to the multi-modal speech processing method.
[0083] In some optional implementations, as shown in Figure 4 Step 204 includes:
[0084] Step 2041, determining whether there is currently a speech segment signal being processed according to the current output channel corresponding speech processing mode.
[0085] The output channel refers to a process of processing a speech signal according to a speech processing mode and outputting a speech segment signal. In this embodiment, two output channels are included. One is a process of extracting a first speech segment from the speech signal by a first speech processing mode and outputting the first speech segment, that is, a single-mode speech processing process. The other is a process of extracting a second speech segment from the speech signal by a second speech processing mode and outputting the first speech segment, that is, a multi-mode speech processing process. Generally, the speech segment signals are arranged in sequence frame by frame. The starting frame of each speech segment signal includes a starting mark, and the ending frame includes an ending mark. The electronic device can monitor in real time whether the speech segment signal obtained by the speech signal processing mode corresponding to the current output channel includes the ending mark. If the ending mark is included, it is determined that there is no speech segment signal being processed. If the ending mark is not included, it is determined that there is a speech segment signal being processed. Alternatively, since the speech segment signal is output frame by frame in sequence, the electronic device can detect in real time whether the frame being output by the current output channel includes the ending mark. If the ending mark is included, it is determined that the current output speech segment signal has been output completely, and there is no speech segment signal being processed at this time. If the frame being output does not include the ending mark, it is determined that there is a speech segment signal being processed.
[0086] Generally, if it is determined that there is a speech segment signal being processed, the current output channel is kept unchanged, and the speech signal is continuously processed and the speech segment signal is output.
[0087] In step 2042, in response to the fact that there is no speech segment signal being processed according to the speech processing mode corresponding to the current output channel, it is determined that the current speech signal processing state meets the speech signal output condition.
[0088] If there is no speech segment signal being processed at present, steps 205-206 can be continuously executed. After one speech segment signal is output completely, the target speech segment signal is determined from the first speech segment signal and the second speech segment signal according to the image quality information, that is, the output channel of the speech segment signal to be output next is switched after the complete speech segment signal is output, so that the integrity of the output speech segment signal is maintained, and the influence of switching the source of the speech segment signal on subsequent speech recognition is avoided in the process of outputting the speech segment signal.
[0089] In some optional implementations, as shown in Figure 5 Step 205 includes:
[0090] In step 2051, for each frame of image in the image sequence, it is determined whether the target part of the user is included in the image.
[0091] The target part can include, but is not limited to, at least one of the following: the entire face, lips, eyes, hands, etc. Generally, whether the target part is included in each frame of image can be determined according to a neural network-based target detection method.
[0092] In step 2052, first image quality information indicating that the image quality of the image is unqualified is generated in response to the image not containing the target part.
[0093] In step 2053, the recognizability of the target part is determined in response to the image containing the target part.
[0094] The recognizability can be determined by determining the sharpness of the target part, the proportion of the area of the target part to the total area of the image, a completeness value of the target part, etc.
[0095] Optionally, the electronic device can determine the sharpness of the target part in the image according to a method of determining image sharpness, and determine the sharpness as the recognizability. The electronic device can also determine the proportion of the area of the target part to the total area of the image, and determine the proportion as the recognizability. The electronic device can also determine a completeness value of the target part as the recognizability, and the completeness value of the target part can be obtained by determining the number of key points on the target part, i.e., the ratio of the number of key points included in the currently displayed target part to the number of key points included in the target part when it is completely displayed as the completeness value.
[0096] Optionally, the electronic device can also weight and sum the above-mentioned sharpness, area proportion, completeness value, etc. according to a preset weight to obtain the recognizability.
[0097] In step 2054, second image quality information indicating that the image quality of the image is qualified is generated in response to the recognizability meeting the recognizability condition.
[0098] Generally, when the recognizability is greater than or equal to a preset threshold, it is determined that the recognizability condition is met.
[0099] In step 2055, first image quality information indicating that the image quality of the image is unqualified is generated in response to the recognizability not meeting the recognizability condition.
[0100] Generally, when the recognizability is less than a preset threshold, it is determined that the recognizability condition is not met.
[0101] In step 2056, the image quality information of the image sequence is determined based on the number of obtained first image quality information and the number of second image quality information.
[0102] Optionally, when the number of the first image quality information is greater than the number of the second image quality information, or the number of the first image quality information is greater than or equal to a preset number threshold, it indicates that the number of unqualified images is relatively large, and image quality information indicating that the overall quality of the image sequence is unqualified can be generated; when the number of the first image quality information is less than or equal to the number of the second image quality information, or the number of the second image quality information is greater than or equal to the preset number threshold, it indicates that the number of qualified images is relatively large, and image quality information indicating that the overall quality of the image sequence is qualified can be generated.
[0103] It should be understood that the above steps 2051-2055 are performed for each frame of image in the image sequence, that is, steps 2051-2055 are performed once for one frame of image, and step 2056 is performed after the image quality determination for each frame of image in the image sequence.
[0104] Optionally, since the images included in the image sequence are acquired in real time in sequence, the above steps 2051-2055 can be performed on each image in real time after the image is acquired, and step 2056 is performed when it is determined that the voice signal processing state meets the voice signal output condition, thereby reducing the delay of determining the image quality information of the image sequence.
[0105] The embodiment determines whether the image quality of each frame of image in the image sequence is qualified by performing target part detection on each frame of image in the image sequence and determining the recognizable degree of the target part, thereby accurately determining whether the overall image quality of the image sequence is qualified, and further improving the accuracy of multi-modal voice recognition.
[0106] In some optional implementations, as shown in Figure 6 Step 2053 includes:
[0107] Step 20531, determining a target region containing the target part from the image.
[0108] Specifically, a target detection method can be used to determine a target region containing the target part from the image. Optionally, the target region can be a rectangular region containing the target part, or a contour containing the target part segmented from the image, etc.
[0109] Step 20532, performing image quality detection on the target region by using a pre-trained target part quality detection model to obtain the recognizable degree of the target part.
[0110] The target part quality detection model can be a model trained in advance based on a machine learning method, and the type of the target part quality detection model can include but is not limited to at least one of the following: a key point detection model, a posture detection model, a gaze detection model, etc.
[0111] As an example, when the target part is a face, the electronic device can utilize a key point detection model to detect key points of the face displayed in the image, and then determine the recognizability according to the detected face key points. For example, the electronic device can determine a ratio of the number of detected face key points to the number of normally displayed face key points as the recognizability, which represents the completeness of the face displayed in the image.
[0112] As another example, when the target part is a face, the electronic device can utilize a pose detection model to detect a pose angle of the face, and then determine the recognizability according to the pose angle. For example, the electronic device can determine the inverse of the deviation of the detected pose angle from the pose angle in the ideal case (i.e., the face directly facing the lens of the image acquisition device) as the recognizability, that is, the smaller the deviation, the greater the recognizability.
[0113] The embodiment can effectively utilize the high-precision target part quality detection model trained by the machine learning method, thereby helping to improve the accuracy of quality determination of the image.
[0114] Exemplary Apparatus
[0115] Figure 7 is a structural schematic diagram of a voice signal processing apparatus provided by an example embodiment of the present disclosure. The embodiment can be applied on an electronic device, such as a mobile phone, a tablet computer, a smart speaker, etc. Figure 7 As shown in the figure, the voice signal processing apparatus includes: an acquisition module 701 configured to acquire a voice signal and an image sequence in a target space; a first extraction module 702 configured to extract a first voice segment signal from the voice signal by a first voice processing manner based on the voice signal; a second extraction module 703 configured to extract a second voice segment signal from the voice signal by a second voice processing manner based on the voice signal and the image sequence; a first determination module 704 configured to determine whether a current voice signal processing state meets a voice signal output condition; a second determination module 705 configured to determine image quality information of the image sequence in response to the voice signal processing state meeting the voice signal output condition; and an output module 706 configured to determine a target voice segment signal from the first voice segment signal and the second voice segment signal based on the image quality information of the image sequence, and output the target voice segment signal.
[0116] In the embodiment, the acquisition module 701 can acquire a voice signal and an image sequence in a target space. The target space can be any type of space, such as a vehicle interior space, a room interior space, etc. The voice signal and the image sequence can be acquired for a user in the target space.
[0117] The above image sequence can be acquired by, for example,Figure 1 The image acquisition device 104 shown in the figure is used to capture images, and the image acquisition device 104 can capture images for a certain position to obtain the image sequence. The voice signal can be captured by the voice acquisition device 105 shown in the figure. Figure 1
[0118] It should be noted that the voice signal and the image sequence correspond in the time scale, that is, the start time and the end time of the voice signal are the same as the start time and the end time of the image sequence.
[0119] In this embodiment, the first extraction module 702 can extract a first voice segment signal from the voice signal by a first voice processing manner based on the voice signal.
[0120] The first voice processing manner is a manner of processing the voice signal alone, and can also be referred to as a single-mode voice processing manner. The first voice processing manner can intercept the input single-mode voice signal, cut the continuous audio stream form of the voice signal into a plurality of voice segment signals, and set a start mark and an end mark at the start point and the end point of each voice segment signal.
[0121] In this embodiment, the second extraction module 703 can extract a second voice segment signal from the voice signal by a second voice processing manner based on the voice signal and the image sequence.
[0122] The second voice processing manner is a manner of combining the image sequence and the voice signal for voice recognition, and is also referred to as a multi-mode voice processing manner. Generally, the execution process of the second voice processing manner is as follows: inputting the voice signal into a voice feature extraction network to obtain voice feature data, inputting the images included in the image sequence into an image feature extraction network to obtain image sequence feature data, the image sequence feature data representing the motion features of the target part (such as the lip, eyeball, etc.) of the user; then inputting the voice feature data and the image sequence feature data into a feature fusion network after splicing to obtain a mask value; finally, multiplying the voice signal by the mask value, so that the second voice segment signal can be obtained.
[0123] In this embodiment, the first determination module 704 can determine whether the current voice signal processing state meets the voice signal output condition. The voice signal processing state can include a busy state and an idle state, and generally, if the current voice signal processing state is the idle state, it can be determined that the current voice signal processing state meets the voice signal output condition.
[0124] The busy state indicates that the electronic device is currently processing some information and cannot output the voice segment signal. For example, if the electronic device is currently processing a voice signal, that is, the operation of extracting a first voice segment signal or a second voice segment signal from the voice signal has not ended, it is determined that the voice signal processing state is the busy state. If the voice signal is not currently being processed, that is, the complete first voice segment signal or second voice segment signal has been obtained, it is determined that the voice signal processing state is the idle state. Alternatively, if the electronic device is currently outputting a voice segment signal, it is determined that the voice signal processing state is the busy state. If no voice segment signal is currently being output, it is determined that the voice signal processing state is the idle state. Alternatively, if the current target resource occupation rate (for example, CPU occupation rate) of the electronic device for performing voice recognition exceeds a preset occupation rate, it is determined that the voice signal processing state is the busy state. Otherwise, it is determined that the voice signal processing state is the idle state.
[0125] In this embodiment, the second determination module 705 can determine the image quality information of the image sequence in response to the voice signal processing state meeting the voice signal output condition.
[0126] The image quality information can be information indicating whether the overall quality of the images included in the image sequence is qualified, for example, the number 1 indicates qualified, and the number 0 indicates unqualified. Generally, the electronic device can determine the image quality score of each image in the image sequence, and then determine the image quality information of the image sequence according to the image quality score of each image according to a preset statistical method. For example, the average value (or median value, etc.) of each image quality score can be determined. If the average value is greater than or equal to a preset value, the image quality information indicating that the image quality is qualified is generated. If the average value is less than the preset value, the image quality information indicating that the image quality is unqualified is generated.
[0127] In this embodiment, the output module 706 can determine the target voice segment signal from the first voice segment signal and the second voice segment signal based on the image quality information of the image sequence, and output the target voice segment signal. The output target voice segment signal can be further used for voice recognition to obtain a recognition result, and the recognition result can be further used for human-computer interaction.
[0128] Referring to Figure 8 , Figure 8 is a structural schematic diagram of a voice signal processing device provided by another exemplary embodiment of the present disclosure.
[0129] In some optional implementation manners, the second extraction module 703 comprises: a first determination unit 7031, configured to determine audio feature data of the speech signal based on a preset audio feature extraction network; a second determination unit 7032, configured to determine image sequence feature data of the image sequence based on a preset image sequence feature extraction network; a fusion unit 7033, configured to combine the audio feature data and the image sequence feature data, and input the combined data into a pre-trained feature fusion network to obtain mask data; and an extraction unit 7034, configured to extract a second speech segment signal from the speech signal based on the mask data.
[0130] In some optional implementation manners, the first determination module 704 comprises: a third determination unit 7041, configured to determine whether there is currently a speech segment signal being processed in a manner corresponding to the current output channel; and a fourth determination unit 7042, configured to determine that the current speech signal processing state meets the speech signal output condition in response to there being no speech segment signal being output in the manner corresponding to the current output channel.
[0131] In some optional implementation manners, the second determination module 705 comprises: a fifth determination unit 7051, configured to determine, for each image in the image sequence, whether the image contains a target part of the user; a first generation unit 7052, configured to generate first image quality information indicating that the image quality of the image does not meet a quality requirement in response to the image not containing the target part; a sixth determination unit 7053, configured to determine an identifiable degree of the target part in response to the image containing the target part; a second generation unit 7054, configured to generate second image quality information indicating that the image quality of the image meets the quality requirement in response to determining that the identifiable degree meets an identifiable condition; a third generation unit 7055, configured to generate the first image quality information indicating that the image quality of the image does not meet the quality requirement in response to determining that the identifiable degree does not meet the identifiable condition; and a seventh determination unit 7056, configured to determine image quality information of the image sequence based on a quantity of the first image quality information and a quantity of the second image quality information.
[0132] In some optional implementation manners, the sixth determination unit 7053 comprises: a determination subunit 70531, configured to determine a target region containing the target part from the image; and a detection subunit 70532, configured to perform image quality detection on the target region by using a pre-trained target part quality detection model to obtain the identifiable degree of the target part.
[0133] In some optional implementation, the output module 706 includes: an eighth determination unit 7061, configured to determine the second speech segment signal as the target speech segment signal in response to the image quality information of the image sequence indicating that the overall quality of the image sequence is qualified; and a ninth determination unit 7062, configured to determine the first speech segment signal as the target speech segment signal in response to the image quality information of the image sequence indicating that the overall quality of the image sequence is unqualified.
[0134] The speech signal processing apparatus provided by the above-mentioned embodiments of the present disclosure extracts the first speech segment signal from the speech signal through the first speech processing manner, extracts the second speech segment signal from the speech signal through the second speech processing manner based on the speech signal and the image sequence, determines the image quality information of the image sequence in response to the current speech signal processing state meeting the speech signal output condition, and finally determines the target speech segment signal from the first speech segment signal and the second speech segment signal based on the image quality information of the image sequence, and outputs the target speech segment signal. The embodiments of the present disclosure realize automatic selection of the first speech segment signal extracted by the single-mode speech processing manner or the second speech segment signal extracted by the multi-mode speech processing manner as the target speech segment signal required by speech recognition according to the image quality of the captured image sequence, so that the source of the output speech segment signal is selected according to the image quality when the output target speech segment signal is recognized, and thus the accuracy of speech recognition is improved.
[0135] Exemplary electronic device
[0136] Hereinafter, an electronic device according to embodiments of the present disclosure will be described with reference to the accompanying drawings. Figure 9 The electronic device can be any one or both of a terminal device 101 and a server 103 as shown in Figure 1 or a stand-alone device independent of the terminal device 101 and the server 103, which can communicate with the terminal device 101 and the server 103 to receive input signals collected therefrom.
[0137] Figure 9 A block diagram of an electronic device according to embodiments of the present disclosure is shown.
[0138] As shown in Figure 9 the electronic device 900 includes one or more processors 901 and a memory 902.
[0139] The processor 901 can be a central processing unit (CPU) or other form of processing unit having data processing and / or instruction execution capabilities, and can control other components in the electronic device 900 to perform desired functions.
[0140] The memory 902 can include one or more computer program products that can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory, for example, can include random access memory (RAM), cache memory, and / or the like. The non-volatile memory, for example, can include read-only memory (ROM), hard disk drives, solid-state drives, and / or the like. The computer-readable storage media can store one or more computer program instructions executable by the processor 901 to implement the voice signal processing method of various embodiments of the present disclosure described above and / or other desired functions. Various contents such as image sequences, voice signals, and the like can also be stored in the computer-readable storage media.
[0141] In one example, the electronic device 900 can further include an input device 903 and an output device 904, which are interconnected through a bus system and / or other forms of connection mechanisms (not shown).
[0142] For example, when the electronic device is the terminal device 101 or the server 103, the input device 903 can be a camera, a microphone, a mouse, a keyboard, and the like, for inputting image sequences, voice signals, various commands, and the like. When the electronic device is a stand-alone device, the input device 903 can be a communication network connector for receiving inputted image sequences, voice signals, various commands from the terminal device 101 and the server 103.
[0143] The output device 904 can output various information including the target voice segment signal and the like to the outside. The output device 904 can include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto, and the like.
[0144] Of course, in order to simplify, Figure 9 Only some of the components of the electronic device 900 related to the present disclosure are shown in FIG. 9, and components such as buses, input / output interfaces, and the like are omitted. In addition, the electronic device 900 can further include any other appropriate components according to specific application cases.
[0145] Exemplary Computer Program Product and Computer-Readable Storage Medium
[0146] In addition to the above-described method and device, an embodiment of the present disclosure can be a computer program product including computer program instructions that, when executed by a processor, cause the processor to perform the steps of the voice signal processing method according to various embodiments of the present disclosure described in the above "Exemplary Methods" section of the specification.
[0147] The computer program product can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computing device, partly on the user's device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device or entirely on the remote computing device or server.
[0148] Furthermore, embodiments of the present disclosure can also be a computer readable storage medium, having stored thereon computer program instructions which, when executed by a processor, cause the processor to carry out the steps described in the above "Exemplary Method" section of the present specification for the voice signal processing method according to various embodiments of the present disclosure.
[0149] The computer readable storage medium can be any combination of one or more computer readable media. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. The computer readable storage medium can include, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. More specific examples (a non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0150] The above describes the basic principles of the present disclosure in combination with specific embodiments, but it should be noted that the advantages, benefits, effects and the like mentioned in the present disclosure are only examples and are not limiting, and these advantages, benefits, effects and the like cannot be considered as the must-have of each embodiment of the present disclosure. In addition, the above specific details are only for the purpose of example and understanding, and the above details do not limit the present disclosure to the must-use specific details.
[0151] Each embodiment in the present specification is described in a progressive manner, and each embodiment focuses on the difference from other embodiments, and the same or similar parts between each embodiment can be referred to each other. For system embodiments, since they are basically corresponding to method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.
[0152] The block diagrams of devices, apparatuses, equipment, systems referred to in this disclosure are merely illustrative examples and are not intended to require or imply that the connection, arrangement, configuration must be as shown in the block diagrams. These devices, apparatuses, equipment, systems can be connected, arranged, configured in any manner as will be appreciated by those skilled in the art. Words such as "include," "contain," "have," and the like are open-ended words that are to be interpreted to mean "including but not limited to," and are not to be interpreted as limiting the described embodiment to features, elements, and / or steps disclosed herein. The words "or" and "and" as used herein are to be interpreted as the word "and / or," and are not to be interpreted as requiring both features, elements, and / or steps disclosed herein. The word "such as" as used herein is to be interpreted as the phrase "such as but not limited to," and is not to be interpreted as limiting the described embodiment to features, elements, and / or steps disclosed herein.
[0153] The methods and apparatuses of this disclosure can be implemented in a number of ways. For example, the methods and apparatuses of this disclosure can be implemented using software, hardware, firmware, or any combination of these. The above described order of steps for the methods is merely illustrative, and the steps of the methods of this disclosure are not limited to the order specifically described above unless otherwise specifically stated. Furthermore, in some embodiments, the disclosure can also be implemented as a program recorded in a recording medium, which includes machine readable instructions for implementing the methods according to the disclosure. Thus, the disclosure also covers a recording medium storing a program for executing the methods according to the disclosure.
[0154] It is also important to note that the devices, equipment, and methods of this disclosure can be embodied in a variety of ways. These variations are contemplated as being within the scope of the present disclosure.
[0155] The above description of the disclosed aspects is given for illustrative purposes and is not intended to limit the scope of the disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0156] The above description has been given for illustrative and descriptive purposes. In addition, this description is not intended to limit the embodiments of the disclosure to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those of skill in the art will recognize certain variations, modifications, changes, additions, and sub-combinations thereof.
Claims
1. A voice signal processing method, comprising: obtaining a voice signal and an image sequence in a target space; based on the voice signal, extracting a first voice segment signal from the voice signal by a first voice processing manner; based on the voice signal and the image sequence, extracting a second voice segment signal from the voice signal by a second voice processing manner; determining whether a current voice signal processing state meets a voice signal output condition; in response to the voice signal processing state meeting the voice signal output condition, determining image quality information of the image sequence; based on the image quality information of the image sequence, determining a target voice segment signal from the first voice segment signal and the second voice segment signal, and outputting the target voice segment signal; wherein the determining of the target voice segment signal from the first voice segment signal and the second voice segment signal based on the image quality information of the image sequence comprises: in response to the image quality information of the image sequence indicating that the overall quality of the image sequence is qualified, determining the second voice segment signal as the target voice segment signal; in response to the image quality information of the image sequence indicating that the overall quality of the image sequence is unqualified, determining the first voice segment signal as the target voice segment signal.
2. The method of claim 1, wherein, the extracting of the second voice segment signal from the voice signal based on the voice signal and the image sequence by the second voice processing manner comprises: determining audio feature data of the voice signal based on a preset audio feature extraction network; determining image sequence feature data of the image sequence based on a preset image sequence feature extraction network; merging the audio feature data and the image sequence feature data, and inputting the merged data into a pre-trained feature fusion network to obtain mask data; extracting the second voice segment signal from the voice signal based on the mask data.
3. The method of claim 1, wherein, the determining of whether the current voice signal processing state meets the voice signal output condition comprises: determining whether there is currently a voice segment signal being processed according to a voice processing manner corresponding to a current output channel; in response to there being no voice segment signal being processed according to the voice processing manner corresponding to the current output channel, determining that the current voice signal processing state meets the voice signal output condition.
4. The method of claim 1, wherein, the determining of the image quality information of the image sequence comprises: for each frame of image in the image sequence, determining whether the image contains a target part of a user; in response to the image not containing the target part, generating first image quality information indicating that the image quality of the image is unqualified; in response to the image containing the target part, determining an identifiable degree of the target part; in response to determining that the identifiable degree meets an identifiable condition, generating second image quality information indicating that the image quality of the image is qualified; in response to determining that the identifiable degree does not meet the identifiable condition, generating first image quality information indicating that the image quality of the image is unqualified; Determine image quality information of the image sequence based on the obtained quantity of first image quality information and quantity of second image quality information.
5. The method of claim 4, wherein, The determination of the recognizability of the target part includes: Determining a target region containing the target part from the image; Using a pre-trained target part quality detection model to perform image quality detection on the target region to obtain the recognizability of the target part.
6. A speech signal processing apparatus, comprising: An acquisition module configured to acquire a speech signal and an image sequence in a target space; A first extraction module configured to extract a first speech segment signal from the speech signal by a first speech processing manner based on the speech signal; A second extraction module configured to extract a second speech segment signal from the speech signal by a second speech processing manner based on the speech signal and the image sequence; A first determination module configured to determine whether a current speech signal processing state meets a speech signal output condition; A second determination module configured to determine image quality information of the image sequence in response to the speech signal processing state meeting the speech signal output condition; An output module configured to determine a target speech segment signal from the first speech segment signal and the second speech segment signal based on the image quality information of the image sequence, and output the target speech segment signal; The output module includes: an eighth determination unit configured to determine the second speech segment signal as the target speech segment signal in response to the image quality information of the image sequence indicating that the overall quality of the image sequence is qualified; and a ninth determination unit configured to determine the first speech segment signal as the target speech segment signal in response to the image quality information of the image sequence indicating that the overall quality of the image sequence is unqualified.
7. The apparatus of claim 6, wherein, The second extraction module includes: A first determination unit configured to determine audio feature data of the speech signal based on a preset audio feature extraction network; A second determination unit configured to determine image sequence feature data of the image sequence based on a preset image sequence feature extraction network; A fusion unit configured to combine the audio feature data and the image sequence feature data, and input the combined data into a pre-trained feature fusion network to obtain mask data; An extraction unit configured to extract the second speech segment signal from the speech signal based on the mask data.
8. A computer-readable storage medium, the storage medium storing a computer program for executing the method of any one of claims 1-5.
9. An electronic device, comprising: A processor; A memory for storing executable instructions of the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method of any one of claims 1-5.
Citation Information
Patent Citations
Model driving method and device based on emotion recognition
CN115049016A