Voice operation method of device, apparatus, and electronic device

By autonomously detecting and registering the voiceprint information of multiple users in the device, the problem of poor speech recognition performance in multi-user scenarios is solved, and efficient speech separation and recognition are achieved in noisy environments.

CN115050375BActive Publication Date: 2026-01-02HUAWEI TECH CO LTD

Patent Information

Application Number
CN202110218188.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-26
Publication Date
2026-01-02
Estimated Expiration
2041-02-26

AI Technical Summary

Technical Problem

Existing speech recognition technologies perform poorly in noisy environments and multi-user speaking scenarios, especially in multi-user speaking scenarios where speech recognition performance deteriorates significantly. Furthermore, existing solutions require pre-registration of voiceprints, resulting in a poor user experience and insufficient adaptability and usability.

Method used

The system captures video through a camera, detects facial images, extracts lip and facial features, determines time intervals, extracts audio segments, performs speech recognition using voiceprint information, autonomously registers and manages the voiceprint information of multiple users, and separates speech one by one to adapt to multi-user scenarios.

Benefits of technology

When multiple users use the device simultaneously, the system autonomously extracts the voiceprint information of each user and separates the speech one by one, which improves the accuracy of speech recognition and user experience, reduces power consumption, and adapts to the speech recognition effect in multi-user scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115050375B_ABST
    Figure CN115050375B_ABST
Patent Text Reader

Abstract

A voice operation method of a device, comprising: acquiring a video collected by a camera; acquiring voice information collected by a microphone; detecting a face image in the video; extracting a lip feature and a face feature of the face image; determining a time interval according to the lip feature; intercepting a corresponding audio segment in the voice information according to the time interval; acquiring voiceprint information according to the face feature; and performing voice recognition on the audio segment according to the voiceprint information to acquire voice information. The voice operation method of the device does not need to pre-determine a target user and pre-record voiceprint information of the target user, can autonomously extract voiceprint information of multiple users when multiple users use the device at the same time, and can separate voices one by one. Meanwhile, the voiceprint can be autonomously updated and registered. The voice operation method of the device can significantly improve the voice recognition effect of the device in a noisy or multi-user speaking scene.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech recognition, and in particular to a speech operation of a device, an apparatus and an electronic device. BACKGROUND

[0002] Speech recognition technology, also known as automatic speech recognition (ASR), aims to convert the lexical content in human speech into computer-readable input, such as key presses, binary codes or character sequences. Speech recognition is different from speaker recognition and speaker verification, which attempt to identify or verify the speaker who utters the speech rather than the lexical content and meaning contained in the speech. The greatest advantage of speech recognition technology is to make human-computer interaction more natural and make the interactive device easier to use.

[0003] In recent years, speech recognition technology in a single-person speaking scenario has been relatively mature, and existing speech recognition ASR technology can effectively recognize the speech signal of a single speaker as text. However, in a noisy environment, especially in a multi-user speaking scenario or when the use environment has background speech, the speech recognition effect of the existing speech recognition system will be greatly reduced.

[0004] The performance of a speech recognition system depends on the quality of the speech signal and whether it is a single speaker or multiple speakers; among them, the quality of the speech signal can be controlled and optimized by existing technologies or methods such as echo cancellation, noise suppression, automatic gain control adjustment, etc., while the problem of speech recognition in a multi-user speaking scenario needs to be solved urgently.

[0005] SUMMARY

[0006] To solve the above problems, embodiments of the present application provide a speech operation method of a device, an apparatus and an electronic device.

[0007] In a first aspect, the embodiments of the present application provide a speech operation method of a device, which comprises: acquiring a video collected by a camera; acquiring speech information collected by a microphone; detecting a face image in the video; extracting a lip feature and a face feature of the face image; determining a time interval according to the lip feature; extracting a corresponding audio segment in the speech information according to the time interval; acquiring voiceprint information according to the face feature; and performing speech recognition on the audio segment according to the voiceprint information to obtain speech information. In this way, the voiceprint information of multiple users can be registered and managed by the device autonomously, and the voiceprint information of each user can be extracted autonomously in the process of using the device for speech operation by multiple users, the speech can be separated one by one according to the user identity, and the speech information can be obtained.

[0008] In an embodiment, the acquiring voice information further includes: detecting whether a face image exists in the video; and recording the voice information when the face image exists. In this way, power consumption of the device can be reduced.

[0009] In an embodiment, the number of users is multiple, and the detecting the face image in the video and extracting the lip feature and the face feature of the face image include: detecting a first face image and a second face image in the video, extracting a first lip feature and a first face feature from the first face image, and extracting a second lip feature and a second face feature from the second face image. In this way, the identity and the speaking state of the user can be identified.

[0010] In an embodiment, the determining a time interval according to the lip feature and the intercepting a corresponding audio segment in the voice information according to the time interval include: determining a first time interval according to the first lip feature, intercepting a first audio segment in the voice information according to the first time interval, determining a second time interval according to the second lip feature, and intercepting a second audio segment in the voice information according to the second time interval. In this way, the time interval of each audio segment and the time interval of each face image in the video in a speaking state can be aligned.

[0011] In an embodiment, the acquiring voiceprint information of the user includes: searching for corresponding voiceprint information in the device according to the face feature, and acquiring the voiceprint information of the user. In this way, the voiceprint information of the user can be automatically extracted in the process of voice operation using the device.

[0012] In an embodiment, the acquiring voiceprint information of the user includes: searching for corresponding voiceprint information in the device according to the first face feature, acquiring first voiceprint information, and searching for corresponding voiceprint information in the device according to the second face feature, and acquiring second voiceprint information. In this way, the voiceprint information of each user can be automatically extracted in the process of voice operation using the device by multiple users.

[0013] In an embodiment, the acquiring the voiceprint information of the user comprises: searching for corresponding voiceprint information in the device according to the face feature, and if the search result is that the corresponding voiceprint information does not exist, determining a third time interval in which the user is in a single speech state according to the time interval; and acquiring a third audio segment from the voice information according to the third time interval, extracting a voiceprint feature of the third audio segment, and obtaining the voiceprint information of the user. In this way, the device can automatically register the voiceprint of a new user and acquire the voiceprint information of the new user in the process of using the device for voice operation, without increasing additional registration operations. In an embodiment, the acquiring the voiceprint information of the user comprises: searching for corresponding voiceprint information in the device according to the face feature, and if the search result is that the corresponding voiceprint information does not exist, determining a third time interval in which the user is in a single speech state according to a difference between the first time interval and the second time interval; and acquiring a third audio segment from the voice information according to the third time interval, extracting a voiceprint feature of the third audio segment, and obtaining the voiceprint information of the user. In this way, the device can automatically register the voiceprint of a new user and ensure that the clarity of the voiceprint registered by the new user can meet the use requirements, without increasing additional registration operations.

[0014] In an embodiment, the acquiring the voiceprint information of the user comprises: searching for corresponding voiceprint information in the device according to the face feature, and if the search result is that the corresponding voiceprint information does not exist, determining a third time interval in which the user is in a single speech state according to the time interval; and acquiring a third audio segment from the voice information according to the third time interval, extracting a voiceprint feature of the third audio segment, and obtaining the voiceprint information of the user. In this way, the device can automatically register the voiceprint of a new user and ensure that the clarity of the voiceprint registered by the new user can meet the use requirements, without increasing additional registration operations.

[0015] In an embodiment, the acquiring the voiceprint information of the user comprises: searching for corresponding voiceprint information in the device according to the face feature, and if the search result is that the corresponding voiceprint information does not exist, determining a third time interval in which the user is in a single speech state according to the time interval; and acquiring a third audio segment from the voice information according to the third time interval, extracting a voiceprint feature of the third audio segment, and obtaining the voiceprint information of the user. In this way, the device can automatically register the voiceprint of a new user and ensure that the clarity of the voiceprint registered by the new user can meet the use requirements, without increasing additional registration operations.

[0016] In an implementation, the voice recognition on the audio segment according to the voiceprint information comprises: performing voice separation on the first audio segment according to the first voiceprint feature to obtain first audio information; performing voice recognition on the first audio information to obtain first voice information; performing voice separation on the second audio segment according to the second voiceprint feature to obtain second audio information; and performing voice recognition on the second audio information to obtain second voice information. In this way, the device can autonomously extract the voiceprint information of each user, separate the voice information of each user one by one, recognize the intention or instruction of the voice information of each user, and meanwhile avoid the problem that the intention or instruction of voice recognition is confused due to the overlapping of multiple pieces of voice information.

[0017] In an implementation, the method further comprises: obtaining a registration time of the voiceprint feature, and if the registration time of the voiceprint feature exceeds a preset valid time, updating the voiceprint information to a third voiceprint feature extracted from an audio segment.

[0018] In an implementation, the updating of the voiceprint information to the third voiceprint feature comprises: comparing the third voiceprint feature and the voiceprint feature, and retaining one of the voiceprint features that meets a use requirement; wherein the voiceprint feature that meets the use requirement is a voiceprint feature that is relatively clean and clear. In this way, the old voiceprint feature of the user can be updated to maintain the effect of voice recognition.

[0019] In a second aspect, the embodiments of the present application provide a voice operation device of a device, and the beneficial effects can be referred to the description in the first aspect. The device comprises: a master control module configured to obtain a video collected by a camera and voice information collected by a microphone; a face detection module configured to detect a face image in the video; an image recognition module configured to extract a lip feature and a face feature of the face image; a lip feature recognition module configured to determine a time interval according to the lip feature, and cut a corresponding audio segment from the voice information according to the time interval; a voiceprint extraction module configured to obtain voiceprint information according to the face feature; and a voice recognition module configured to perform voice recognition on the audio segment according to the voiceprint information to obtain voice information.

[0020] In a third aspect, the embodiments of the present application provide an electronic device, and the beneficial effects can be referred to the description in the first aspect. The electronic device comprises a memory and a processor; the processor is configured to execute computer execution instructions stored in the memory, and the processor executes the computer execution instructions to perform the voice operation method of the device according to any one of the embodiments.

[0021] In a fourth aspect, the embodiments of the present application provide a storage medium, and the advantages can refer to the description in the first aspect. The storage medium comprises a readable storage medium and a computer program stored in the readable storage medium, and the computer program is used to implement the voice operation method of the device in any one of the above-mentioned embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the multiple embodiments disclosed in the specification, the drawings required to be used in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only the multiple embodiments disclosed in the specification, and other drawings can also be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0023] The drawings required to be used in the embodiment description or the prior art description will be briefly introduced as follows.

[0024] Figure 1 The application scenario diagram of the voice operation method of the device provided by the embodiments of the present application;

[0025] Figure 2 The system framework diagram of the voice operation method of the device provided by the present application;

[0026] Figure 3A The flowchart of the voice operation method of the device provided by the embodiment 1 of the present application;

[0027] Figure 3B The flowchart of the new user self voiceprint extraction;

[0028] Figure 4 The voice operation method module interaction diagram of the voice operation method of the device provided by the embodiment 1 of the present application;

[0029] Figure 5 The flowchart of the multi-user parallel voice separation of the voice operation method of the device provided by the embodiment 2 of the present application;

[0030] Figure 6 The module interaction diagram of the multi-user parallel voice separation of the voice operation method of the device provided by the embodiment 2 of the present application;

[0031] Figure 7 The flowchart of the device self-updating of the voice operation method of the device provided by the embodiment 3 of the present application;

[0032] Figure 8 The module interaction diagram of the device self-updating of the voice operation method of the device provided by the embodiment 3 of the present application;

[0033] Figure 9 An electronic device schematic diagram is provided for the embodiments of the present application. DETAILED DESCRIPTION

[0034] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application.

[0035] In one scheme, a speech separation method is used for speech recognition, the voiceprint of a target speaker is extracted from a reference speech of the target speaker for registration, the voiceprint is used as a clue to extract the speech of the target speaker in mixed speech, and the subsequent speech recognition is performed in the activated manner of a wake-up word. This scheme is suitable for a scene where personnel are relatively fixed and voiceprint registration is relatively convenient. However, in the multi-speaker scene, the speech of the target speaker is separated, the scheme is easily disturbed by the speech activity of surrounding people, the separated speech is difficult to recognize, and the subsequent speech recognition effect is not ideal.

[0036] Another deficiency is that the scheme needs to pre-register the voiceprint of the target speaker, extract the speech information of the target speaker corresponding to the registered voiceprint from the mixed speech containing multiple users speaking at the same time, and perform speech recognition on the extracted speech information as separated speech. In this process, the original speech information corresponding to the voiceprint for registration needs to meet certain requirements of time length and speech intelligibility. Special users such as the elderly and children who are not familiar with the voiceprint registration requirements and operation methods need to be guided to complete the registration, the user experience is poor, and the participation is low.

[0037] Another deficiency is that the target speaker must be pre-selected when the speech is separated by using the scheme. If the identities of the speakers in the mixed speech cannot be determined in advance in a multi-user scene where the speakers change, the target speaker whose speech needs to be separated cannot be determined in advance, and the registered voiceprint of the speaker whose speech needs to be separated cannot be obtained in advance, resulting in poor usability.

[0038] In addition, with the change of time, the speech of the target speaker participating in the registration will change due to the change of physiological characteristics, causing the voiceprint characteristics to drift, and causing the effect of using the pre-registered voiceprint of the target speaker for speech separation to be poor.

[0039] To solve the above problems, the embodiment of the present application provides a voice operation method of a device, which can independently register and manage voiceprint information of multiple users, independently extract voiceprint information of each user in the process of using the device for voice operation by multiple users, and separate the voice according to the identity of the user. The device using the method for voice recognition and operation does not need to determine the target user in advance, does not need to record the voiceprint information of the target user in advance, and does not need to select the voiceprint information of the target user to be separated, which can significantly improve the voice recognition effect in a noisy or multi-user speaking scene.

[0040] First, the voice operation method of the device provided by the present application is introduced in combination with the accompanying drawings. Figure 1 And 2 , the voice operation method of the device provided by the present application is introduced.

[0041] Figure 1 The application scenario diagram of the voice operation method of the device provided by the embodiment of the present application is shown in FIG. 1. Figure 1 As shown in FIG. 1, the voice operation method of the device provided by the embodiment of the present application is especially suitable for a scene with variable users such as a conference room or an exhibition booth, and is also suitable for a fixed scene with single and stable users.

[0042] The intelligent device is a device for video and voice interaction with a user, and the method for voice operation of the device provided in the embodiments of the present application needs to have some physical components that can be communicatively connected on the intelligent device in these scenarios. Exemplarily, the physical components can include a vehicle terminal used in a vehicle, a smart television used in a home scenario, a smart screen used in a conference room or an exhibition stand, and a smart phone capable of being connected with the above-mentioned intelligent devices, etc. The vehicle terminal device is also called an electronic control unit (ECU), a domain control unit (DCU), or a mobile data center (MDC), etc. The physical components include a processor 10, a camera 11, a sound receiving device 12, a display 13, a feature memory 14, a remote controller 15, and a loudspeaker 16, etc. The processor 10 can be a CPU, an MCU, an NPU, a GPU, a DSP, an ISP, etc. The processor 10 executes part or all of the application programs of the method for voice operation of the device provided in the embodiments of the present application. The camera 11 can be a camera communicatively connected inside the vehicle terminal or the intelligent device, or a camera communicatively connected outside the vehicle terminal or the intelligent device, and is used to shoot a video in a certain range interval in the current scenario. The video format output by the camera 11 includes MPEG, AVI, nAVI, ASF, MOV, 3GP, WMV, DivX, XviD, RM, RMVB, FLV / F4V, etc. The sound receiving device 12 records the original voice information of the user in a certain sound receiving range in the current scenario. The sound receiving device 12 includes an integrated microphone, a microphone or a microphone array connected with the intelligent device, or a microphone or a microphone array communicatively connected with the intelligent device through short-distance connection technology, etc. The format of the voice information output by the sound receiving device 12 includes WAV, MP3, AAC, OGG, WMA, OPUS, FLAC, APE, M4A, AMR, etc. The display 13 displays the text information recognized according to the user voice, the accepted user instruction, and the user identity recognition result, etc. The feature memory 14 stores the registration information of the user, including the face feature of the human face, the voiceprint feature, the user portrait, and the personalized information. The user portrait includes the identity, the gender, the age range, and the user group feature of the user; the personalized information includes the personal preferences and the historical use records according to the historical use information statistics, etc. The remote controller 15 starts the camera 11 or the sound receiving device 12 connected with the intelligent device through short-distance connection technology. The loudspeaker 16 performs voice interaction between the device and the user, including playing the voice recognition result, the suggestion, and the interactive operation prompt.

[0043] Figure 2 The system framework diagram of the method for voice operation of the device provided in the present application is shown in FIG. 1. As shown in FIG. 1, the system framework diagram of the method for voice operation of the device provided in the present application includes a processor 10, a camera 11, a sound receiving device 12, a display 13, a feature memory 14, a remote controller 15, and a loudspeaker 16, etc. Figure 2As shown, the logic function modules required by the processor 10 to perform the method include a master control module 20, a face detection module 21, an image recognition module 22, a state judgment module 23, a face feature matching module 231, a lip feature recognition module 232, a video codec 24, a voiceprint registration module 25, a speech separation module 26, a speech recognition module 27, an intent selection module 28, a speech codec 29, and a voiceprint extraction module 30. The state judgment module 23 includes the face feature matching module 231 and the lip feature recognition module 232.

[0044] The specific scheme includes: the master control module 20 of the device starts to shoot a video according to the camera 11, obtains a continuous image stream, performs face detection on the continuous image stream through the face detection module 21, or decodes the video that needs to be format-converted through the video codec 24, obtains a continuous image stream, and then performs face detection on each picture of the continuous image stream. It is detected whether there is a user using the device in the current environment, and after detecting the user, the master control module 20 starts the sound recording device 12 to record voice files in real time, and records the voice information of the user. The face image of each user in the video / continuous image stream is recognized through the image recognition module 22, and the identity of each user is recognized. Through the face feature matching module 231, it is identified that the current user is an existing user who has registered a voiceprint feature in the device or a new user who has not registered a voiceprint feature in the device; for multiple existing users using the device, the registration information of each existing user is determined respectively. Then, according to the registration information of each existing user, the voiceprint feature previously stored in the device is obtained. Through the lip feature recognition module 232, the time interval of the position change of the lip feature on the face image in the video is obtained. Through the voiceprint registration module 25, the audio information is obtained from the sound recording device 12, the continuous audio segment corresponding to the time interval of the position change of the lip feature is intercepted, the voiceprint feature of a single new user speaking alone is obtained, and the voiceprint feature of the new user is saved. Through the voiceprint extraction module 30, the voiceprint feature of each user saved in the device is extracted, the speech information that needs to be format-converted is decoded through the speech codec 29 to obtain continuous audio information, and the voiceprint feature of each user and the corresponding audio information are sent into the speech separation module 26 in parallel, and the voiceprint feature corresponding to each existing user is used for speech separation. The mixed speech of multiple users can be sent into one speech separation module and executed in parallel, or sent into multiple speech separation modules and executed simultaneously. Then, the audio information of each existing user separated is subjected to speech recognition respectively, and through the speech recognition module 27, the voice information of each user operating the current device is obtained according to the voiceprint information. Finally, the intent selection module 28 executes the intent and corresponding instruction of multiple users operating the current device according to the identified voice information according to the intent priority or matching correlation.

[0045] The above scheme can automatically extract the voiceprint information of each user when multiple users simultaneously issue voice commands to the device, separate the voice information of each user one by one, and identify the intention or instruction of each user's voice information, thereby avoiding the problem of incorrect voice recognition intention or instruction caused by the mutual overlap of multiple pieces of voice information, and further optimizing the effect of voice recognition and voice operation of the device in a multi-user simultaneous speaking scenario.

[0046] For each new user in the device without registered voiceprint information, the voiceprint registration module 25 selects the corresponding voice information when the user speaks alone, obtains the voiceprint feature and saves it, thereby realizing the self-registration of the voiceprint feature of the new user by the device, and the voiceprint feature can be directly used for separating the mixed voice of multiple users later.

[0047] Further, in view of the problem that the voiceprint feature of the registered user may drift over time, the identity of the user currently using the device can be identified during the use of the device by the user, and the condition of the voiceprint feature corresponding to the user is checked. If the registration time of the corresponding voiceprint feature is relatively early and needs to be updated, the new voiceprint feature of the user is obtained by using the voiceprint registration module 25 during the use of the device by the user, and the old voiceprint feature of the user is updated, so as to maintain the effect of voice recognition.

[0048] In the present application, the voiceprint is a voiceprint feature obtained by extracting the original voice information of a certain user through a traditional algorithm, and the voiceprint feature is used to separate the individual voice information from the mixed voice. The original voice information includes WAV, MP3, AAC, OGG, WMA, OPUS, FLAC, APE, M4A, AMR, etc. format files, and the traditional algorithm includes machine learning and / or deep learning algorithm. In the following description of the present application, the voiceprint, voiceprint information, voiceprint feature, voiceprint data, and voiceprint feature information are consistent in meaning, and all refer to the voiceprint feature obtained by extracting the original voice information of a certain speaker through a traditional algorithm.

[0049] The technical solutions in Embodiment 1 of the present application will be described below with reference to the accompanying drawings. Figure 3A 、 3B and the accompanying Figure 4 . Figure 3A The flowchart of the new user autonomous voiceprint extraction process of the voice operation method of the device provided in Embodiment 1 of the present application is shown in FIG. 1. Figure 3A As shown in FIG. 1, the process includes the following steps:

[0050] S301, obtaining the video of the user in the environment around the device collected by the camera 11, and obtaining the voice information of the user collected by the sound collecting device 12.

[0051] In a possible implementation, the processor 10 issues an instruction to start the camera 11, and detects whether there is a user in the current environment according to whether there is a face image in the video captured by the camera 11; in the case where there is a user, the processor 10 issues an instruction to start the sound recording device 12, and synchronously records voice information of the user.

[0052] Optionally, the camera 11 includes an embedded camera of the smart device, such as a built-in camera of a smart phone, a smart screen, or the like, or an external camera connected to the smart device.

[0053] Optionally, the video output by the camera 11 can be in one of the following formats, such as MPEG, AVI, nAVI, ASF, MOV, 3GP, WMV, DivX, XviD, RM, RMVB, and FLV / F4V, and the like; the processor 10 can decode the video that needs to be converted in format by a video codec 24 to obtain a continuous image stream.

[0054] Optionally, the voice information recorded by the sound recording device 12 can be in one of the following formats, such as WAV, MP3, AAC, OGG, WMA, OPUS, FLAC, APE, M4A, AMR, and the like; the processor 10 can decode the voice information that needs to be converted in format by a voice codec 29 to obtain a continuous audio stream.

[0055] In a possible implementation, the processor 10 can also first start the sound recording device 12 to record voice information in the current environment, detect whether there is a user who speaks in the current environment by voice recognition technology, and if there is, start the camera 11 to record a video of the user.

[0056] In a possible implementation, the processor 10 can also detect in a certain space by an infrared sensor, detect whether there is a user who uses the device by infrared recognition technology, if there is, start the camera 11 to record a video file of the user, and start the sound recording device 12 to record a voice file in real time, and record voice information of the user.

[0057] In a possible implementation, the processor 10 can also detect in a certain space by a proximity sensor, detect whether there is a user who uses the device at a position close to the current device, if there is, start the camera 11 to record a video file of the user, and start the sound recording device 12 to record a voice file in real time, and record voice information of the user.

[0058] Exemplarily, the sound recording device 12 can be a microphone, and the processor 10 gives an instruction to start the microphone to record voice information of a user in a certain sound recording range in the current scene.

[0059] Optionally, the sound collecting device 12 can be a microphone array, and the processor 10 gives an instruction to start the microphone array to collect sound of the user in a certain sound collecting range in the current environment based on sound source positioning technology of the microphone array. The directional sound collection can improve the signal-to-noise ratio (i.e., the ratio of signal to noise, the larger the ratio, the higher the useful signal-to-noise ratio compared to noise), the clarity and the recognition of the voice information.

[0060] Optionally, the sound source positioning technology of the traditional microphone array can be used to collect sound of the user, and exemplarily, the time of arrival (tdoa), the controllable wave velocity formation or the positioning technology based on the subspace can be used for sound source positioning, and the multi-sound source positioning technology based on the music algorithm can be used to collect sound of multiple users.

[0061] When there are multiple speakers near the device in the current environment, the voice information recorded in real time by the sound collecting device 12 is mixed voice information.

[0062] In S302, a face image in the video is detected, and a facial feature and a lip feature of the face image are extracted; a time interval is determined according to the lip feature; and a corresponding audio segment in the voice information is intercepted according to the time interval.

[0063] Optionally, the number of different faces can be counted according to the face images in the continuous image stream, so as to determine the number of users in front of the device, and the number of users can be one or more; and the number of users in a speaking state in front of the device can also be counted according to the number of face images with a relative position change of the lip feature in the continuous image stream, and the number of users in the speaking state can be one or more.

[0064] In a possible implementation, the processor 10 obtains a continuous image stream and a corresponding time interval from the video, identifies a face image of each user in the continuous image stream, and extracts a facial feature and a lip feature of each user from the face image of each user.

[0065] Exemplarily, when the number of users is multiple, a face image of one user in the video can be taken as a first face image, and a face image of any other user can be taken as a second face image, the first face image and the second face image in the video are detected, a first lip feature and a first facial feature are extracted from the first face image, and a second lip feature and a second facial feature are extracted from the second face image.

[0066] In a possible implementation, the face images corresponding to each user in the continuous image stream are detected in parallel, the time interval in which each user is in the speaking state is obtained according to the relative position change of the lip feature of each face image in the continuous image stream, and each audio segment corresponding to each time interval is cut from the recorded voice information, so that the time interval of each audio segment and the time interval in which each face image in the video is in the speaking state can be aligned.

[0067] For example, the first time interval can be determined according to the relative position change of the first lip feature on the first face image in the continuous image stream, the first audio segment is cut from the voice information according to the first time interval, the second time interval is determined according to the relative position change of the second lip feature on the second face image in the continuous image stream, and the second audio segment is cut from the voice information according to the second time interval.

[0068] For example, the processor 10 performs image recognition on each frame of image in the video, and the time interval in which user A is in the speaking state is 7:05-7:20 am, and the audio segment corresponding to the period of 7:05-7:20 am is cut from the voice information recorded in real time by the sound collecting device 12. The time interval in which user B is in the speaking state is 7:10-7:25 am, and the audio segment corresponding to the period of 7:10-7:25 am is cut from the voice information recorded in real time by the sound collecting device 12.

[0069] S303, obtaining the voiceprint information according to the face feature.

[0070] In a possible implementation, whether the face feature of each user exists in the feature storage 14 of the device and matches the voiceprint feature is found according to the face feature of each user, the finding result is that the face feature does not exist, the user corresponding to the face feature is a new user, the voiceprint feature of the new user needs to be registered, and then S304 is performed; the finding result is that the face feature exists, the user corresponding to the face feature is an existing user, and S305 is performed.

[0071] For example, the existing user in the device whose face feature and voiceprint feature have been registered has the face feature and the voiceprint feature stored in the feature storage 14 of the device, and the voiceprint feature can be found according to the face feature; the user in the device whose face feature and voiceprint feature are not saved is a new user, and the face feature and the voiceprint feature of the new user need to be registered in the device and stored in the feature storage 14 of the device.

[0072] S304, extracting the voiceprint feature according to the audio segment, obtaining the voiceprint information of the new user, and registering the voiceprint information.

[0073] The precondition for performing the voiceprint feature extraction is that the new user in the sound receiving range of the sound receiving device is in the individual speaking state to ensure that the quality of the voiceprint feature is clean and clear, and can meet the use requirements of the speech separation and speech recognition.

[0074] As a preferred implementation, the third time interval in which the user is in the individual speaking state can be determined according to the difference between the first time interval and the second time interval.

[0075] In a possible implementation, if all the users are in the camera detection range, such as Figure 3B As shown, the following steps are performed to obtain the voiceprint information of the new user.

[0076] S3041, the video is divided into several time intervals according to the time interval of the relative position change of the lip feature in the face image. For example, in the video segment from 7:05 to 7:20 in the morning, there are face images of user A and user B, the time interval of the relative position change of the lip feature of the face image of user A is from 7:05 to 7:10, and the time interval of the appearance of the relative position change of the lip feature of the face image of user B is from 7:10 to 7:15.

[0077] S3042, the time period in which the user is in the individual speaking state is determined according to the difference between the first time interval and the second time interval.

[0078] For example, if user A is a new user, the time interval of the relative position change of the lip feature of the face image of user A is the first time interval from 7:05 to 7:10, and the time interval of the appearance of the relative position change of the lip feature of the face image of user B is the second time interval from 7:08 to 7:15, it can be determined that the time period in which user A is in the individual speaking state is from 7:05 to 7:08, and the time period in which user B is in the individual speaking state is from 7:10 to 7:15.

[0079] If the first time interval in which the face image of user A appears is from 7:05 to 7:10, and the second time interval in which the face image of user B appears is from 7:03 to 7:15, the time period in which user A is in the individual speaking state cannot be determined, and S302 can be performed to determine the time period in which user A is in the individual speaking state according to the next frame of video.

[0080] S3043, the time interval in which the user is in the individual speaking state is determined according to the time period in which the user is in the individual speaking state in the video. The time interval is recorded as the third time interval.

[0081] Optionally, if there are multiple time periods in which the new user is in the individual speaking state in the video, the earliest time period can be selected as the third time interval according to the chronological order.

[0082] Optionally, if there are multiple new users in the video, determine the time period in which each new user is in the single speaking state.

[0083] S3044, according to the time period in which the user in the video is in the single speaking state, cut the third audio segment from the voice information.

[0084] In a possible implementation, in order to ensure that the quality of the voiceprint is clear and can meet the use requirements, it is necessary to determine whether the number of users in the speaking state in the third audio segment is 1, if the determination result is "yes", the audio segment is the audio segment of the new user in the single speaking state, and the voiceprint extraction condition is met, then S3045 is executed; if the audio segment is not in the single speaking state, S302 is executed to determine the time period in which the user is in the single speaking state according to the next frame of video.

[0085] In a possible implementation, if there is a user speaking outside the camera range, a top-down auditory attention (abbreviated as TDAA) algorithm can be used for speaker number inference. For the third audio segment recorded by the sound collecting device within the sound collecting range, the audio signals corresponding to the intensity from large to small are extracted from the audio segment from top to bottom according to the intensity of the audio features such as volume, signal-to-noise ratio, intelligibility or recognition degree, and the extraction is recursively performed until it cannot be extracted any more, and the number of users in the speaking state is inferred according to the number of extracted audio signals.

[0086] Exemplarily, the TDAA algorithm is used for the third audio segment from 7:05 to 7:08 to recursively extract the audio signal with the strongest intelligibility from the third audio segment from large to small, and the number of users in the speaking state is inferred. When there is 1 user speaking in the third audio segment from 7:05 to 7:08, the time period is the time period in which the new user is in the single speaking state, and the voiceprint extraction condition is met, then S3045 is executed; if the number of users in the speaking state inferred by the TDAA algorithm in this period is greater than 1, S302 is executed.

[0087] Optionally, if there are multiple new users in the video, the corresponding third audio segment is cut from the voice information according to the time period in which each new user is in the single speaking state.

[0088] Optionally, if there are multiple time periods in which a new user is in the single speaking state in the video, multiple third audio segments corresponding to each single speaking state period can be cut from the voice information.

[0089] S3045, extract the voiceprint feature of the third audio segment to obtain the voiceprint information of the new user.

[0090] Optionally, for multiple new users, the corresponding voiceprint features can be extracted from the third audio segments corresponding to each new user.

[0091] Optionally, if a new user has multiple third audio segments in a single speech state, the corresponding voiceprint features can be extracted from each third audio segment, the quality of each voiceprint feature is compared, and the voiceprint feature with the cleanest and clearest quality is selected as the voiceprint information of the new user, which can meet the use requirements of speech separation and speech recognition.

[0092] S3046, the correspondence between the face feature and the voiceprint feature of the new user is established, and the face feature and the voiceprint information are saved to the feature storage 14, and the registration of the face feature and the voiceprint feature of the new user in the device is completed.

[0093] For example, the face feature and the voiceprint information of the new user are set as the same ID or the associated ID.

[0094] Optionally, the face feature and the voiceprint information of the new user are stored in the user's personal registration information in the feature storage 14.

[0095] S305, the voiceprint feature registered in the device according to the face feature of the user is obtained.

[0096] In one possible implementation, when the number of users is multiple, the face feature of each user is extracted from the face image of each user, and the face feature and the corresponding voiceprint feature of each user are queried from the feature storage 14 of the device according to the face feature of each user.

[0097] For example, the corresponding voiceprint information can be found in the device according to the face feature of user A, and the first voiceprint information corresponding to user A is obtained; the corresponding voiceprint information is found in the device according to the face feature of user B, and the second voiceprint information corresponding to user B is obtained.

[0098] Optionally, the registration information of each user can be obtained from the feature storage 14 of the device according to the face feature of each user, the registration information including the face feature, the voiceprint feature, the user portrait and the personalized information of the user; the user portrait including the identity, the gender, the age group and the group characteristics of the user; the personalized information including the personal preferences and the historical use records of the user according to the historical use information.

[0099] S306, the speech recognition is performed on the first audio segment according to the voiceprint information, and the speech information is obtained.

[0100] In a possible implementation, when the number of users is multiple, the first audio segment can be subjected to speech separation according to the voiceprint feature of each user, audio information corresponding to each user is extracted, speech recognition is performed on the audio information of each user, and voice information issued by each user to the device is obtained.

[0101] Exemplarily, the first audio segment can be subjected to speech separation according to the first voiceprint feature corresponding to user A, and first audio information is obtained; speech recognition is performed on the first audio information, and first voice information issued by user A is obtained; the second audio segment is subjected to speech separation according to the second voiceprint feature corresponding to user B, and second audio information is obtained; and speech recognition is performed on the second audio information, and second voice information issued by user B is obtained.

[0102] Exemplarily, the voice information issued by the user to the device includes an intention or instruction of the user to operate the device.

[0103] In the method proposed in Embodiment 1 of the present application, the interaction process of the function module of the new user self-service voiceprint registration is as shown in Figure 4 , and will not be described herein again.

[0104] When multiple users use the device, each face image and facial feature in the current image is judged, when the voiceprint extraction condition is met, for a new user without voiceprint feature information in the device, voice information corresponding to a speech of the new user is subjected to voiceprint feature extraction and saved; the voice information can be directly used for speech separation in the future, and the problem that the speech recognition effect is greatly deteriorated in the current ASR scheme under a noisy environment, especially in a multi-user speech scenario, is solved.

[0105] For a smart screen / Car terminal and the like, multi-user simultaneous use is one of common use scenarios of the device, multiple users are simultaneously in the detection range of the camera 11 and the sound receiving range of the microphone 12 of the device, and common voice / noise interference exists in the environment, when the user performs voice interaction with the intelligent device, the device automatically identifies the user, and automatically completes voiceprint recording and extraction of a new user at a suitable time, completes registration of the voiceprint feature of the new user, reduces the use threshold of speech separation, and improves the ease of use of the speech separation method. When speech separation needs to be performed on the new user, since the device can automatically obtain the voiceprint information of the new user, the fluency of the interaction process is not interrupted, and the speech separation effect is improved.

[0106] The method for voice operation of the device provided in Embodiment 1 of the present application collects voice information of a user when the camera 11 identifies that the user has not registered in the current smart device and the user speaks in the current environment and only in the current environment, and subsequently directly uses the voiceprint feature of the user for voice separation in a multi-user scenario. This scheme does not need to specify a target user to pre-register voiceprint feature information, avoids introducing an additional step, and solves the problem that special groups such as the elderly and children have low participation due to complex operations in the process of using the device.

[0107] The method for voice operation of the device provided in Embodiment 1 of the present application acquires the facial features of the user through the camera 11, uses the facial features as the identity of the user, and can directly judge and select the corresponding voiceprint feature for the user speaking in the current scenario in the process of using the device, thereby solving the problem that the user identity cannot be judged and the voiceprint cannot be selected in a multi-user scenario.

[0108] For voice separation of multiple users in a multi-user scenario, the existing smart device directly performs voice recognition on the voice information, resulting in that when there are multiple speakers in the current environment, the text corresponding to the voice information of each speaker is mixed together and the voice information cannot be distinguished.

[0109] Embodiment 2 of the present application further extracts the audio information of each user in Embodiment 1, and sends the separated voice information of multiple users into the voice recognition module 27 respectively, and respectively acquires the intention or instruction of each user operating the current device.

[0110] The technical solutions in Embodiments of the present application will be described below with reference to the accompanying drawings. Figures 5-6 The technical solutions in Embodiments of the present application will be described below with reference to the accompanying drawings.

[0111] Figure 5 is a flowchart of multi-user parallel voice separation of the method for voice operation of the device provided in Embodiment 2 of the present application.

[0112] As shown in Figure 5 , the following steps are included:

[0113] S501, acquire the video of the user in the environment around the device collected by the camera 11, and acquire the voice information of the user collected by the sound collecting device 12. The specific implementation manner is as described in S301 of Embodiment 1, and will not be repeated here.

[0114] S502, detecting a face image in the video, extracting a face feature and a lip feature of the face image; determining a time interval of a relative position change of the lip feature on the face image in the video, and cutting a corresponding audio segment in the voice information according to the time interval. The specific implementation manner is as shown in S302 of Embodiment 1, and details are not repeated here.

[0115] S503, identifying the user according to the face feature, and obtaining a voiceprint information of the user.

[0116] In a possible implementation, the number of users is multiple, and the face feature of each user can be matched with the face feature stored in the current device to determine whether the user is an existing user, and the voiceprint feature registered by the existing user in the device can be obtained. The implementation can be achieved by performing the following steps.

[0117] S5031, determining whether the face feature in the face image corresponds to a user existing in the current device and having the face feature and the corresponding voiceprint information, if the determination result is "yes", the user is an existing user, and S5032 is performed; if the determination result is "no", the user is a new user, and S304 is performed.

[0118] Exemplarily, it can be determined whether the face feature matches the voiceprint feature in the feature storage 14 of the device according to the face image, if the result of the determination is "yes", the user corresponding to the face is an existing user, and the voiceprint feature of the user does not need to be registered, and S5032 is performed; if the result of the determination is "no", the user corresponding to the face is a new user, and the voiceprint feature of the user needs to be registered, and S304 is performed.

[0119] S5032, obtaining the voiceprint feature registered by each user in the device according to the face feature of each user.

[0120] In a possible implementation, the face feature of each user is extracted from the face image of each user, and the face feature and the corresponding voiceprint feature of each user are queried from the feature storage 14 of the device according to the face feature of each user.

[0121] Optionally, the registration information of the user can be obtained from the feature storage 14 of the device according to the face feature of each user, the registration information including the registered face feature, the voiceprint feature, the user portrait and the personalized information of the user; the user portrait includes the identity, the gender, the age range and the group feature of the user; the personalized information includes the personal preference and the historical use record of the user according to the historical use information.

[0122] S504, performing voice recognition on the voiceprint feature and the corresponding audio segment of each user in parallel, and obtaining the voice information issued by each user to the device.

[0123] In a possible implementation, the speech intention or instruction of each user can be obtained by performing S5041-S5043 as follows:

[0124] S5041, the voiceprint feature of each user and the corresponding audio segment are sent to a speech separation module 26 for speech separation.

[0125] Optionally, the processor 10 can send the voiceprint feature of each user and the corresponding audio segment to one speech separation module 26 for parallel processing, or send the voiceprint feature of each user and the corresponding audio segment to a corresponding number of speech separation modules 26 for parallel speech separation processing.

[0126] S5042, the speech separation module 26 extracts the audio information of each user from the corresponding audio segment according to the voiceprint feature of each user.

[0127] S5043, speech recognition is performed on the audio information of each user respectively to obtain the speech information of each user to the device.

[0128] In a possible implementation, the speech recognition module 27 performs speech recognition on the audio information of each user respectively to obtain the recognition result corresponding to each piece of audio information.

[0129] In a possible implementation, the speech recognition result of each user can be input into a semantic recognition model, and the semantic recognition model performs generalization search according to the user portrait of each user and the input speech recognition result to obtain the text corresponding to the speech recognition result of each user, and obtains the speech information of each user to the device according to the corresponding text.

[0130] The user portrait includes the identity, gender, age group, user group information, language, etc. of the user.

[0131] Due to the difference in geographical location and region, many regions have their own voice characteristics, and the intonation and speed of each person's speech also differ. The semantic recognition model can be personalized trained to adapt to the accent, dialect, habitual words, etc. with personal characteristics, so as to improve the accuracy of speech recognition.

[0132] For example, a user in Shaanxi Province issues an instruction "I want to buy ***" to a smart device, and the voice recognition result can be "I want to sell ***". The processor 10 of the device obtains the user portrait of the user as Zhang, male, 49 years old, and native of Shaanxi Province, a farmer according to the registration information of the user, and then inputs the voice recognition result "I want to sell ***" of the user into a trained semantic recognition model. The semantic recognition model performs a generalization search on the language of the male farmers under 50 years old who speak Chinese and live in Shaanxi Province according to the voice recognition result of the user combined with the user portrait, determines the search field as the range of Chinese-Shaanxi dialect, and can further recognize the voice information with local characteristics in the audio signal to obtain the voice information "I want to buy ***" issued by the user to the device.

[0133] In a possible implementation, the registration information of each user can be obtained from the device according to the facial features, the individualized information of each user can be obtained according to the registration information of each user, the voice recognition of each user can be individualized based on the individualized information, and the voice information issued by each user to the device can be obtained.

[0134] For example, the registration information includes a user name, a user id, and a user identification code; the individualized information includes the preferences, interested categories, and historical use records of each user; and the voice information issued by each user to the device includes the intention or instruction of each user to operate the device.

[0135] For example, the interested categories, preferences, and historical information of each user can be obtained from the apps used by the user according to the registration information of each user, the voice recognition of each user can be individualized based on the historical information, and the voice information issued by each user to the device can be obtained.

[0136] For example, the processor 10 of the device can obtain the historical information of the interested app from the app used by the user according to the registration information of the user, match the recognition result of the audio information based on the historical information, and quickly obtain the voice information "please open ***" issued by the user to the device.

[0137] Optionally, the voice information issued by the user to the device includes the intention or instruction of the user to operate the current device, and the intention or instruction of each user to operate the current device can be converted into corresponding text and displayed on the display screen of the smart device.

[0138] The individualized recognition of the voice information issued by each user to the device in the scheme can increase the accuracy and relevance of voice recognition.

[0139] S505, for each user's voice information, a strategy selection is performed, and the voice information of each user is executed according to the result of the strategy selection.

[0140] In a possible implementation, the training of the strategy selection of the intent decision model by using the voice information set with the priority identifier enables the voice information to be sorted according to the priority or matching relationship of the voice information and to be output in sequence.

[0141] In a possible implementation, the voice information of each user can be input into the pre-trained intent decision model, and the strategy selection of the intent decision model includes:

[0142] a. Preferentially executing the instruction of the high-priority intent. For example, in a home scenario, user 1 and user 2 use the smart screen device at the same time, at this time, user 1's mobile phone rings and needs to be muted, and the voice instruction of user 1 is "please mute"; and user 2 wants to change the channel, and the voice instruction of user 2 is "please change the channel"; the processor 10 inputs the mute and channel change instructions into the intent decision model, because the mute intent has higher real-time performance, the voice information sequence output by the intent decision model is mute-channel change, and the main control module 20 preferentially executes the mute instruction and then executes the channel change instruction.

[0143] b. Preferentially executing the instruction corresponding to the intent with high matching relationship. For example, in a home scenario, user 1 and user 2 use the smart screen device at the same time, the voice instruction recognition effect of user 1 is good, the voice recognition result is clear and matched, the execution intent is clear or the recognition confidence is high, and the instruction can be directly executed; the voice instruction of user 2 is not clear because of the speed, distance, accent and the like, the processor 10 inputs the voice instruction of user 1 and the voice instruction of user 2 into the intent decision model, the voice information sequence output by the intent decision model is the voice instruction of user 1-the voice instruction of user 2, and the main control module 20 preferentially executes the voice instruction of user 1 and then executes or does not execute the instruction of user 2, and preferentially presents the intent of user 1 on the display screen of the smart screen device and then presents or does not present the instruction of user 2.

[0144] c. Preferentially executing the instruction of a specific user. For example, in a vehicle terminal, the priority of the owner, driver or parent of the vehicle terminal is higher than that of the visitor, passenger and child, when multiple users control the vehicle terminal at the same time, the processor 10 inputs the voice instructions of the multiple users into the intent decision model, the intent decision model preferentially outputs the instruction of the user with high priority, and the main control module 20 preferentially executes the instruction of the user with high priority.

[0145] In the method proposed in Embodiment 2 of the present application, the functional module interaction process of multi-user parallel voice separation is as shown in Figure 6 , and will not be described here.

[0146] In the embodiment 2 of the present application, the intentions or instructions of multiple users can be executed according to the intention priority or matching relevance order, so as to solve the problem that multiple commands overlap each other when multiple users give voice commands at the same time, resulting in confusion of the recognized intentions. For example, user 1 and user 2 give voice commands at the same time in front of the device, the text of the voice command given by user 1 is A1A2A3A4, and the text of the voice command given by user 2 is B1B2B3B4. If user 1 and user 2 speak at the same time, the traditional voice recognition method may recognize A1B1A2B2A3B3A4B4, resulting in completely incorrect semantic recognition result. However, the scheme of the embodiment 2 of the present application can effectively separate the intentions and instructions of user 1 and user 2 speaking at the same time, solve the problem of poor semantic recognition effect when multiple users speak at the same time in a multi-user scenario, greatly improve the effect of voice intention executed by the device, and improve the experience of human-computer interaction voice operation.

[0147] The method of voice operation of the device provided in the embodiment 2 of the present application is particularly suitable for a fixed personnel scenario, such as an intelligent device in a car, a family, a conference room, and the like. The method also has good voice interaction experience for a multi-user scenario with unfixed personnel.

[0148] If the voice of the user changes due to physiological characteristic change over time, resulting in drift of the voiceprint characteristics, then in the process of using the device by the user, the original voiceprint data of the user in the device needs to be updated by extracting new voiceprint characteristics. The method of voice operation of the device provided in the embodiment 3 of the present application judges whether the voiceprint characteristics of the user saved in the feature storage 14 of the device meet the use requirement. If the voiceprint characteristics do not meet the use requirement, then the voiceprint information of the user is recorded when the user speaks alone, and the previous voiceprint information is updated, so that the voiceprint characteristics of the user saved in the device meet the use requirement, and the accuracy of voice separation and voice recognition is ensured. The following will be described in combination with the accompanying drawings. Figures 7-8 On the basis of the embodiment 1, the technical scheme in the embodiment 3 of the present application is described.

[0149] Figure 7 The flowchart of the method of voice operation of the device provided in the embodiment 3 of the present application is shown in FIG. 7. Figure 7 As shown in FIG. 7, the steps are as follows:

[0150] S701, the video of the user in the environment around the device is acquired by the camera 11, and the voice information of the user recorded by the sound recording device 12 is acquired. The specific implementation manner is as shown in S301 of the embodiment 1, which will not be described herein again.

[0151] S702: Detect face images in the video, extract facial and lip features from the face images; determine the time interval of the relative position change of the lip features on the face image in the video, and extract the corresponding audio segment from the speech information according to the time interval. The specific implementation method is the same as S302 in Embodiment 1, and will not be repeated here.

[0152] S703 identifies users based on facial features and obtains the user's voiceprint information and registration time.

[0153] In one possible implementation, the system determines whether the user is an existing user based on the facial features of each user. If the user is an existing user, the system obtains the existing user's voiceprint features and registration time, and executes S704; otherwise, it executes S303.

[0154] The specific implementation method for determining whether the corresponding user is an existing user is referred to S3031 of Embodiment 1, and will not be repeated here.

[0155] S704 If the user corresponding to a certain facial feature has a voiceprint feature registration time that exceeds the preset valid time, then execute S705; if the voiceprint feature registration time is within the preset valid time and does not need to be updated, then execute S305 or S504.

[0156] S705, extract voiceprint features based on audio segments to obtain the user's voiceprint information. For specific implementation methods, refer to S3041-3046 of Embodiment 1, which will not be repeated here.

[0157] S706: Compare the new voiceprint features and the old voiceprint features of the user to determine whether the new voiceprint features meet the quality requirements for voiceprint extraction. If they do, save the new voiceprint features and delete the old voiceprint features. Otherwise, retain the old voiceprint features.

[0158] In one possible implementation, the new voiceprint feature and the old voiceprint feature are compared, and the voiceprint feature that is cleaner and clearer and can meet the usage requirements is selected as the new voiceprint feature for the user, while the old voiceprint feature is deleted.

[0159] In the method proposed in Embodiment 3 of this application, the interactive process of the function module for autonomously updating multiple user voiceprint information is as follows: Figure 8 As shown, further details are omitted. When there are existing users registered in the device within the current environment, the voiceprint information quality and registration time of each user determine whether the current voiceprint features need to be updated. Therefore, when each existing user speaks individually, the user's voiceprint information is extracted and the previous voiceprint features are updated to ensure the timeliness of the user's voiceprint features in the system, thereby ensuring the effectiveness of speech separation.

[0160] The voiceprint characteristics of a person can slightly drift over time, and the voiceprint characteristics of the registered user in the embodiment 3 of the present application are used to automatically detect the validity period of the voiceprint features, and the voiceprint features are automatically updated according to the validity period when the registration conditions are met, thereby ensuring the timeliness of the user face features and the voiceprint features stored in the device, and ensuring the voice separation effect, and avoiding the problem of poor recognition effect caused by voiceprint features not meeting the use requirements. At the same time, the user does not need to be reminded to re-record the voiceprint or check regularly. Instead, the device automatically manages and updates the voiceprint feature information during the use of the device by the user.

[0161] The voice operation method of the device provided in the embodiments 1-3 of the present application can automatically complete the extraction, registration and update of the voiceprint features of the user. In the subsequent use of the device, when multiple users simultaneously give voice control instructions to the device, the corresponding relationship between the identity and the voiceprint information established previously can be used to separate the mixed voice to obtain audio information for each user in the scene. This is conducive to subsequent voice recognition, improves the recognition accuracy, and in addition, the intention decision can be made for the voice recognition results corresponding to the separated multiple pieces of voice information, and the voice instructions are executed according to the decision.

[0162] The voice operation method of the device provided in the embodiments 1-3 of the present application can better extract and process the voice instructions given by multiple users at the same time, improve the use experience of human-computer voice interaction and reduce the use threshold of the user, and finally the voiceprint information can be better updated and managed by using the method of the present application without the additional cooperation of the user.

[0163] The embodiments 1-3 in the present application mainly take the smart device in the home use scenario as the description object, and other suitable scenarios also include the in-vehicle environment. With the popularization of new energy vehicles, more and more vehicle terminals themselves are provided with cameras 11 and microphones 12. The method of the present application can help the driver to more conveniently realize the voice control of the vehicle terminal, and separate and exclude the interference, thereby improving the ease of use and safety of the driving operation. Similar use scenarios also include video conference rooms and exhibition booths where multiple users are present.

[0164] The embodiment of the present application provides a voice operation device of a device. A main control module 20 of the device acquires video collected by a camera and acquires voice information collected by a microphone. A face detection module 21 detects a face image in the video. An image recognition module 22 extracts a lip feature and a face feature of the face image. A lip feature recognition module 232 determines a time interval according to the lip feature, and extracts a corresponding audio segment in the voice information according to the time interval. A voiceprint extraction module 30 acquires voiceprint information according to the face feature. A voice recognition module 27 performs voice recognition on the audio segment according to the voiceprint information, and acquires voice information.

[0165] The embodiment of the present application provides an electronic device 900, as shown in the figure, comprising a memory 91, a processor 10; the processor 10 is used for executing computer execution instructions stored in the memory 91, and the processor 10 executes the computer execution instructions to execute the voice operation method of the device of any one of the above-mentioned embodiments. Figure 9 The embodiment of the present application provides a storage medium 92, comprising a readable storage medium and a computer program stored in the readable storage medium, and the computer program is used for realizing the voice operation method of the device of any one of the above-mentioned embodiments.

[0166] The embodiment of the present application provides a storage medium 92, comprising a readable storage medium and a computer program stored in the readable storage medium, and the computer program is used for realizing the voice operation method of the device of any one of the above-mentioned embodiments.

[0167] Those skilled in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in the present application can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software mode depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of the present application.

[0168] In addition, various aspects or features of the embodiments of the present application can be implemented as methods, apparatuses or products using standard programming and / or engineering techniques. The term "product" used in the present application covers computer programs that can be accessed from any computer readable device, carrier or medium. For example, the computer readable medium can include, but is not limited to: magnetic storage devices (such as hard disks, floppy disks or magnetic tapes, etc.), optical discs (such as compact discs (CD), digital versatile discs (DVD), etc.), smart cards and flash memory devices (such as erasable programmable read-only memory (EPROM), cards, sticks or key drives, etc.). In addition, various storage media described in the present application can represent one or more devices and / or other machine readable media for storing information. The term "machine readable medium" can include, but is not limited to, wireless channels and various other media capable of storing, containing and / or carrying instructions and / or data.

[0169] It should be understood that in various embodiments of the embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of the processes should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0170] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system, device and unit described above can refer to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0171] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.

[0172] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0173] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present application essentially or the parts that contribute to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or an access network device, etc.) execute all or part of the steps of the various embodiments of the method of the present application. The foregoing storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, and various program code storage media.

[0174] The above is only a specific implementation of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical range disclosed in the embodiments of the present application, which should be covered within the protection scope of the embodiments of the present application.

Claims

1. A voice operation method of a device, characterized by, The method comprises: acquiring a video collected by a camera; acquiring a voice collected by a microphone; in the case of multiple users using the device simultaneously, the voice is a mixed voice; detecting a face image in the video, extracting a lip feature and a face feature of the face image; including: detecting a first face image in the video, extracting a first lip feature and a first face feature from the first face image; determining a time interval according to the lip feature; including: determining a first time interval according to the first lip feature; extracting a corresponding audio segment from the voice information according to the time interval; including: extracting a first audio segment from the mixed voice according to the first time interval; acquiring voiceprint information according to the face feature, including: the multiple users include at least one new user; extracting a corresponding third audio segment from the mixed voice according to a period in which each new user is in a single speaking state; extracting a corresponding voiceprint feature from the third audio segment corresponding to each new user; the corresponding voiceprint feature includes first voiceprint information; establishing a correspondence between the face feature and the voiceprint information of each new user, and saving the face feature and the voiceprint feature; the correspondence between the face feature and the voiceprint information of each new user includes the correspondence between the first face feature and the first voiceprint information; acquiring voiceprint information according to the face feature, further comprising: searching for corresponding voiceprint information in the device according to the first face feature, and acquiring the first voiceprint information; performing voice recognition on the audio segment according to the voiceprint information to acquire voice information, including: performing voice separation on the first audio segment according to the first voiceprint information to obtain first audio information; performing voice recognition on the first audio information to acquire first voice information.

2. The voice operation method of the device according to claim 1, characterized by, The acquisition of voice information further comprises: detecting whether a face image exists in the video; recording the voice information in the case of the face image.

3. The voice operation method of the device according to claim 1 or 2, characterized in that, The number of users is multiple, and the detection of the face image in the video and the extraction of the lip feature and the face feature of the face image comprise: detecting a second face image in the video, extracting a second lip feature and a second face feature from the second face image.

4. The voice operation method of the device according to claim 3, characterized by, The determination of the time interval according to the lip feature; extracting a corresponding audio segment from the voice information according to the time interval, including: determining a second time interval according to the second lip feature; extracting a second audio segment from the voice information according to the second time interval.

5. The voice operation method of the device according to claim 4, characterized by, The acquisition of the voiceprint information of the user includes: searching for corresponding voiceprint information in the device according to the second face feature, and acquiring second voiceprint information.

6. The voice operation method of the device according to claim 1, characterized by, The extraction of a corresponding third audio segment from the mixed voice according to a period in which each new user is in a single speaking state includes: searching for corresponding voiceprint information in the device according to the face feature, and determining that the user is a new user if the search result is non-existent; determining a third time interval in which the user is in a single speaking state according to the time interval; According to the third time interval, a third audio segment is cut from the mixed voice.

7. The voice-operated method of operating an appliance according to any one of claims 4-6, characterized in that, The third audio segment is cut from the mixed voice according to the third time interval, and the method further comprises: According to the face feature, corresponding voiceprint information is searched in the device, and if the search result is that the voiceprint information does not exist, then According to the difference between the first time interval and the second time interval, a third time interval in which the user is determined to be in the individual speech state is determined. According to the third time interval, a third audio segment is cut from the mixed voice.

8. The voice operation method of the device according to claim 6, wherein, The third audio segment is cut from the mixed voice according to the third time interval, and the method further comprises: An audio signal with the maximum intensity of audio features is extracted from the third audio segment from top to bottom in turn by using a TDAA algorithm, and the number of users in the speech state is inferred according to the number of extracted audio signals. When the number of users in the speech state is 1, the third audio segment is extracted.

9. The voice operation method of the device according to claim 5, wherein, According to the voiceprint information, voice recognition is performed on the audio segment to obtain voice information, and the method further comprises: According to the second voiceprint information, voice separation is performed on the second audio segment to obtain second audio information. The second audio information is subjected to voice recognition to obtain second voice information.

10. The voice operation method of the device according to claim 1, wherein, The method further comprises: The registration time of the voiceprint feature is obtained, and if the registration time of the voiceprint feature exceeds a pre-set valid time, then According to the audio segment, a third voiceprint feature is extracted, and the voiceprint information is updated to the third voiceprint feature.

11. The voice operation method of the device according to claim 10, wherein, The voiceprint information is updated to the third voiceprint feature, and the method further comprises: The third voiceprint feature and the voiceprint feature are compared, and one voiceprint feature that meets the use requirement is retained from the two voiceprint features. The voiceprint feature that meets the use requirement is a clean and clear voiceprint feature.

12. A voice operated device of an apparatus, characterized by The device comprises: A master control module is configured to acquire a video collected by a camera and acquire voice collected by a microphone; in a case where multiple users use the device at the same time, the voice is mixed voice; A face detection module is configured to detect a face image in the video; the face detection module comprises a first face detection module configured to detect a first face image in the video; An image recognition module is configured to extract a lip feature and a face feature of the face image; the image recognition module comprises a first image recognition module configured to extract a first lip feature and a first face feature from the first face image; A lip feature recognition module is configured to determine a time interval according to the lip feature and cut a corresponding audio segment from the voice information according to the time interval; the lip feature recognition module comprises a first lip feature recognition module configured to determine a first time interval according to the first lip feature and cut a first audio segment from the mixed voice according to the first time interval; The voiceprint extraction module is configured to acquire voiceprint information according to the face feature; the multi-user includes at least one new user, a third audio segment corresponding to each new user is intercepted from the mixed voice according to a period in which each new user is in a single speech state; a corresponding voiceprint feature is extracted from the third audio segment corresponding to each new user; the corresponding voiceprint feature includes first voiceprint information; a corresponding relationship between the face feature and the voiceprint information of each new user is established, and the face feature and the voiceprint feature are stored; the corresponding relationship between the face feature and the voiceprint information of each new user includes a corresponding relationship between the first face feature and the first voiceprint information; The voiceprint extraction module is further configured to search for corresponding voiceprint information in the device according to the first face feature, and acquire first voiceprint information; The voice recognition module is configured to perform voice recognition on the audio segment according to the voiceprint information, and acquire voice information; the voice recognition module is configured to perform voice separation on the first audio segment according to the first voiceprint information, and obtain first audio information; and the voice recognition module is configured to perform voice recognition on the first audio information, and acquire first voice information.

13. An electronic device, comprising: The device includes a memory and a processor; the processor is configured to execute computer execution instructions stored in the memory; and the processor is configured to execute the computer execution instructions to execute the voice operation method of the device according to any one of claims 1 to 11.

14. A storage medium, characterized by The device includes a readable storage medium and a computer program stored in the readable storage medium; and the computer program is configured to implement the voice operation method of the device according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Conference summary generation method and device, computer device and storage medium

    CN110298252A

  • Sound source localization method and device with high accuracy, voice recognition method and system, storage device and terminal

    CN110767226A

  • Speech processing method, apparatus, system and device and medium

    CN110875053A

  • A voice detection apparatus, a method thereof and a vehicle including the same

    KR1020180119929A

Cited By

  • Speech operation method for device, apparatus, and electronic device

    WO2022179253A1