User registration method and related equipment

By collecting the user's image set and voice, using the interactive willingness value model and consistency judgment, the intelligent device dynamically decides the registration timing, solving the problem of low accuracy of contactless registration and achieving contactless registration with high accuracy and simplified operation.

CN115609596BActive Publication Date: 2025-09-05HUAWEI TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202110734850.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-30
Publication Date
2025-09-05
Estimated Expiration
2041-06-30

AI Technical Summary

Technical Problem

The accuracy of contactless registration of existing smart devices is low and misregistration is prone to occur, which is especially unsuitable for operation by children or the elderly.

Method used

Smart devices collect user images and voices, dynamically decide on the timing of registration, and use interaction willingness value models and consistency judgments to improve the accuracy of contactless registration.

Benefits of technology

It improves the accuracy of contactless registration, avoids the mistaken registration of incorrect biometric information or non-target users, simplifies the registration process, and facilitates operation for children and the elderly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115609596B_ABST
    Figure CN115609596B_ABST
Patent Text Reader

Abstract

This application provides a user registration method and related devices for use with smart devices. The method comprises: collecting a user's image set and voice; obtaining the user's interaction willingness value based on the image set, where the interaction willingness value describes the user's willingness to interact with the smart device; and registering the user's face and / or voiceprint if the interaction willingness value is greater than or equal to an interaction willingness threshold. This application dynamically determines the registration timing during the user's interaction with the smart device, allowing for seamless user registration and improving the accuracy of seamless registration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information security technology, and in particular to a user registration method and related equipment. Background Art

[0002] At present, most smart devices (such as smart robots) perform biometric registration such as face and voiceprint, which is based on a preset established process. That is, the registration is performed based on a preset established process (for example, the user needs to enter basic user information such as user name and mobile phone number according to the prompts, and then enter biometric information such as face or voiceprint), and the user can perceive the registration process. The accuracy rate of sensory registration is high and it is not easy to register by mistake. However, the process of sensory registration is relatively cumbersome and not conducive to operation by children or the elderly. Some smart devices are already capable of senseless registration, which means that the user registers while the user is using the smart device, and the user will not be aware of the registration process. However, the accuracy rate of senseless registration is not high at present, and it is easy to register by mistake, that is, register the wrong biometric information or non-target users. Summary of the Invention

[0003] The embodiment of the present application discloses a user registration method and related devices, which can dynamically decide the registration timing during the interaction between the user and the smart device, perform contactless registration for the user, and improve the accuracy of contactless registration.

[0004] In a first aspect, the present application discloses a user registration method, which is applied to a smart device. The method includes: collecting a user's image set and voice; obtaining the user's interaction willingness value based on the image set, wherein the interaction willingness value is used to describe the strength of the user's willingness to interact with the smart device; if the interaction willingness value is greater than or equal to an interaction willingness threshold, registering the user's face and / or voiceprint.

[0005] According to the user registration method provided in the embodiment of the present application, the smart device registers the user's face and / or voiceprint during a conversation with a stranger (such as chatting or playing games), dynamically decides the registration timing, and improves the accuracy of contactless registration.

[0006] In some optional implementations, obtaining the user's interaction willingness value based on the image set includes: acquiring the user's behavior information based on the image set; and obtaining the interaction willingness value based on the user's behavior information.

[0007] In some optional embodiments, obtaining the user's behavior information based on the image set includes: obtaining the user's facial angle and the distance between the user and the smart device in each image of the image set; inputting the facial angle and the distance into an interaction willingness value model to obtain the user's initial interaction willingness value; identifying the user's actions based on the image set; and obtaining the user's interaction willingness value based on the user's initial interaction willingness value and actions.

[0008] In some optional embodiments, obtaining the user's interaction willingness value based on the user's initial interaction willingness value and actions includes: if the user's actions include preset actions, adding the preset interaction willingness value on the basis of the initial interaction willingness value to obtain the user's interaction willingness value; if the user's actions do not include preset actions, using the initial interaction willingness value as the user's interaction willingness value.

[0009] In some optional embodiments, the actions include facial movements, head movements, and body movements.

[0010] In some optional embodiments, before registering the user's face and / or voiceprint, the method further includes: determining whether the image set and the voice are consistent; when the interaction willingness value is greater than or equal to the interaction willingness threshold, and the image set and the voice are consistent, performing the registration of the user's face and / or voiceprint.

[0011] By judging the consistency between the collected image set and the voice, it is possible to effectively avoid misrecognition or misjudgment caused by the voice of users who are not in the image set. For example, in a scenario where a user is in front of a smart device but does not speak, and a large-screen device next to it is playing audio, it is possible to effectively avoid the smart device from misidentifying the audio played on the large-screen device as the voice of the user.

[0012] In some optional implementations, determining the consistency of the image set and the voice includes: determining whether the image set and the voice correspond to the same user; and / or determining whether the image set and the voice correspond to the same language content.

[0013] In some optional embodiments, determining whether the image set and the voice correspond to the same user includes: performing lip movement recognition based on the image set to obtain a first speaker position; performing sound source localization based on the voice to obtain a second speaker position; and determining whether the first speaker position and the second speaker position are consistent.

[0014] In some optional embodiments, determining whether the image set and the speech correspond to the same language content includes: performing lip reading recognition on the image set to obtain a first language content; performing speech recognition on the speech to obtain a second language content; and determining whether the first language content and the second language content are consistent.

[0015] In some optional embodiments, obtaining the user's interaction willingness value based on the image set includes: if the image set includes multiple users, obtaining the interaction willingness value of each of the multiple users; the method also includes: if the interaction willingness value of more than one user among the multiple users is greater than or equal to the interaction willingness threshold, selecting the user with the largest interaction willingness value among the multiple users as the target user.

[0016] In some optional embodiments, registering the user's voiceprint includes: collecting at least three segments of speech from the user, each segment of speech being greater than or equal to a preset duration; pre-registering the voiceprint for each segment of speech; based on each segment of speech, performing voiceprint verification on other segments of speech in the at least three segments of speech; if the voiceprint verification of other segments of speech passes based on each segment of speech, the user's voiceprint registration is successful.

[0017] In some optional embodiments, before obtaining the user's interaction willingness value based on the image set, the method further includes: obtaining the field of view angle corresponding to the user in the image set; if the field of view angle corresponding to the user is less than or equal to the field of view angle threshold, executing the method of obtaining the user's interaction willingness value based on the image set.

[0018] In some optional implementations, after registering the user's face and / or voiceprint, the method further includes: collecting interaction information between the user and the smart device; and actively interacting with the user based on the interaction information.

[0019] A second aspect of the present application discloses a computer-readable storage medium comprising computer instructions. When the computer instructions are executed on a smart device, the smart device executes the user registration method as described in the first aspect.

[0020] The third aspect of the present application discloses a smart device, which includes a processor and a memory, wherein the memory is used to store instructions, and the processor is used to call the instructions in the memory, so that the smart device executes the user registration method described in the first aspect.

[0021] The fourth aspect of the present application discloses a chip system, which is applied to a smart device; the chip system includes an interface circuit and a processor; the interface circuit and the processor are interconnected through lines; the interface circuit is used to receive signals from the memory of the smart device and send signals to the processor, the signals including computer instructions stored in the memory; when the processor executes the computer instructions, the chip system executes the user registration method described in the first aspect.

[0022] In a fifth aspect, the present application discloses a computer program product. When the computer program product is run on a computer, the computer is caused to execute the user registration method as described in the first aspect.

[0023] A sixth aspect of the present application discloses a device capable of implementing the intelligent device behavior described in the method of the first aspect. This functionality can be implemented through hardware or through hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the aforementioned functionality.

[0024] It should be understood that the computer-readable storage medium described in the second aspect, the smart device described in the third aspect, the chip system described in the fourth aspect, the computer program product described in the fifth aspect, and the apparatus described in the sixth aspect all correspond to the method of the first aspect. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 This is a schematic diagram of an application scenario of the user registration method provided in an embodiment of the present application.

[0026] Figure 2 This is a flowchart of the user registration method provided in an embodiment of the present application.

[0027] Figure 3 This is a schematic diagram of an application scenario in which multiple users exist around a smart device.

[0028] Figure 4 This is a schematic diagram of voiceprint registration.

[0029] Figure 5 This is a flowchart of a user registration method provided by another embodiment of the present application.

[0030] Figure 6 This is a flowchart of a user registration method provided by another embodiment of the present application.

[0031] Figure 7 The smart device obtains the user's interaction willingness value based on the collected image set ( Figure 2 202) of the detailed flow chart.

[0032] Figure 8 It is a schematic diagram of the structure of the smart device provided in the embodiment of the present application. DETAILED DESCRIPTION

[0033] To facilitate understanding, some illustrations of concepts related to the embodiments of the present application are given for reference.

[0034] It should be noted that in this application, "at least one" means one or more, and "more than one" means two or more than two. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and drawings of this application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0035] In order to better understand the user registration method and related devices disclosed in the embodiments of the present application, the application scenarios of the user registration method of the present application are first described below.

[0036] Figure 1 This is a schematic diagram of an application scenario of the user registration method provided in an embodiment of the present application.

[0037] like Figure 1 As shown, the user registration method provided in the embodiment of the present application can be applied to intelligent robots. Currently known intelligent devices (such as intelligent robots) perform registration of biometric information such as face and voiceprint, mostly based on a preset established process. The user can perceive the registration process, which is called sensory registration. For example, the user needs to enter basic user information such as username and mobile phone number according to the prompts, and then enter biometric information such as face or voiceprint. The registration process is relatively cumbersome and not conducive to operation by children or the elderly. Currently, there are also some intelligent devices that can achieve sensorless registration. User registration is performed while the user is using the intelligent device, and the user will not perceive the registration process. However, the accuracy rate of sensorless registration is currently not high, and it is easy to register incorrect biometric information or non-target users (i.e., misregistration). According to the user registration method provided in the embodiment of the present application, during the process of a conversation between a user and an intelligent robot (such as chatting with the intelligent robot or playing a dialogue game), the intelligent robot dynamically decides the registration time and registers the user's face and voiceprint, avoiding the registration of incorrect biometric information or non-target users, thereby improving the accuracy rate of sensorless registration. During the registration process, users do not need to enter or speak basic user information such as username and mobile phone number as prompted, nor do they need to enter biometric information such as face and voiceprint as prompted.

[0038] The intelligent robot includes an image acquisition device (e.g., a camera) and a voice acquisition device (e.g., a microphone array). The image acquisition device is used to acquire images, and the voice acquisition device is used to acquire voice. The intelligent robot may have functions such as visual tracking, face recognition, sound source localization, lip movement recognition, lip reading recognition, voice recognition, and interaction intention (engagement) recognition. The intelligent robot implements the user registration method provided in the embodiments of the present application based on these functions.

[0039] The intelligent robot can collect images and voices in real time, and determine whether the registration conditions are met based on the collected images and voices. When the registration conditions are met, the user will be registered.

[0040] In the embodiment of the present application, the intelligent robot can identify the user's willingness to interact with the intelligent robot based on images collected over a period of time, determine the consistency of the images and voices collected during this period, and register the user when the user's willingness to interact and the consistency of the images and voices meet the requirements. Identifying the user's willingness to interact is to determine the strength of the user's willingness to interact with the intelligent robot. The strength of the interaction willingness can be represented by an interaction willingness value, and the user's interaction willingness value can be obtained through an interaction willingness value model. Judging the consistency of image and voice is to judge whether the collected image and voice are consistent, including judging whether the collected image and voice correspond to the same user, judging whether the collected image and voice correspond to the same language content, and the consistency of image and voice can be judged through sound source localization, lip movement recognition, lip language recognition, voice recognition and other technologies. The specific process of user registration by the intelligent robot will be described in detail in the following sections. Figure 2 Described in .

[0041] During the conversation between the user and the intelligent robot, the user can actively initiate interaction with the intelligent robot. For example, after the user says the intelligent robot's wake-up word (such as "XX classmate"), the intelligent robot can be woken up to start a conversation with the intelligent robot. After waking up the intelligent robot, the user can express his or her needs. Figure 1 As shown in the figure, the user says "What's the weather like today?" The intelligent robot can analyze the user's voice to obtain the user's semantic meaning and then provide feedback to the user, such as "Today's weather in Beijing is cloudy, 18℃-25℃."

[0042] Alternatively, during a conversation between a user and an intelligent robot, the intelligent robot can proactively interact with the user. For example, the intelligent robot can include a distance sensor or a proximity light sensor to detect the presence of a user nearby. When the intelligent robot detects the presence of a user nearby, it can proactively interact with the user. For example, the intelligent robot can proactively play a voice message saying, "Hi, host, let me play a song for you." If the user responds, "OK," the intelligent robot will begin playing the song.

[0043] During the conversation between the user and the intelligent robot, the user can stay or walk in front of the intelligent robot, and the intelligent robot can follow the user through visual following technology to turn towards the user.

[0044] Figure 1 The user registration method provided in the embodiment of the present application is applied to the scenario of an intelligent robot. In other embodiments of the present application, the user registration method provided in the embodiment of the present application can also be applied to other intelligent devices that interact with users, such as smart speakers and terminal devices.

[0045] The terminal device in the embodiments of the present application may refer to a user device, an access terminal, a user unit, a user station, a remote terminal, a mobile device, a user terminal, a wireless communication device, a user agent, or a user apparatus. The terminal device may be a mobile phone, a tablet computer, a computer with wireless transceiver function, a Session Initiation Protocol (SIP) phone, a personal digital assistant (PDA), a handheld device with wireless communication function, a computer or other processing device, an in-vehicle device, a wearable device, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal in a smart home, a terminal device in a future 5G network, or a terminal device in a future evolved public land mobile network (PLMN), etc., and the embodiments of the present application do not limit this.

[0046] In other embodiments of the present application, a smart device (e.g., a smart robot) may also be connected to a server. The server may provide content services to the smart device, such as providing the smart device with songs, stories, movies, etc. that the user has requested. The smart device may also send collected images and voice to the server, which may process the collected images and voice. For example, the smart device may send collected images to the server for face recognition and send collected voice to the server for voice recognition.

[0047] Figure 2 The user registration method provided in the embodiment of the present application is applied to smart devices, such as Figure 1 Intelligent robots in the.

[0048] like Figure 2 As shown, the user registration method provided in the embodiment of the present application includes the following steps:

[0049] 201, the smart device collects the user's image set and voice.

[0050] The image set and voice collected by the smart device correspond to the same time period.

[0051] As mentioned above, the intelligent device (such as an intelligent robot) includes an image acquisition device and a voice acquisition device. The intelligent device acquires images through the image acquisition device and acquires voice through the voice acquisition device.

[0052] The smart device can capture images in real time, and if a user is identified in the captured image, a set of images of the user is captured. Exemplarily, the smart device may include a motor, which can drive the smart device to rotate so that the smart device faces the user and tracks the user, and / or the smart device may include a motion device (such as wheels) that allows the smart device to follow the user's movement and track the user. After identifying the user from the image, the smart device can track the user to obtain a set of images of the user. The smart device can extract user features (such as color features, such as color histograms) from the captured image, and track the user based on the user features using a Kalman filter algorithm, a particle filter algorithm, or the like. For tracking users using a Kalman filter algorithm or a particle filter algorithm, reference can be made to related technologies, which will not be described here.

[0053] In one embodiment of the present application, a smart device can use a user recognition model to identify whether a user is present in an image. The user recognition model represents the correspondence between user characteristics and images, and can be a neural network model. Using this user recognition model, the smart device can determine whether a user is present in a captured image. If a user is present in the image, the device captures a collection of images and voice of the user.

[0054] A user's image collection includes multiple images of the user. These images may include the user's face, full-body, or half-body images. Therefore, the smart device can capture the user's head movements, body movements, and other information based on the collected image collection.

[0055] In one embodiment of the present application, in order to save power of the smart device, the smart device may start collecting the user's image set and voice after detecting the presence of a user nearby. The smart device may detect whether there is a user nearby using a distance sensor or a proximity light sensor.

[0056] In an embodiment of the present application, a user's image set may include a preset number of images. Specifically, after detecting the presence of a user around, the smart device may capture an image of the user every preset time period to obtain a preset number of images. For example, the preset time period is 60ms, and the preset number is 5. After detecting the presence of a user around, the smart device captures the user's image at 60ms, 120ms, 180ms, 240ms, and 300ms (here, the time when the user is detected around the smart device is the starting time 0ms), and collects 5 images to obtain the user's image set. It should be understood that the images in the image set in the embodiment of the present application include images of the user's face and limbs. Alternatively, the smart device may also start recording a video after detecting the presence of a user around, and then obtain the user's image set by intercepting multiple frames of images from the recorded video.

[0057] 202. The smart device obtains the user's interaction willingness value based on the collected image set.

[0058] The interaction willingness value indicates the user's willingness to interact with the smart device. A higher interaction willingness value indicates a stronger user's willingness to interact with the smart device; a lower interaction willingness value indicates a weaker user's willingness to interact with the smart device. For example, the interaction willingness value can range from 0 to 1, with values ​​closer to 0 indicating a weaker user's willingness to interact and values ​​closer to 1 indicating a stronger user's willingness to interact.

[0059] In an embodiment of the present application, the smart device can obtain the user's behavior information based on the collected image set, and obtain the user's interaction willingness value based on the user's behavior information. The image set includes multiple images, and the user's behavior parameters corresponding to each image constitute the user's behavior information. The user's behavior parameters may include the user's facial angle, the distance between the user and the smart device, the user's movements, etc. Among them, the user's movements may include the user's head movements and body movements. For example, the closer the user's facial angle is to facing the smart device, the closer the distance between the user and the smart device is, and the more the user's movements tend to interact with the smart device, the stronger the user's interaction willingness value.

[0060] The specific process of obtaining the user's interaction willingness value based on the collected image set will be described in Figure 5 Described in .

[0061] In an embodiment of the present application, when there are multiple users around the smart device, the image captured by the smart device may include multiple users, and the smart device may obtain the interaction willingness value of each user.

[0062] Figure 3 This is a schematic diagram of an application scenario where there are multiple users around a smart device. Figure 3As shown, there are two users around the smart device, one is a child and the other is an adult. The image collected by the smart device includes these two users. The smart device obtains the interaction willingness values ​​of the child and the adult according to the above method.

[0063] 203 : The smart device determines whether the user's interaction willingness value is greater than or equal to an interaction willingness threshold (eg, 0.7).

[0064] If the user's interaction willingness value is greater than or equal to the interaction willingness threshold, it indicates that the user has a strong willingness to interact with the smart device, and subsequent processing is performed at this time.

[0065] Otherwise, if the user's interaction willingness value is less than the interaction willingness threshold, it indicates that the user's willingness to interact with the smart device is weak, and 201 is returned.

[0066] In one embodiment of the present application, if there are multiple users in the image and the interaction willingness values ​​of more than one user are greater than or equal to the interaction willingness threshold, the smart device determines the user with the largest interaction willingness value as the target user and performs subsequent processing on the target user.

[0067] Continue with Figure 3 For example, the image captured by the smart device includes two users, a child and an adult. After the smart device obtains the interaction willingness values ​​of the child and the adult according to the above method, assuming that the interaction willingness values ​​of the child and the adult are both greater than the interaction willingness threshold, the smart device determines that the interaction willingness value of the child is larger and determines that the child is the target user. The smart device then performs subsequent user registration processing for the child.

[0068] In one embodiment of the present application, before obtaining the user's interaction willingness value based on the collected image set (step 202), the smart device can obtain the user's corresponding field of view angle in the image set based on the collected image set. The smart device determines whether the user's corresponding field of view angle is greater than or equal to the field of view angle threshold. If the user's corresponding field of view angle is greater than or equal to the field of view angle threshold, it returns to 201. For example, the field of view angle threshold can be set to 45 degrees. When the user's corresponding field of view angle is greater than or equal to 45 degrees, it returns to 201 and the user is not registered. The user's corresponding field of view angle represents the user's horizontal position in the image. The larger the user's corresponding field of view angle, the closer the user is to the edge of the image, and the smaller the user's corresponding field of view angle, the closer the user is to the center of the image. The user's corresponding field of view angle can be determined based on the field of view angle of the image acquisition device. For example, if the field of view angle of the image acquisition device is 60 degrees, if the user is at the edge of the image, the user's corresponding field of view angle is 60 degrees. If the user is at the center of the image, the user's corresponding field of view angle is 0 degrees. By judging the field of view angle corresponding to the user, users who accidentally enter the field of view of the image acquisition device of the smart device in the collected image set can be excluded, thereby avoiding registration of users who accidentally enter.

[0069] 204 , if the user's interaction willingness value is greater than or equal to the interaction willingness threshold (eg, 0.7), the smart device determines whether the collected image set and voice are consistent.

[0070] In the embodiment of the present application, the smart device can determine whether the collected image set and voice are consistent through lip movement recognition and sound source localization, that is, determine whether the collected image set and voice correspond to the same user through lip movement recognition and sound source localization. Lip movement recognition processes the collected image set, and the speaker's position (which may be called the first speaker position) can be determined through lip movement recognition. Sound source localization processes the collected voice, and the speaker's position (which may be called the second speaker position) can be determined through sound source localization. Determining whether the collected image set and voice are consistent through lip movement recognition and sound source localization is to determine whether the first speaker position and the second speaker position are consistent.

[0071] The smart device may include a microphone array, and the smart device may perform sound source localization based on the microphone array. Specific details of performing sound source localization based on the microphone array may be referred to the prior art and will not be described in detail here.

[0072] The smart device can calculate the size of the lip region in each image in the image set and determine whether lip movement has occurred based on the difference in lip region area between images. Alternatively, the smart device can extract the open and closed state of the lips in each image in the image set and detect whether lip movement has occurred based on the opening and closing amplitude.

[0073] Alternatively, the smart device can determine whether the collected image set and voice are consistent through lip reading recognition and voice recognition, that is, whether the collected image set and voice correspond to the same language through lip reading recognition and voice recognition. Lip reading recognition processes the collected image set and determines the language content (which may be called the first language content) through lip reading recognition. Voice recognition processes the collected voice and determines the language content (which may be called the second language content) through voice recognition. Determining whether the collected image set and voice are consistent through lip reading recognition and voice recognition is to determine whether the first language content and the second language content are consistent.

[0074] Smart devices can perform lip reading recognition based on image processing and pattern recognition. The recognition process can include: lip area positioning, lip region of interest feature extraction, and lip reading content recognition. Among them, lip region of interest feature extraction and lip reading content recognition are difficult. Principal Component Analysis (PCA), Discrete Time Cosine Transformation (DCT), Singular Value Decomposition (SVD), and Independent Component Correlation Algorithm (ICA) can be used to extract lip region of interest features. Classifiers such as Hidden Markov Models (HMM), Artificial Neural Networks (ANN), Template Matching Algorithm (TMA), and Support Vector Machine (SVM) can be used for lip reading content recognition.

[0075] The acoustic model, language model, and decoder are the core components of a speech recognition system. The acoustic model constructs a probabilistic mapping between input speech and output acoustic units; the language model describes the probabilistic collocation relationships between different words, making the recognized sentences more natural-sounding. The decoder combines the acoustic model's probability values ​​with the language model's scores on different collocations to filter and ultimately obtain the most likely recognition result. Speech recognition results can be obtained by calculating the maximum a posteriori probability of the speech signal based on the text. There are generally two decoding methods: dynamic decoding and static decoding. Dynamic decoding compiles the dictionary into a state network to form a search space. The general process for compiling a dictionary into a state network is as follows: all words in the dictionary are connected in parallel to form a parallel network; words in the parallel network are replaced with phoneme strings; each phoneme is split into a state sequence based on its context; and the beginning and end of the state network are connected based on the principle of consistent phoneme context to form a loop. The network compiled by dynamic decoding is generally called a linear lexicon. Its characteristic is that the state sequence of each word is strictly independent, and there is no shared node between the states of different words. This results in high memory usage and a high level of repeated computation during the decoding process. Speech recognition solutions based on static decoding are primarily implemented using finite state transducer (FST) networks. For example, a weighted finite state transducer (WFST) network integrates most components of the speech recognition process (including the pronunciation dictionary, acoustic model, grammatical information, etc.) to create a finite state transition graph. By decoding tokens, the optimal speech recognition result is obtained by searching within this finite state transition graph.

[0076] If the collected image set and voice are inconsistent, return 201.

[0077] By executing step 204 to determine the consistency of the collected image set and the voice, it is possible to effectively avoid misrecognition or misjudgment caused by the voice of a user who is not in the image set. For example, in a scenario where a user is in front of a smart device but does not speak and a large-screen device next to the user is playing audio, it is possible to effectively avoid the smart device from misidentifying the audio played on the large-screen device as the voice of the user.

[0078] In other embodiments of the present application, the consistency between the collected image set and the voice may not be judged. When the user's interaction willingness value is greater than or equal to the interaction willingness threshold, the step of judging whether the user's face has been registered based on the collected image set is executed.

[0079] 205. If the collected image set and the voice are consistent, the smart device determines whether the user's face has been registered based on the collected image set.

[0080] In one embodiment of the present application, a smart device tracks faces based on captured images. The smart device can identify a face frame in each image in an image set, combine the overlap ratio and local feature information of the face frames in multiple images, and keep tracking each face frame.

[0081] exist Figure 1 In the illustrated embodiment, a face frame to be tracked can be specified, and the body and head of the intelligent robot can be driven to rotate so that the robot faces the face frame to be tracked, thereby achieving visual tracking.

[0082] The smart device can determine whether the time of continuous face tracking exceeds a first preset time (e.g., 2 seconds). If the time of continuous face tracking exceeds the first preset time, the smart device determines whether the user's face has been registered.

[0083] Smart devices can determine whether a user's face has been registered through processes such as image preprocessing, face detection, face tracking, face alignment, and face feature extraction.

[0084] Smart devices can perform image preprocessing on each image in the image set through grayscale transformation, histogram equalization, normalization, etc. to eliminate the impact of factors such as uneven lighting, low resolution and motion blur on recognition, and achieve grayscale correction and noise filtering.

[0085] When performing face detection, the smart device can input each image in the image set into the face detection model to obtain information such as the location of the face frame in the image and the face confidence.

[0086] When tracking a face, the smart device can determine the movement trajectory and size change of the face in the image set.

[0087] When performing face alignment, the smart device can locate key facial feature points, such as eyes, nose tip, mouth corners, eyebrows, etc., based on the face image to obtain a set of facial feature points.

[0088] When extracting facial features, smart devices can input images into a neural network for convolution calculation to obtain facial features.

[0089] When performing face recognition, the smart device can measure the similarity between the extracted facial features and the facial features in the face library to determine whether the user's face has been registered.

[0090] If the user's face has been registered, return 201.

[0091] 206. If the user's face has not been registered, the smart device registers the user's face.

[0092] Smart devices can store facial features and corresponding user information in a face library.

[0093] In one embodiment of the present application, if the user's face has not been registered, the smart device can also perform facial attribute recognition based on the collected image set and store the recognized facial attributes in the face library. The facial attribute recognition performed by the smart device can include age group recognition, expression recognition, and gender recognition. For example, the recognized age groups can include 0-6 years old, 7-12 years old, 13-18 years old, 19-35 years old, 36-50 years old, and over 50 years old. The recognized expressions can include neutral, happy, surprised, angry, sad, disgusted, ashamed, and afraid. The recognized genders include male and female.

[0094] In one embodiment of the present application, face registration is independent and not affected by voiceprint registration.

[0095] 207. If the collected image set and the voice are consistent, the smart device determines whether the user's voiceprint has been registered based on the collected voice.

[0096] Determining whether a user's voiceprint has been registered based on the collected speech is voiceprint recognition. During voiceprint recognition, the smart device extracts the user's voiceprint features from the collected speech and inputs them into the voiceprint recognition model to obtain the voiceprint recognition result. The smart device can utilize the short-term stationary nature of speech and use the Mel Cepstral Transform method to extract voiceprint features from the collected speech.

[0097] In one embodiment of the present application, the smart device can use a dynamic time warping (DTW) model, a Gaussian mixture model-universal background model (GMM-UBM), a total variability model (TVM, such as x-vector), a neural network model, etc. for voiceprint recognition. The smart device can extract the voiceprint features of each speaker and perform identity recognition based on each speaker's voiceprint features.

[0098] If the user's voiceprint has been registered, return 201.

[0099] 208. If the user's voiceprint has not been registered, the smart device registers the user's voiceprint.

[0100] Smart devices can store the user's voiceprint features and corresponding user information in the voiceprint library.

[0101] In this embodiment of the present application, to improve the accuracy of voiceprint registration and avoid registering incorrect speech or noise, the smart device collects multiple (at least three) segments of speech and performs voiceprint registration based on the collected multiple segments. Each collected segment of speech is greater than or equal to a preset duration (for example, 2 seconds).

[0102] Specifically, for each voice segment, the smart device pre-registers the voiceprint for that segment and then verifies the voiceprint of other segments based on the pre-registered voice segment, that is, verifies whether the other segments correspond to the same user as the pre-registered voice segment. If, for each voice segment, the other segments correspond to the same user as the pre-registered voice segment (i.e., the voiceprint verification passes), the voiceprint registration is successful. If, for a certain voice segment, the other segments do not correspond to the same user as the pre-registered voice segment (i.e., the voiceprint verification fails), the voiceprint registration fails.

[0103] Figure 4 This is a schematic diagram of voiceprint registration.

[0104] like Figure 4 As shown in the figure, in the voiceprint registration process, the smart device first collects four voice segments A, B, C, and D. For example, voice segment A can be the 0th to 2nd second of the voice, voice segment B can be the 2nd to 4th second of the voice, voice segment C can be the 4th to 6th second of the voice, and voice segment D can be the 6th to 8th second of the voice. The duration of each voice segment can be 2 seconds. Then, the smart device uses A, B, C, and D for pre-registration, respectively, to obtain pre-registered voice segments A, B, C, and D. Next, use pre-registered voice segment A to perform voiceprint verification on voice segments B, C, and D to determine whether voice segments B, C, and D correspond to the same user as voice segment A. Use pre-registered voice segment B to perform voiceprint recognition on voice segments A, C, and D to determine whether voice segments A, C, and D correspond to the same user as voice segment B. Use pre-registered voice segment C to perform voiceprint recognition on voice segments A, B, and D to determine whether voice segments A, B, and D correspond to the same user as voice segment C. Use pre-registered voice segment D to perform voiceprint recognition on voice segments A, B, and C to determine whether voice segments A, B, and C correspond to the same user as voice segment D. If voiceprint verification passes for pre-registered voice segments A, B, C, and D, voiceprint registration is successful.

[0105] In this embodiment, voiceprint registration is affected by face registration. If face registration fails, voiceprint registration also fails, and the voiceprint feature and corresponding user information will not be stored.

[0106] It should be understood that in other embodiments of the present application, voiceprint registration may not be affected by face registration, and voiceprint registration may be performed even if face registration and recognition are performed.

[0107] In other embodiments of the present application, the smart device may only register the user's face, or only register the user's voiceprint.

[0108] Currently, existing smart devices register biometric information such as faces and voiceprints based on pre-set processes or existing audio and video. They lack a mechanism for dynamic and seamless interaction with robots, registration, and decision-making.

[0109] The existing contactless registration process does not judge the collected images and voices, nor does it determine the registration timing, making it easy to use the wrong images and / or voices for registration. For example, the images and voices collected by the smart device may not correspond to the same user, and the existing contactless registration process may register different images and voices as the biometric information of the same user. For another example, a user may accidentally enter the field of view of the smart device and be captured in the image. This user is not the target registration user, and the existing contactless registration process may incorrectly register this user.

[0110] According to the user registration method provided in the embodiment of the present application, the smart device determines whether the user's interaction willingness value is greater than or equal to the interaction willingness threshold during the conversation with the user, and determines whether the collected image set and voice meet the consistency. When the user's interaction willingness value is greater than or equal to the interaction willingness threshold, and the collected image set and voice meet the consistency, the user's face and / or voiceprint are registered. The present application can register users naturally and accurately, and improve the accuracy of non-sensing registration. In addition, the embodiment of the present application determines whether the user meets the conditions for non-sensing registration based on the collected image set rather than a single image, which can improve the accuracy of registration.

[0111] In an embodiment of the present application, during the user registration process, the smart device can obtain relevant user information based on the collected image set and voice, such as the user's username, age group, and gender. When registering the user's face and / or voiceprint, the relevant user information is associated with the user's facial features and / or voiceprint features and stored in the smart device. For example, the smart device can ask the user for their name, age, gender, etc. during a conversation and obtain this information based on the user's answers. In another example, the smart device can identify the user's age, gender, etc. by using a facial image to obtain this information.

[0112] Alternatively, when registering a user's face and / or voiceprint, the smart device can automatically assign a username or user ID to the user, associate the assigned username or user ID with the user's facial features and / or voiceprint features, and store them in the smart device.

[0113] After the contactless registration is completed, the smart device stores the user's facial features and voiceprint features. It can also store relevant information about the user, such as the user name, age group, gender, etc. obtained during the conversation with the user. Based on the stored user's facial features and voiceprint features, the smart device can authenticate the user and provide relevant services to registered users. It should be noted that if the user's relevant information (such as the user name) is not obtained during the user conversation, smart recognition can be set.

[0114] For example, for a registered user, the smart device can collect interaction information between the user and the smart device. When the user is detected using / logging in to the smart device again, the smart device can interact with the user based on the user's historical interaction information.

[0115] In one embodiment of the present application, the smart device can actively interact with the user based on the user's historical interaction information. Active interaction means that the smart device actively outputs interactive content to the user. Different users may interact with different smart devices with different content. For example, when children interact with smart devices historically, they often demand to play fairy tales; when teenagers interact with smart devices historically, they often demand to play pop music; when middle-aged people interact with smart devices historically, they often demand to play current affairs news. Smart devices can output different interactive content for different registered users when actively interacting with users. The output content can be content that the user is interested in, thereby increasing the success rate of active interaction with users and improving the stickiness with users.

[0116] For registered users, the smart device can also determine the content of interaction with the user based on stored user information (such as age group, gender, etc.).

[0117] In an embodiment of the present application, the smart device can determine the interaction type of the registered user (for example, determine the interaction type of the registered user based on the historical interaction information of the registered user) and preset the active interaction content corresponding to each interaction type. The interaction type may include a greeting interaction type, a riddle guessing interaction type, a music playing interaction type, a topic chat initiation interaction type, a dance interaction type, etc. The active interaction content of each interaction type may also include multiple different contents. For example, the active interaction content of the greeting interaction type may include: "Hi, the weather is nice today, how are you feeling?", "What's up, man", or "It's another beautiful day today, let's relax", etc.

[0118] The smart device can output the current interaction content corresponding to the user's preferred interaction type based on the user's preferred interaction type. For example, if the user prefers a riddle-guessing interaction type, a riddle can be played. The user's preferred interaction type can be the interaction type that interacts with the user the most.

[0119] The smart device can also play the current interactive content of the interaction type that the user interacted with most recently. For example, if the interaction type that the user interacted with most recently was a music playing interaction type, the current interactive content of the music playing interaction type can be outputted.

[0120] It should be noted that, according to different needs, the execution order of each step of the user registration method provided in the embodiment of the present application can be changed, some steps can be omitted, and some steps can be changed.

[0121] Figure 5 This is a flowchart of a user registration method provided by another embodiment of the present application.

[0122] Figure 2 In the illustrated embodiment, the smart device first determines whether the user's interaction willingness value is greater than or equal to the interaction willingness threshold, and determines whether the collected image set and voice are consistent. When the user's interaction willingness value is greater than or equal to the interaction willingness threshold, and the collected image set and voice are consistent, it determines whether the user's face and voiceprint have been registered.

[0123] Figure 5 In the illustrated embodiment, the smart device first determines whether the user's face and / or voiceprint has been registered. If the user's face and / or voiceprint has not been registered, the smart device then determines whether the user's interaction willingness value is greater than or equal to the interaction willingness threshold, and also determines whether the collected image set and voice are consistent. If the user's interaction willingness value is greater than or equal to the interaction willingness threshold, and the collected image set and voice are consistent, the user's face and / or voiceprint is registered.

[0124] Smart devices can register only face or voiceprint, or both face and voiceprint. If the smart device registers both face and voiceprint, when it determines that the user's face has been registered but the voiceprint has not been registered, or the user's voiceprint has been registered but the face has not been registered, it can be as follows: Figure 5 The steps of determining whether the user's interaction willingness value is greater than or equal to the interaction willingness threshold and determining whether the collected image set and voice are consistent are performed. When the user's interaction willingness value is greater than or equal to the interaction willingness threshold and the collected image set and voice are consistent, the user's voiceprint or face is registered.

[0125] In another embodiment of the present application, when it is determined that the user's face has been registered but the voiceprint has not been registered, or the user's voiceprint has been registered but the face has not been registered, the user's voiceprint or face can be directly registered.

[0126] Figure 6 This is a flowchart of a user registration method provided by another embodiment of the present application.

[0127] Figure 2In the embodiment shown, the smart device executes the steps in sequence. In other embodiments of the present application, the user registration method may include multiple sub-processes, and the smart device may process some sub-processes in parallel.

[0128] like Figure 6 As shown, the user registration method provided by the embodiment of the present application includes an interaction willingness recognition sub-process, a face registration sub-process, and a voiceprint registration sub-process. The smart device can process the face registration sub-process and the voiceprint registration sub-process in parallel. The interaction willingness recognition sub-process may include: collecting the user's image set and voice (601); obtaining the user's interaction willingness value based on the collected image set (602); judging whether the user's interaction willingness value is greater than or equal to the interaction willingness threshold (603); if the user's interaction willingness value is greater than or equal to the interaction willingness threshold, judging whether the collected image set and voice are consistent (604). The face registration sub-process may include: if the collected image set and voice are consistent, judging whether the user's face has been registered based on the collected image set (605); if the user's face has not been registered, the smart device registers the user's face (606). The voiceprint registration sub-process may include: if the collected image set and voice are consistent, the smart device judges whether the user's voiceprint has been registered based on the collected voice (607); if the user's voiceprint has not been registered, the smart device registers the user's voiceprint (608).

[0129] According to the user registration method provided in the embodiment of the present application, users can complete the registration without any feeling in a variety of scenarios. The following describes the working process of the user registration method provided in the embodiment of the present application through some specific scenarios.

[0130] Scenario 1: The user wakes up the smart device to interact:

[0131] User: Classmate XX.

[0132] Smart device: Hi, what’s your name?

[0133] User: I’m Xiaohong, nice to meet you.

[0134] Smart device: Nice to meet you, Xiaohong. What can I do for you?

[0135] User: What’s the weather like today?

[0136] Smart device: Today in Beijing it will be cloudy, 18℃-25℃.

[0137] In one implementation of the present invention, a smart device may collect a user's image set and voice after the smart device finishes a current speech and before the smart device begins its next speech. For example, the smart device may collect the user's image set and voice via its camera and microphone after saying "Hi, hello, what's your name?" and before saying "Nice to meet you, Xiaohong. What can I do for you?" and / or after saying "Nice to meet you, Xiaohong. What can I do for you?" and before saying "Today's weather in Beijing is cloudy, 18-25°C." Based on the collected image set, the smart device determines whether the user's interaction willingness value is greater than or equal to an interaction willingness threshold, determines whether the collected image set and voice are consistent, and determines whether the user's face and voiceprint have been registered. If the user's interaction willingness value is greater than or equal to the interaction willingness threshold, the collected image set and voice are consistent, and the user's face and / or voiceprint have not been registered, the smart device registers the user's face and / or voiceprint.

[0138] In another implementation of the present embodiment, after the user wakes up the smart device by saying "XX classmate," the smart device may capture an image of the user and determine whether the user's face has been registered based on the captured image. If the user's face has not been registered, the smart device may capture a set of images and voice from the user via its camera and microphone after saying "Hi, hello, what's your name?" and before saying "Nice to meet you, Xiaohong. What can I do for you?" and / or after saying "Nice to meet you, Xiaohong. What can I do for you?" and before saying "Today's weather in Beijing is cloudy, 18°C-25°C." Based on the captured image set, the smart device determines whether the user's interaction willingness value is greater than or equal to an interaction willingness threshold, and determines whether the captured image set and voice are consistent. If the user's interaction willingness value is greater than or equal to the interaction willingness threshold, the captured image set and voice are consistent, and the user's face is registered. The smart device may also determine whether the user's voiceprint has been registered based on the captured voice. If the user's voiceprint has not been registered, the smart device may register the user's voiceprint.

[0139] In another implementation of the present embodiment, the smart device may begin capturing images of the user after the user wakes up the device by saying "XX classmate." Based on the captured images, the user's corresponding field of view angle is obtained. The smart device determines whether the user's corresponding field of view angle is greater than or equal to a field of view angle threshold. If so, the smart device determines whether the user's face has been registered. If the user's face has not been registered, the smart device captures the user's image set and voice via the camera and microphone after saying "Hi, hello, what's your name?" and before saying "Nice to meet you, Xiaohong. What can I do for you?" and / or after saying "Nice to meet you, Xiaohong. What can I do for you?" and before saying "Today's weather in Beijing is cloudy, 18°C-25°C." Based on the captured images, the smart device obtains the user's interaction willingness value, determines whether the user's interaction willingness value is greater than or equal to the interaction willingness threshold, and determines whether the captured image set and voice are consistent. If the user's interaction willingness value is greater than or equal to the interaction willingness threshold, the captured image set and voice are consistent, and the user's face is registered. The smart device can also determine whether the user's voiceprint has been registered based on the collected voice. If the user's voiceprint has not been registered, it can also register the user's voiceprint.

[0140] In other implementations of the embodiments of the present application, the smart device may start collecting images of the user after the user says "XX classmate" to wake up the smart device. Based on the collected image, the field of view angle corresponding to the user in the image is obtained. The smart device determines whether the field of view angle corresponding to the user is greater than or equal to the field of view angle threshold. If the field of view angle corresponding to the user is greater than or equal to the field of view angle threshold, it determines whether the user's face has been registered. If the user's face has not been registered, the smart device collects the user's image set and voice through the camera and microphone. Based on the collected image set, the smart device obtains the user's interaction willingness value, determines whether the user's interaction willingness value is greater than or equal to the interaction willingness threshold, and determines whether the collected image set and voice meet consistency. If the user's interaction willingness value is greater than or equal to the interaction willingness threshold, the collected image set and voice meet consistency, and the user's face is registered. The smart device can also determine whether the user's voiceprint has been registered based on the collected voice. If the user's voiceprint has not been registered, it can also register the user's voiceprint.

[0141] Scenario 2: When the smart device detects the presence of a user around the smart device, the smart device actively interacts with the user.

[0142] Smart device: Hi, master, good morning!

[0143] User: Good morning!

[0144] Smart device: Today is another beautiful day, let's relax!

[0145] User: Let’s play some music.

[0146] Smart device: What kind of music does the owner want to listen to?

[0147] User: Play "Going Down the Mountain".

[0148] In one implementation of the embodiment of the present application, the smart device may collect the user's image set and voice through a camera and microphone after saying "Hi, master, good morning!" and before saying "Today is another beautiful day, let's relax!" and / or after saying "Today is another beautiful day, let's relax!" and before saying "What music would you like to listen to?" The smart device obtains the user's interaction willingness value based on the collected image set, determines whether the user's interaction willingness value is greater than or equal to an interaction willingness threshold, determines whether the collected image set and voice are consistent, and determines whether the user's face and voiceprint have been registered. If the user's interaction willingness value is greater than or equal to the interaction willingness threshold, the collected image set and voice are consistent, and the user's face and / or voiceprint have not been registered, the smart device registers the user's face and / or voiceprint.

[0149] In other implementations of the embodiments of the present application, upon detecting the presence of a user near the smart device, the smart device may capture an image of the user and determine whether the user's face has been registered based on the captured image. If the user's face has not been registered, the smart device may capture a set of images and voice from the user after saying "Hi, master, good morning!" and before saying "It's another beautiful day, let's relax!" and / or after saying "It's another beautiful day, let's relax!" and before saying "What music would you like to listen to?" Based on the captured image set, the smart device obtains the user's interaction willingness value, determines whether the user's interaction willingness value is greater than or equal to an interaction willingness threshold, and determines whether the captured image set and voice are consistent. If the user's interaction willingness value is greater than or equal to the interaction willingness threshold, the captured image set and voice are consistent, and the user's face is registered. The smart device may also determine whether the user's voiceprint has been registered based on the captured voice. If the user's voiceprint has not been registered, the smart device may register the user's voiceprint.

[0150] It should be noted that a smart device can register a user's face and voiceprint in a single interaction, or in two interactions. The smart device can register the user's face in the first interaction and the voiceprint in the second. Alternatively, the smart device can register the user's voiceprint in the first interaction and the face in the second. The following uses a scenario where a user interacts with a robot as an example to illustrate registering a user's face and voiceprint in two steps.

[0151] The user's face is registered during the first interaction, and the user's voiceprint is registered during the second interaction:

[0152] The user interacts with the robot for the first time: During this first interaction, the user does not speak, but stares at the robot for a while. The robot determines that the user's willingness to interact is greater than or equal to the interaction willingness threshold, and since the user's face has not been registered before, it registers the user's face. The robot creates a new user data entry. Since no speech was spoken during this interaction, the voiceprint information in this user data entry is temporarily missing.

[0153] The user interacts with the robot for the second time: After a while, the user interacts with the robot for the second time. The robot matches the user's registered face and completes the user's voiceprint registration based on the collected voice, completing the voiceprint information in the previously created user data.

[0154] The user's voiceprint is registered during the first interaction, and the user's face is registered during the second interaction:

[0155] The user interacts with the robot for the first time while wearing a mask. The robot determines that the user's willingness to interact is greater than or equal to the threshold, and that the user's voiceprint has not been registered. Therefore, the robot registers the user's voiceprint. However, since the user is wearing a mask, the robot does not register the user's face. Therefore, the robot creates a new user data entry, temporarily missing the user's facial information.

[0156] The user interacts with the robot for the second time: After a while, the user interacts with the robot again, this time without a mask. The robot matches the collected voice to the user's previously registered voiceprint. The robot then registers the collected face with the previously created user data, completing the previously missing facial information.

[0157] Figure 7 It is a detailed flowchart of how smart devices obtain the user's interaction willingness value based on the collected image set.

[0158] Generally speaking, if a person's face is closer to a smart device and facing the smart device over a period of time, it indicates that the user has a stronger willingness to interact with the smart device, and vice versa. Figure 7 The process shown is to obtain the user's interaction willingness value based on the user's face angle and the distance between the user and the smart device.

[0159] like Figure 7 As shown in FIG, the smart device obtains the user's interaction willingness value based on the collected image set, including:

[0160] 701 , the smart device obtains the user's face angle and the distance between the user and the smart device in each image of the image set.

[0161] The face angle of the user in the image refers to the angle between the face in the image and the face of the user when the user is facing the smart device. In an embodiment of the present application, images containing faces at various angles can be pre-stored, and the smart device determines the face angle of the user in each image in the image set by comparing the images containing the user's face in the image set with the stored images containing faces at various angles. The pre-stored images containing faces at various angles (i.e., face angles) can be images of any user. Multiple images containing faces can be pre-stored, each image corresponding to a face angle, and marked with the corresponding face angle.

[0162] In an embodiment of the present application, the correspondence between the distance between the user and the smart device and the pixel size of the user's head can also be pre-stored, and the distance between the user and the smart device in each image in the image set can be determined based on the pixel size of the user's head in the image in the image set and the correspondence.

[0163] It should be understood that the method for obtaining the facial angle of the user in the image and the distance between the user and the smart device is merely an example. In other embodiments of the present application, the smart device may obtain the facial angle of the user in the image and the distance between the user and the smart device in other ways, without limitation. For example, the smart device may include a distance sensor (such as an ultrasonic sensor, an infrared sensor, a proximity sensor, etc.), and the smart device may use the distance sensor to measure the distance between the user and the smart device.

[0164] At 702 , the smart device inputs the user's facial angle in each image and the distance between the user and the smart device into an interaction willingness value model to obtain the user's initial interaction willingness value.

[0165] For example, let's take the image set in the above embodiment as an example. The user's facial angles in these five images are 5°, 6°, 5°, 7°, and 5°, respectively. The distances between the user and the smart device in these five images are 600mm, 600mm, 590mm, 610mm, and 580mm, respectively. A 1×10 vector x0, x0 = [5, 600, 6, 600, 5, 590, 7, 610, 5, 580], can be constructed based on 5°, 6°, 5°, 7°, 5°, 600mm, 600mm, 590mm, 610mm, and 580mm. This vector x0 is then input into the interaction willingness model. In vector x0, 5, 6, 5, 7, and 5 represent the user's facial angles in the five images, respectively, and 600, 600, 590, 610, and 580mm represent the distances between the user and the smart device in the five images, respectively.

[0166] The interaction willingness value model in the embodiment of the present application can adopt a three-layer neural network model with a small number of parameters. Among them, the interaction willingness value model can include a first hidden layer, a second hidden layer, and an output layer. The first hidden layer can have 10 neurons. The matrix calculation of the first layer is x0×H1+b1, where H1 is a 10×10 matrix and b1 is a 1×10 vector. Then, after tanh() calculation, a new 1×10 vector x1 is obtained. The second hidden layer can have 6 neurons. The matrix calculation of the second layer is x1×H2+b2, where H2 is a 10×6 matrix and b2 is a 1×10 vector. Then, after tanh() calculation, a new 1×6 vector x2 is obtained. The calculation of the output layer is x2×H3+b3, where H3 is a 6×1 matrix. Among them, H1, H2, H3, b1, b2, and b3 are parameters of the interaction willingness value model. The result of the output layer is calculated by tanh() to obtain the output value of the interaction willingness value model, that is, the user's initial interaction willingness value. The user's initial interaction willingness value is between 0 and 1. The larger the user's initial interaction willingness value is, the higher the user's willingness to interact with the smart device.

[0167] Before inputting the user's facial angle and the distance between the user and the smart device in each image into the interaction willingness value model, the interaction willingness value model needs to be trained using training data.

[0168] The training data includes a plurality of training samples with labeling information (for example, labeled as 0 or 1). Each training sample includes the user's facial angle and the distance between the user and the smart device, and the labeling information indicates whether the subject has the willingness to interact. Exemplarily, in an embodiment of the present application, training samples can be collected by a smart device in a simulated experimental site. A process for collecting training data is as follows: a button is pre-set on the smart device. When the subject has the willingness to interact, the subject can press the button, so that the smart device can mark the training sample as 1, indicating that the collected data is a training sample in which the subject has the willingness to interact. When the subject has no willingness to interact, the subject can short press or not press the button, so that the smart device can mark the training sample as 0, indicating that the collected data is a training sample in which the subject has no willingness to interact.

[0169] In the embodiments of the present application, collecting training samples in a simulated experimental field can collect more realistic and accurate data, improving the accuracy of the interaction willingness value model. In addition, in the embodiments of the present application, the collected training samples are directly annotated on-site, eliminating the need for subsequent user analysis and annotation, saving subsequent manpower.

[0170] The interaction willingness value model is trained using training samples to obtain a trained interaction willingness value model. The training process is not described in detail in the embodiments of this application. It should be understood that the interaction willingness value model in the embodiments of this application represents the mapping relationship between the angle of the face, the distance between the user and the smart device, and the interaction willingness value. In other words, the user's facial angle and the distance between the user and the smart device in an image are input into the interaction willingness value model to obtain the user's interaction willingness value.

[0171] In another implementation of the embodiment of the present application, the user's interaction willingness value can be determined not by training the interaction willingness value model, but by manually formulated rules. The above-mentioned manually formulated rules can predefine a certain facial angle of the user or a certain interval, and / or a certain distance between the user and the smart device or a certain interval, corresponding to a certain interaction willingness value or a certain interval. Thus, after the smart device obtains the user's facial angle and the distance between the user and the smart device, the user's interaction willingness value can be obtained according to the above-mentioned manually formulated rules. Exemplarily, the manually specified rules may include: if the user's facial angle is less than 10 degrees and the distance between the user and the smart device is less than 600mm, the user's interaction willingness value is greater than 0.7.

[0172] 703 , the smart device recognizes the user's actions based on the collected image set.

[0173] Recognizing the user's actions based on the collected image set is to identify the user's actions corresponding to each image in the image set.

[0174] The user's actions can include facial movements, head movements, and body movements. Facial movements can include grinning, sticking out the tongue, pouting, glaring, etc. Head movements can include nodding, shaking the head, etc. Body movements can include waving hands, making a heart shape with the hands, and making an OK gesture, etc.

[0175] A deep learning network-based classifier can be used to classify images in an image collection to identify user actions. Alternatively, key points of the human body can be extracted from the collected image collection to identify user actions based on the key points.

[0176] The specific content of identifying the user's action based on the collected image set can be referred to the existing technology and will not be described in detail here.

[0177] 704 , the smart device obtains the user's interaction willingness value according to the user's initial interaction willingness value and the user's action.

[0178] In one embodiment of the present application, if the user's action includes a preset action, the preset interaction willingness value can be added to the initial interaction willingness value to obtain the user's interaction willingness value; if the user's action does not include a preset action, the initial interaction willingness value can be used as the user's interaction willingness value.

[0179] For example, the preset action may be waving, nodding, etc. In the embodiment of the present application, based on the initial interaction willingness value of the user obtained according to the above interaction willingness value model, if it is detected that the user performs the above preset action, it can be determined that the user has a strong interaction willingness, and a certain value can be added to the initial interaction willingness value to obtain the user's interaction willingness value.

[0180] In an embodiment of the present application, the smart device uses the aforementioned interaction willingness value model to determine a user's initial interaction willingness value. To more accurately determine whether the user is willing to actively interact, the user's actions in the image can be combined to further determine the user's interaction willingness value. Preset actions can be configured in the smart device. If the user's actions identified in the image collection include the preset actions, the preset interaction willingness value is added to the initial interaction willingness value to obtain the user's interaction willingness value. If the user's actions do not include the preset actions, the initial interaction willingness value is used as the user's interaction willingness value. Preset actions can include waving, nodding, or making an "OK" gesture. For example, if the preset interaction willingness value is 0.2, and the smart device uses the interaction willingness value model to obtain the user's initial interaction willingness value of 0.7, and recognizes that the user's actions include waving, the user's interaction willingness value is 0.7 + 0.2, which is 0.9. In other words, in an embodiment of the present application, the smart device uses the interaction willingness value model to obtain the user's initial interaction willingness value and then combines the identified user actions to obtain the user's final interaction willingness value.

[0181] In another implementation of the embodiment of the present application, the process of the smart device obtaining the user interaction willingness value may not include steps 703 and 704, but directly uses the initial interaction willingness value obtained in step 702 based on the user's facial angle and the distance between the user and the smart device as the final user interaction willingness value.

[0182] Figure 8 This is a schematic diagram of the structure of a smart device provided in an embodiment of the present application. Figure 8 As shown, the smart device 80 may include: a processor 801, a memory 802, a wireless communication module 803, an audio module 804, a microphone 805, a sensor 806, a camera 807, a display screen 808, etc. It is understood that the structure shown in this embodiment does not constitute a specific limitation on the smart device 80. In other embodiments of the present application, the smart device 80 may include more or fewer components than shown, or combine or split certain components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0183] The processor 801 may include one or more processing units. For example, the processor 801 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, a display processing unit (DPU), and / or a neural-network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors.

[0184] Processor 801 may be provided with a memory for storing instructions and data. In some embodiments, the memory in processor 801 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 801. If processor 801 needs to use the same instruction or data again, it can directly access the memory, avoiding repeated access, reducing the waiting time of processor 801, and improving the efficiency of smart device 80.

[0185] In some embodiments, the processor 801 may include one or more interfaces. The interface may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface and / or a universal serial bus (USB) interface, etc. It will be understood that the interface connection relationship between the modules illustrated in the embodiment of the present invention is only a schematic illustration and does not constitute a structural limitation of the smart device 80. In other embodiments of the present application, the smart device 80 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.

[0186] The memory 802 can be used to store one or more computer programs, which include instructions. The processor 801 can enable the smart device 80 to perform relevant actions in the embodiments of the present application by running the instructions stored in the memory 802. The memory 802 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system; the program storage area can also store one or more applications. The data storage area can store data (such as photos, contacts) created during the use of the smart device 80. In addition, the memory 802 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. In some embodiments, the processor 801 can enable the smart device 80 to perform various functional applications and data processing by running instructions stored in the memory 802, and / or instructions stored in a memory provided in the processor 801.

[0187] The wireless communication function of the smart device 80 can be implemented through the wireless communication module 803. The wireless communication module 803 can provide wireless communication solutions including wireless local area networks (WLAN), Bluetooth, global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. applied on the smart device 80. The wireless communication module 803 can be one or more devices integrating at least one communication processing module. The wireless communication module 803 in the embodiment of the present application is used to implement the transceiver function of the electronic device.

[0188] The smart device 80 can implement audio functions such as music playback and recording through the audio module 804, microphone 805, etc. Among them, the audio module 804 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 804 can also be used to encode and decode audio signals. In some embodiments, the audio module 804 can be set in the processor 801, or some functional modules of the audio module 804 can be set in the processor 801. The smart device 80 can be provided with at least one microphone 805. In other embodiments, the smart device 80 can be provided with two microphones 805, which can not only collect sound signals but also implement noise reduction functions. In other embodiments, the smart device 80 can also be provided with three, four or more microphones 805 to collect sound signals, reduce noise, identify sound sources, implement directional recording functions, etc.

[0189] Sensor 806 may include a pressure sensor 806A, a distance sensor 806B, a proximity light sensor 806C, and the like. Pressure sensor 806A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 806A may be provided on display screen 808, and smart device 80 detects the intensity of a touch operation based on pressure sensor 806A. Distance sensor 806B is used to measure distance. Smart device 80 may measure distance using infrared or laser. Proximity light sensor 806C may include a light-emitting diode (LED) and a photodetector. The LED may be an infrared LED. The photodetector may be a photodiode. Smart device 80 emits infrared light through the LED. Smart device 80 uses the photodetector to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it may be determined that an object is near smart device 80. When insufficient reflected light is detected, it may be determined that no object is near smart device 80.

[0190] The smart device 80 can implement a shooting function through one or more cameras 807. In addition, the smart device 80 can implement a display function through a display screen 808. The display screen 808 is used to display images, videos, etc. The display screen 808 includes a display panel. The display panel can use a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active matrix organic light-emitting diode or an active matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), Mini LED, Micro LED, Micro OLED, quantum dot light emitting diodes (QLED), etc. In some embodiments, the smart device 80 may include one or more display screens 808.

[0191] In the embodiment of the present application, the distance sensor 806B and the proximity light sensor 806C can be used to detect whether there is a user around the smart device 80, the camera 807 can collect an image set of the user, and the processor 801 is used to perform the actions in the above embodiment. Figure 8 The smart device 80 shown can implement the user registration method in the above embodiment. It should be understood that Figure 8 The structural description in FIG. 8 is an example of the smart device 80 .

[0192] This embodiment further provides a computer storage medium, which stores computer instructions. When the computer instructions are executed on a smart device, the smart device executes the above-mentioned related method steps to implement the user registration method in the above-mentioned embodiment.

[0193] This embodiment further provides a computer program product. When the computer program product is run on a smart device, the smart device executes the above-mentioned related steps to implement the user registration method in the above-mentioned embodiment.

[0194] In addition, an embodiment of the present application also provides a device, which can specifically be a chip, component or module, and the device may include a connected processor and memory; wherein the memory is used to store computer execution instructions, and when the device is running, the processor can execute the computer execution instructions stored in the memory to enable the chip to execute the user registration method in the above-mentioned method embodiments.

[0195] Among them, the smart device, computer storage medium, computer program product or chip provided in this embodiment is used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be repeated here.

[0196] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0197] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0198] The units described as separate components may or may not be physically separate, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0199] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0200] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0201] The above is only a specific embodiment of the present application, but the scope of protection of this application is not limited to this. Any changes or substitutions within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A user registration method, applied to a smart device, characterized in that: The method comprises: Collect user's image set and voice; Obtaining the field of view angle corresponding to the user in the image set; If the field of view angle corresponding to the user is less than or equal to the field of view angle threshold, obtaining an interaction willingness value of the user based on the image set, wherein the interaction willingness value is used to describe the strength of the user's willingness to interact with the smart device; If the interaction willingness value is greater than or equal to the interaction willingness threshold, the user's face and / or voiceprint is registered.

2. The user registration method according to claim 1, wherein: Obtaining the user's interaction willingness value according to the image set includes: acquiring behavior information of the user according to the image set; The interaction willingness value is obtained according to the user's behavior information.

3. The user registration method according to claim 2, wherein: The acquiring the user's behavior information according to the image set includes: Obtaining the facial angle of the user and the distance between the user and the smart device in each image of the image set; Inputting the face angle and the distance into an interaction willingness value model to obtain the user's initial interaction willingness value; recognizing an action of the user based on the image set; The interaction willingness value of the user is obtained according to the initial interaction willingness value and action of the user.

4. The user registration method according to claim 3, wherein: The acquiring of the user's interaction willingness value according to the user's initial interaction willingness value and action includes: If the user's action includes a preset action, adding the preset interaction willingness value to the initial interaction willingness value to obtain the user's interaction willingness value; If the user's action does not include a preset action, the initial interaction willingness value is used as the user's interaction willingness value.

5. The user registration method according to claim 3, wherein: The movements include facial movements, head movements and body movements.

6. The user registration method according to claim 1, wherein: Before registering the user's face and / or voiceprint, the method further includes: Determining whether the image set and the speech are consistent; When the interaction willingness value is greater than or equal to an interaction willingness threshold, and the image set is consistent with the voice, the registration of the user's face and / or voiceprint is performed.

7. The user registration method according to claim 6, wherein: Determining the consistency between the image set and the speech includes: Determining whether the image set and the voice correspond to the same user; and / or It is determined whether the image set and the speech correspond to the same language content.

8. The user registration method according to claim 7, wherein: Determining whether the image set and the voice correspond to the same user includes: Performing lip movement recognition based on the image set to obtain a first speaker position; Performing sound source localization based on the speech to obtain the position of the second speaker; Determine whether the first speaker position and the second speaker position are consistent.

9. The user registration method according to claim 7, wherein: Determining whether the image set and the speech correspond to the same language content includes: performing lip reading recognition on the image set to obtain first language content; performing speech recognition on the speech to obtain second language content; Determine whether the first language content and the second language content are consistent.

10. The user registration method according to any one of claims 1 to 9, characterized in that: Obtaining the user's interaction willingness value according to the image set includes: If the image set includes multiple users, obtaining an interaction willingness value of each user among the multiple users; The method further comprises: If the interaction willingness values ​​of more than one user among the multiple users are greater than or equal to the interaction willingness threshold, the user with the largest interaction willingness value among the multiple users is selected as the target user.

11. The user registration method according to any one of claims 1 to 9, wherein: Registering the user's voiceprint includes: Collect at least three voice segments of the user, each segment being greater than or equal to a preset duration; Pre-register voiceprint for each speech segment; Based on each speech segment, performing voiceprint verification on other speech segments in the at least three speech segments; If the voiceprint verification of other voice segments is passed based on each voice segment, the user's voiceprint registration is successful.

12. The user registration method according to any one of claims 1 to 9, wherein: After registering the user's face and / or voiceprint, the method further includes: Collecting interaction information between the user and the smart device; Actively interact with the user according to the interaction information.

13. A computer-readable storage medium, characterized in that The method comprises computer instructions, which, when executed on a smart device, enable the smart device to execute the user registration method according to any one of claims 1 to 12.

14. A smart device, characterized in that: The smart device includes a processor and a memory, the memory is used to store instructions, and the processor is used to call the instructions in the memory, so that the smart device executes the user registration method according to any one of claims 1 to 12.

15. A chip system, characterized in that: The chip system is applied to a smart device; the chip system includes an interface circuit and a processor; the interface circuit and the processor are interconnected through a line; the interface circuit is used to receive a signal from the memory of the smart device and send a signal to the processor, the signal including a computer instruction stored in the memory; when the processor executes the computer instruction, the chip system executes the user registration method according to any one of claims 1 to 12.

16. A computer program product, characterized in that When the computer program product is run on a smart device, the smart device is enabled to execute the user registration method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • User registration method and apparatus of intelligent robot

    CN106295299A

  • Voiceprint creation and registration method and device

    CN107492379A

  • Identity authentication method and device

    CN111199032A

  • Interaction method and device, electronic equipment and storage medium

    CN111931897A

  • Man-machine interaction control method, intercom calling method and related devices

    CN112102546A