Human-computer interaction information processing method, human-computer interaction device and storage medium
By collecting and processing fingerprint, voiceprint, and dynamic image data, and combining them to calculate similarity, the problem of inaccurate recognition in human-computer interaction devices when there is insufficient data in a single dimension has been solved, thereby improving recognition accuracy and business processing efficiency.
Patent Information
- Application Number
- CN202211651785.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-21
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2042-12-21
AI Technical Summary
Existing human-computer interaction devices fail to accurately identify data when data in a certain dimension is insufficient, leading to difficulties in business processing. Furthermore, existing multi-dimensional data solutions have not been optimized for single-dimensional problems.
The system collects fingerprint data, voiceprint data, and dynamic images of the target user, preprocesses them, and calculates the similarity of each data point. When the amount of data in a certain dimension is insufficient, it combines data from multiple dimensions to calculate the similarity and adjusts the similarity threshold to ensure recognition accuracy.
When the amount of data in one dimension is insufficient, multi-dimensional data combination calculations are used to ensure the recognition accuracy of human-computer interaction devices and improve business processing efficiency.
Smart Images

Figure CN115861658B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of computer processing, and in particular to a human-computer interaction information processing method, a human-computer interaction device, and a storage medium. BACKGROUND
[0002] In the existing human-computer interaction process, if the authentication of the device to the user is only authenticated and identified by one-dimensional data, once the user cannot normally provide the one-dimensional data, it will bring trouble to the user and the service provider, and is not conducive to the handling of business.
[0003] However, in the existing scheme of using multi-dimensional data to authenticate and identify the user, there is no special optimization for the scene where the amount of data of a certain dimension is collected, which is prone to identification errors. SUMMARY
[0004] The embodiments of the present application provide a human-computer interaction information processing method, a human-computer interaction device, and a storage medium, which solve the problem of inaccurate identification of the human-computer interaction device in the case where a certain dimension of data does not meet the requirements in the current technical solution.
[0005] To solve the above technical problems, the present application:
[0006] In a first aspect, a human-computer interaction information processing method is provided, applied to a human-computer interaction device, and the method comprises:
[0007] Collecting fingerprint data, voiceprint data, and dynamic images of a target user;
[0008] Preprocessing the fingerprint data, voiceprint data, and dynamic images to obtain corresponding fingerprint feature sequences, voiceprint feature sequences, and dynamic feature sequences;
[0009] When the amount of data of each sequence in the fingerprint feature sequence, the voiceprint feature sequence, and the dynamic feature sequence meets a pre-designed calculation condition, the similarity between the fingerprint feature sequence, the voiceprint feature sequence, and the dynamic feature sequence and the pre-stored target user feature data is calculated; when the similarity between any feature in the fingerprint feature sequence, the voiceprint feature sequence, and the dynamic feature sequence and the pre-stored target user feature data is lower than a first threshold value, the corresponding similarity is compared with a second threshold value, and if the similarity is greater than the second threshold value, the target user is allowed to perform corresponding operations based on a first execution authority, and a prompt for collecting a third feature is sent to the target user; when the third feature matches the pre-stored target user feature data, the target user is allowed to perform corresponding operations based on a second authority;
[0010] When the data quantity of at least one of the fingerprint feature sequence, the voiceprint feature sequence, and the dynamic feature sequence does not satisfy a pre-designed calculation condition, the fingerprint feature sequence, the voiceprint feature sequence, and the dynamic feature sequence are combined, similarity calculation is performed on the combined features and pre-stored target user feature data, and when the similarity is greater than a fourth threshold value, the target user is allowed to perform corresponding operations based on a second authority.
[0011] In some implementations of the first aspect, the similarity calculation on the combined features and the pre-stored target user feature data includes:
[0012] The target region in the fingerprint feature sequence is extracted multiple times, and the multiple results are averaged to obtain a to-be-recognized fingerprint feature sequence a;
[0013] A GMM Gaussian mixture model is trained based on the voiceprint feature sequence, and a feature parameter sequence of the GMM Gaussian mixture model is obtained as a to-be-recognized voiceprint feature sequence b;
[0014] The dynamic feature sequence is equally divided by different block methods, and a plurality of sub-images of the same size are obtained under each block method. Each sub-image obtained under the same size block method is subjected to spatial structure processing to obtain a gradient value of each sub-image and a pixel mean value of each sub-image under the same size block method. A local differential binary is used to obtain a binary sequence of the to-be-processed image under the same block method according to the gradient value and the pixel mean value of each sub-image under the same size block method. The binary sequences of the to-be-processed images under different block methods are arranged in the same order to obtain a feature sequence c of a to-be-recognized image;
[0015] The to-be-recognized set {to-be-recognized fingerprint feature sequence a, to-be-recognized voiceprint feature sequence b, and feature sequence c of the to-be-recognized image} is composed of a, b, and c;
[0016] The to-be-recognized set {to-be-recognized fingerprint feature sequence a, to-be-recognized voiceprint feature sequence b, and feature sequence c of the to-be-recognized image} is compared with pre-stored target user feature data for similarity calculation.
[0017] In some implementations of the first aspect, the method further includes:
[0018] When the similarity between the fingerprint feature sequence and the pre-stored target user feature data is lower than a first threshold value, the similarity threshold values corresponding to the voiceprint feature sequence and the dynamic feature sequence are increased to a third threshold value, and in a case where the similarity between the voiceprint feature sequence and the dynamic feature sequence and the pre-stored target user feature data is greater than or equal to the third threshold value, the user is allowed to perform corresponding operations based on a second authority;
[0019] When the similarity between the voiceprint feature sequence and the pre-stored target user feature data is lower than the first threshold value, the similarity threshold values corresponding to the fingerprint feature sequence and the dynamic feature sequence are increased to a third threshold value, and in a case where the similarity between the fingerprint feature sequence and the dynamic feature sequence and the pre-stored target user feature data is greater than or equal to the third threshold value, the user is allowed to perform corresponding operations based on the second authority.
[0020] When the similarity between the dynamic feature sequence and the pre-stored target user feature data is lower than the first threshold value, the similarity threshold values corresponding to the fingerprint feature sequence and the voiceprint feature sequence are increased to a third threshold value, and in a case where the similarity between the fingerprint feature sequence and the voiceprint feature sequence and the pre-stored target user feature data is greater than or equal to the third threshold value, the user is allowed to perform corresponding operations based on the second authority.
[0021] In some implementations of the first aspect, the fingerprint data of the target user is a plurality of frames of fingerprint images continuously collected, wherein each frame of fingerprint image contains a plurality of fingerprint pixels;
[0022] The similarity between the fingerprint feature sequence and the pre-stored target user feature data is calculated respectively, including:
[0023] The offset vector of each fingerprint pixel in the positions of adjacent two frames of fingerprint images is detected, and an offset matrix is formed according to the offset vector; the offset matrix is normalized to obtain a feature matrix, which is compared with the pre-stored target user feature data to determine the similarity between the fingerprint feature sequence and the pre-stored target user feature data.
[0024] In some implementations of the first aspect, the voiceprint data of the target user is generated by the user based on a random defined character set a={a1, a2, a3…, an} displayed by the interpersonal interaction device, and the pre-stored target user feature data stores feature data corresponding to the character set a;
[0025] The similarity between the voiceprint feature sequence and the pre-stored target user feature data is calculated respectively, including:
[0026] A plurality of sound signals are extracted from the pre-stored target user feature data, wherein the sound signals are sound signals of the target user corresponding to the displayed random defined character set;
[0027] A plurality of phonemes contained in the plurality of sound signals of the target user are composed into a standard phoneme sequence;
[0028] The target phoneme sequence is recognized in the voiceprint data by using a pre-set speech recognition algorithm, and is compared with the standard phoneme sequence to determine the similarity between the voiceprint feature sequence and the pre-stored target user feature data.
[0029] In some implementations of the first aspect, the dynamic feature sequence includes a plurality of image features of the target user in succession.
[0030] The similarity calculation is performed on the dynamic feature sequence and the pre-stored target user feature data respectively, including:
[0031] The preset recognition model is used to perform feature recognition on the human skeleton features, the body shape features and the facial features in the plurality of image features, and similarity calculation is performed on the recognition results and the pre-stored target user feature data respectively to determine the similarity.
[0032] In some implementations of the first aspect, after allowing the target user to perform the corresponding operation based on the second permission, the method further includes:
[0033] The management personnel corresponding to the human-computer interaction device sends an artificial review request to the target user for updating the data of the target user in the database corresponding to the human-computer interaction device.
[0034] In some implementations of the first aspect, the third feature includes at least one of an identification document of the target user and a signature.
[0035] When the third feature matches the pre-stored target user feature data, the target user is allowed to perform the corresponding operation based on the second permission, including:
[0036] When the information corresponding to the identification document of the target user read matches the pre-stored target user feature data, and / or when the generated track and the image corresponding to the signature read match the pre-stored target user feature data, the target user is allowed to perform the corresponding operation based on the second permission.
[0037] The second aspect provides a human-computer interaction device, including:
[0038] The acquisition module is configured to acquire fingerprint data, voiceprint data and dynamic images of a target user.
[0039] The processing module is configured to pre-process the fingerprint data, the voiceprint data and the dynamic images to obtain corresponding fingerprint feature sequences, voiceprint feature sequences and dynamic feature sequences.
[0040] The processing module is further configured to, when the data quantity of each of the sequence of fingerprint features, the sequence of voiceprint features, and the sequence of dynamic features meets a pre-designed calculation condition, respectively perform similarity calculation on the sequence of fingerprint features, the sequence of voiceprint features, and the sequence of dynamic features and the pre-stored target user feature data; when the similarity between any one of the sequence of fingerprint features, the sequence of voiceprint features, and the sequence of dynamic features and the pre-stored target user feature data is lower than a first threshold value, compare the corresponding similarity with a second threshold value, and if the similarity is greater than the second threshold value, allow the target user to perform a corresponding operation based on a first execution authority, and send a prompt for collecting a third feature to the target user; and when the third feature matches the pre-stored target user feature data, allow the target user to perform a corresponding operation based on a second authority.
[0041] The processing module is further configured to, when the data quantity of at least one of the sequence of fingerprint features, the sequence of voiceprint features, and the sequence of dynamic features does not meet the pre-designed calculation condition, combine the sequence of fingerprint features, the sequence of voiceprint features, and the sequence of dynamic features, perform similarity calculation on the combined features and the pre-stored target user feature data, and when the similarity is greater than a fourth threshold value, allow the target user to perform a corresponding operation based on the second authority.
[0042] In a third aspect, a computer storage medium is provided, and the computer storage medium stores computer program instructions. When the computer program instructions are executed by a processor, the method in the first aspect and some implementation manners of the first aspect are implemented.
[0043] The information processing method for human-computer interaction, the human-computer interaction device, and the storage medium provided by the embodiments of the present application can process the collected multi-dimensional data, and when the data quantity of each kind of data meets a pre-designed calculation condition, perform similarity calculation on each kind of data respectively to set the authority; when the data quantity of at least one kind of data does not meet the pre-designed calculation condition, combine the data of multiple dimensions, perform similarity calculation based on the combined data to set the authority, so that the recognition accuracy of the human-computer interaction device can be ensured in the case that the data quantity obtained in a certain dimension does not meet the requirement. BRIEF DESCRIPTION OF DRAWINGS
[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiments of the present application will be briefly introduced. For those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.
[0045] Figure 1 is a flowchart of an information processing method for human-computer interaction provided by the embodiments of the present application;
[0046] Figure 2Fig. 1 is a structural schematic diagram of a human-computer interaction device according to an embodiment of the present application;
[0047] Figure 3 Fig. 2 is a structural diagram of a computing device according to an embodiment of the present application. DETAILED DESCRIPTION
[0048] The features and exemplary embodiments of various aspects of the present application will be described below in detail, in order to make the purposes, technical solutions and advantages of the present application more clear and apparent, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only configured to explain the present application, and are not configured to limit the present application. The present application can be implemented without some of these specific details by those skilled in the art. The following description of the embodiments is only to provide a better understanding of the present application by showing examples of the present application.
[0049] It should be noted that, in this paper, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the elements defined by the statement "include" do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0050] In the existing human-computer interaction process, the authentication of the device to the user is often only authenticated and identified by one-dimensional data. Once the user cannot normally provide the one-dimensional data, it will cause trouble to the user and the service provider, which is not conducive to the business handling.
[0051] In addition, in the existing user authentication and identification scheme using multi-dimensional data, there is no special optimization for the scene where the problem of a certain dimension of data occurs, which causes inaccurate identification.
[0052] The technical solutions of the embodiments of the present application will be described below in combination with the drawings.
[0053] Figure 1 Fig. 1 is a structural schematic diagram of a human-computer interaction device according to an embodiment of the present application; Figure 1 As shown in Fig. 1, the flow of the human-computer interaction information processing method includes:
[0054] S101: Collects fingerprint data, voiceprint data, and dynamic images of the target user.
[0055] The animated image can represent the user's facial data.
[0056] S102: Preprocess the fingerprint data, voiceprint data, and dynamic image to obtain the corresponding fingerprint feature sequence, voiceprint feature sequence, and dynamic feature sequence.
[0057] In some embodiments, fingerprint data and dynamic images are preprocessed, including fingerprint image normalization, image enhancement, image binarization, fingerprint ridge thinning, and fingerprint feature point parameter extraction; voiceprint data is preprocessed, including noise reduction.
[0058] S103: When the data volume of each of the fingerprint feature sequence, voiceprint feature sequence, and dynamic feature sequence meets the preset calculation conditions, the similarity of the fingerprint feature sequence, voiceprint feature sequence, and dynamic feature sequence with the pre-stored target user feature data is calculated respectively; when the similarity between any feature in the fingerprint feature sequence, voiceprint feature sequence, and dynamic feature sequence and the pre-stored target user feature data is lower than the first threshold, the corresponding similarity is compared with the second threshold. If the similarity is greater than the second threshold, the target user is allowed to perform the corresponding operation based on the first execution permission, and a prompt to collect the third feature is sent to the target user; when the third feature matches the pre-stored target user feature data, the target user is allowed to perform the corresponding operation based on the second permission.
[0059] When the amount of data in at least one of the fingerprint feature sequence, voiceprint feature sequence, and dynamic feature sequence does not meet the preset calculation conditions, the fingerprint feature sequence, voiceprint feature sequence, and dynamic feature sequence are combined, and the combined features are compared with the pre-stored target user feature data. When the similarity is greater than the fourth threshold, the target user is allowed to perform the corresponding operation based on the second permission.
[0060] It should be noted that the amount of data here can be understood as the quantity and richness of the feature sequence data.
[0061] Depend on Figure 1 As can be seen from the disclosed human-computer interaction information processing method, this invention can process collected multi-dimensional data. When the amount of data for each data type meets preset calculation conditions, similarity calculation is performed on each data type separately to set permissions. When the amount of data for at least one data type does not meet the preset calculation conditions, data from multiple dimensions is combined, and similarity calculation is performed based on the combined data to set permissions. This ensures the recognition accuracy of the human-computer interaction device even when the amount of data obtained in a certain dimension does not meet the requirements.
[0062] In some embodiments, the fingerprint feature sequence, the voiceprint feature sequence, and the dynamic feature sequence are combined, and the combined features are subjected to similarity calculation with the pre-stored target user feature data, including:
[0063] The target region in the fingerprint feature sequence is extracted multiple times, and the multiple extracted results are averaged to obtain the to-be-identified fingerprint feature sequence a, wherein the target region can be a part with an identifier in the fingerprint feature sequence.
[0064] The GMM Gaussian mixture model is trained based on the voiceprint feature sequence, and the feature parameter sequence of the obtained GMM Gaussian mixture model is taken as the to-be-identified voiceprint feature sequence b.
[0065] The dynamic feature sequence is equally divided by different size block methods, and a plurality of sub-images of the same size are obtained under each block method. Each sub-image obtained under the same size block method is subjected to spatial structure processing to obtain the gradient value of each sub-image and the pixel mean value of each sub-image under the same size block method. According to the gradient value and the pixel mean value of each sub-image obtained under the same size block method, a local difference binary is used to obtain a binary sequence of the to-be-processed image under the same block method. The binary sequences of the to-be-processed images under different block methods are arranged in the same order to obtain the feature sequence c of the to-be-identified image.
[0066] The to-be-identified set {to-be-identified fingerprint feature sequence a, to-be-identified voiceprint feature sequence b, and feature sequence c of the to-be-identified image} is composed of a, b, and c.
[0067] The to-be-identified set {to-be-identified fingerprint feature sequence a, to-be-identified voiceprint feature sequence b, and feature sequence c of the to-be-identified image} is subjected to similarity calculation with the pre-stored target user feature data.
[0068] Through this embodiment, when the data amount of at least one of the fingerprint feature sequence, the voiceprint feature sequence, and the dynamic feature sequence does not meet the pre-designed calculation condition, the data of each kind is processed first, and then the processed data is combined and subjected to similarity calculation with the pre-stored target user feature data, so that similarity calculation can be normally performed even when the data amount is insufficient.
[0069] In some embodiments, the target region in the fingerprint feature sequence is extracted multiple times, and the multiple extracted results are averaged to obtain the to-be-identified fingerprint feature sequence a, which can be the average of the result sequences after multiple principal component analyses of the same fingerprint, denoted as X2, and the principal component analysis can use PCA operation.
[0070] In some embodiments, the GMM Gaussian Mixture Model is trained based on the voiceprint feature sequence, and a feature parameter sequence of the GMM Gaussian Mixture Model is obtained as the to-be-identified voiceprint feature sequence b. Specifically, the voiceprint feature sequence can be processed into a sequence of voiceprint feature vectors to solve the GMM Gaussian Mixture Model Y1, that is, to obtain the feature parameter sequence λ({μ i ,∑ i ,w i ,M},1≤μ≤M) of the Gaussian Mixture Model, so that the likelihood probability of the feature vector sequence is maximum, where M is the number of mixtures of the GMM Gaussian model, μ i is the mean of the multivariate Gaussian distribution function, ∑ i is the covariance matrix, and w i is the weight distribution of the M multivariate Gaussian models.
[0071] In some embodiments, in order to still be able to identify when the similarity between a certain feature sequence and the pre-stored target user feature data is lower than a certain value, the similarity threshold of other dimensions of data can be increased to perform strong identification based on other types of feature sequences, and then identify the user. Specifically, when the similarity between the fingerprint feature sequence and the pre-stored target user feature data is lower than the first threshold, the similarity threshold corresponding to the voiceprint feature sequence and the dynamic feature sequence is increased to the third threshold, and in the case where the similarity between the voiceprint feature sequence and the dynamic feature sequence and the pre-stored target user feature data is greater than or equal to the third threshold, the user is allowed to perform corresponding operations based on the second authority.
[0072] When the similarity between the voiceprint feature sequence and the pre-stored target user feature data is lower than the first threshold, the similarity threshold corresponding to the fingerprint feature sequence and the dynamic feature sequence is increased to the third threshold, and in the case where the similarity between the fingerprint feature sequence and the dynamic feature sequence and the pre-stored target user feature data is greater than or equal to the third threshold, the user is allowed to perform corresponding operations based on the second authority.
[0073] When the similarity between the dynamic feature sequence and the pre-stored target user feature data is lower than the first threshold, the similarity threshold corresponding to the fingerprint feature sequence and the voiceprint feature sequence is increased to the third threshold, and in the case where the similarity between the fingerprint feature sequence and the voiceprint feature sequence and the pre-stored target user feature data is greater than or equal to the third threshold, the user is allowed to perform corresponding operations based on the second authority.
[0074] In some embodiments, the fingerprint data of the target user is a plurality of frames of continuously collected fingerprint images, and each frame of fingerprint image contains a plurality of fingerprint pixels.
[0075] The similarity between the fingerprint feature sequence and the pre-stored target user feature data is calculated respectively, including:
[0076] detect an offset vector of each fingerprint pixel in the position within two adjacent frames of the fingerprint image, and form an offset matrix according to the offset vector; normalize the offset matrix to obtain a feature matrix, and compare the feature matrix with the pre-stored target user feature data to determine the similarity between the fingerprint feature sequence and the pre-stored target user feature data.
[0077] In some embodiments, the voiceprint data of the target user is generated by the user based on a random defined character set a={a1, a2, a3…, an} displayed by the human-computer interaction device, and the pre-stored target user feature data stores feature data corresponding to the character set a.
[0078] The similarity between the voiceprint feature sequence and the pre-stored target user feature data is calculated respectively, including:
[0079] Extract a plurality of sound signals from the pre-stored target user feature data, wherein the sound signals are sound signals of the target user corresponding to the displayed random defined character set;
[0080] A plurality of phonemes contained in the plurality of sound signals of the target user form a standard phoneme sequence;
[0081] Using a pre-set speech recognition algorithm, the target phoneme sequence is recognized in the voiceprint data, and compared with the standard phoneme sequence to determine the similarity between the voiceprint feature sequence and the pre-stored target user feature data.
[0082] In some embodiments, the dynamic feature sequence includes a plurality of continuous image features of the target user;
[0083] The similarity between the dynamic feature sequence and the pre-stored target user feature data is calculated respectively, including:
[0084] Using a pre-set recognition model, the human body skeleton feature, the body shape feature and the face feature in the plurality of image features are recognized, and the recognition results are compared with the pre-stored target user feature data respectively to determine the similarity.
[0085] In some embodiments, after allowing the target user to perform the corresponding operation based on the second permission, the method further includes:
[0086] The management personnel corresponding to the human-computer interaction device sends a manual review request to the target user for updating the data of the target user in the database corresponding to the human-computer interaction device.
[0087] In some embodiments, the third feature includes at least one of an identification card of the target user and a signature.
[0088] When the third feature matches pre-stored target user feature data, the target user is allowed to perform corresponding operations based on the second permission, including:
[0089] When the information corresponding to the target user's identification document matches the pre-stored target user feature data, and / or when the generated trajectory and image corresponding to the read signature match the pre-stored target user feature data, the target user is allowed to perform the corresponding operation based on the second permission.
[0090] In some embodiments, the method may further include:
[0091] Collect fingerprint data, voiceprint data, and dynamic images of new users;
[0092] Feature extraction and authentication are performed based on new users' fingerprint data, voiceprint data, and dynamic images.
[0093] In some embodiments, in order to provide certain specialized services, the method may further include:
[0094] Collect request information input by the target user;
[0095] The corresponding service is executed based on the request information, wherein the service includes at least one of payment service, service processing and printing service.
[0096] The technical solution of this invention can process collected multi-dimensional data, and when the amount of data for each data type meets preset calculation conditions, perform similarity calculation on each data type separately to set permissions; when the amount of data for at least one data type does not meet the preset calculation conditions, combine data from multiple dimensions, and perform similarity calculation based on the combined data to set permissions. This ensures the recognition accuracy of the human-computer interaction device even when the amount of data obtained in a certain dimension does not meet the requirements.
[0097] and Figure 1 Corresponding to the flowchart of the information processing method for human-computer interaction shown, this invention also discloses a human-computer interaction device, such as... Figure 2 As shown, the human-computer interaction device may include:
[0098] The acquisition module 201 can be used to acquire fingerprint data, voiceprint data and dynamic images of the target user;
[0099] Processing module 202 can be used to preprocess fingerprint data, voiceprint data and dynamic images to obtain corresponding fingerprint feature sequences, voiceprint feature sequences and dynamic feature sequences;
[0100] The processing module 202 can also be configured to perform similarity calculation on the fingerprint feature sequence, the voiceprint feature sequence, and the dynamic feature sequence and the pre-stored target user feature data when the data quantity of each of the fingerprint feature sequence, the voiceprint feature sequence, and the dynamic feature sequence meets a pre-designed calculation condition; when the similarity between any one of the fingerprint feature sequence, the voiceprint feature sequence, and the dynamic feature sequence and the pre-stored target user feature data is lower than a first threshold value, comparing the corresponding similarity with a second threshold value, if the similarity is greater than the second threshold value, allowing the target user to perform a corresponding operation based on a first execution authority, and sending a prompt to collect a third feature to the target user; when the third feature matches the pre-stored target user feature data, allowing the target user to perform a corresponding operation based on a second authority.
[0101] The processing module 202 can also be configured to combine the fingerprint feature sequence, the voiceprint feature sequence, and the dynamic feature sequence when the data quantity of at least one of the fingerprint feature sequence, the voiceprint feature sequence, and the dynamic feature sequence does not meet a pre-designed calculation condition, and perform similarity calculation on the combined feature and the pre-stored target user feature data, and when the similarity is greater than a fourth threshold value, allowing the target user to perform a corresponding operation based on a second authority.
[0102] In some embodiments, the processing module 202 can also be configured to extract a target area in the fingerprint feature sequence multiple times, take the average of the multiple extraction results to obtain a to-be-recognized fingerprint feature sequence a, train a GMM Gaussian mixture model based on the voiceprint feature sequence, and obtain a feature parameter sequence of the GMM Gaussian mixture model as a to-be-recognized voiceprint feature sequence b, divide the dynamic feature sequence by different size block methods, obtain multiple sub-images of the same size under each block method, perform spatial structure processing on each sub-image obtained under the same size block method to obtain the gradient value of each sub-image and the pixel mean value of each sub-image under the same size block method, obtain a binary sequence of the to-be-processed image under the same block method using local differential binary according to the gradient value and the pixel mean value of each sub-image under the same size block method, arrange the binary sequences of the to-be-processed images under different block methods in the same order to obtain a feature sequence c of a to-be-recognized image, and based on a, b, and c, form a to-be-recognized set {to-be-recognized fingerprint feature sequence a, to-be-recognized voiceprint feature sequence b, feature sequence c of the to-be-recognized image}, and perform similarity calculation on the to-be-recognized set {to-be-recognized fingerprint feature sequence a, to-be-recognized voiceprint feature sequence b, feature sequence c of the to-be-recognized image} and the pre-stored target user feature data.
[0103] In some embodiments, the processing module 202 can also be configured to, when the similarity between the fingerprint feature sequence and the pre-stored target user feature data is lower than the first threshold, increase the similarity threshold corresponding to the voiceprint feature sequence and the dynamic feature sequence to a third threshold, and allow the user to perform corresponding operations based on the second authority when the similarity between the voiceprint feature sequence and the dynamic feature sequence and the pre-stored target user feature data is greater than or equal to the third threshold; when the similarity between the fingerprint feature sequence and the pre-stored target user feature data is lower than the first threshold, increase the similarity threshold corresponding to the fingerprint feature sequence and the dynamic feature sequence to the third threshold, and allow the user to perform corresponding operations based on the second authority when the similarity between the fingerprint feature sequence and the dynamic feature sequence and the pre-stored target user feature data is greater than or equal to the third threshold; and when the similarity between the dynamic feature sequence and the pre-stored target user feature data is lower than the first threshold, increase the similarity threshold corresponding to the fingerprint feature sequence and the voiceprint feature sequence to the third threshold, and allow the user to perform corresponding operations based on the second authority when the similarity between the fingerprint feature sequence and the voiceprint feature sequence and the pre-stored target user feature data is greater than or equal to the third threshold.
[0104] In some embodiments, the fingerprint data of the target user is a plurality of frames of continuously collected fingerprint images, wherein each frame of fingerprint image contains a plurality of fingerprint pixels;
[0105] The processing module 202 can also be configured to detect an offset vector of each fingerprint pixel in the positions of two adjacent frames of fingerprint images, and construct an offset matrix according to the offset vector; normalize the offset matrix to obtain a feature matrix, and compare the feature matrix with the pre-stored target user feature data to determine the similarity between the fingerprint feature sequence and the pre-stored target user feature data.
[0106] In some embodiments, the voiceprint data of the target user is generated by the user based on a random defined character set a={a1, a2, a3, …, an} displayed by the interpersonal interaction device, and the pre-stored target user feature data stores feature data corresponding to the character set a;
[0107] The processing module 202 can also be configured to extract a plurality of sound signals from the pre-stored target user feature data, wherein the sound signals are sound signals of the target user corresponding to the displayed random defined character set; group a plurality of phonemes contained in the plurality of sound signals of the target user to form a standard phoneme sequence; use a pre-set speech recognition algorithm to identify a target phoneme sequence in the voiceprint data, and compare the target phoneme sequence with the standard phoneme sequence to determine the similarity between the voiceprint feature sequence and the pre-stored target user feature data.
[0108] In some embodiments, the dynamic feature sequence includes a plurality of image features of the target user in succession;
[0109] The processing module 202 can also be configured to perform feature recognition on the human skeleton features, body shape features and face features in the multi-frame image features using a preset recognition model, perform similarity calculation on the recognition result and the pre-stored target user feature data respectively, and determine the similarity.
[0110] The technical solution of the present application can process the collected multi-dimensional data, and when the data quantity of each type of data meets the preset calculation condition, perform similarity calculation on each type of data respectively to set the permission; when the data quantity of at least one type of data does not meet the preset calculation condition, combine the multi-dimensional data, and perform similarity calculation based on the combined data to set the permission. Thus, the recognition accuracy of the human-computer interaction device can be ensured when the data quantity obtained in a certain dimension does not meet the requirement.
[0111] It should be further noted that, Figure 2 Each module in the device shown in the above Figure 1 The method shown in the above corresponds to the function of each step, and will not be described here.
[0112] Figure 3 is a structural diagram of a computing device provided by an embodiment of the present application. As shown in the above, Figure 3 The computing device 300 includes an input interface 301, a central processor 302, a memory 303 and an output interface 304. The input interface 301, the central processor 302, the memory 303 and the output interface 304 are connected to each other through a bus 310.
[0113] Figure 3 The computing device shown in the above can also be implemented as an execution device of the information processing method of human-computer interaction, and the computing device can include a processor and a memory storing computer executable instructions; the processor can implement the information processing method of human-computer interaction provided by an embodiment of the present application when executing the computer executable instructions.
[0114] An embodiment of the present application further provides a computer readable storage medium, and the computer readable storage medium stores computer program instructions; the computer program instructions are executed by the processor to implement the information processing method of human-computer interaction provided by an embodiment of the present application.
[0115] It should be clear that the present application is not limited to the specific configurations and processes described above and shown in the drawings. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order of the steps, after understanding the spirit of the present application.
[0116] The functional blocks shown in the structural block diagrams described above can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, and the like. When implemented in software, the elements of the present application are program or code segments that are used to perform the required tasks. The program or code segments can be stored in a machine-readable medium, or transmitted through a data signal carried in a carrier wave over a transmission medium or communication link. A "machine-readable medium" includes any medium that can store or transport the information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, read-only memories (ROMs), flash memories, erasable read-only memories (EROMs), floppy disks, compact discs read-only memories (CD-ROMs), optical discs, hard disks, fiber optic media, radio frequency (RF) links, and the like. The code segments can be downloaded via computer networks such as the Internet, intranets, and the like.
[0117] It is also important to note that the examples described herein can be implemented in a variety of systems, including and not limited to a digital electronic circuit, an analog electronic circuit, a computer hardware, firmware, software, or in combinations of them. The example embodiments described herein can be implemented as one or more computer programs or code (e.g., applications) operable to be executed on a computer-based system and / or a processor-based system. As such, embodiments of the present application can also be implemented as program codes containing one or more instructions executable by a processor-based system (e.g., a computer system) to perform the steps described herein.
[0118] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. Alternatively, computer program implemented steps can be implemented by special purpose logic circuitry, e.g., an FPGA or an ASIC, to perform the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0119] The above merely describes specific implementation of the present application, and those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, module and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be described herein again. It should be understood that the protection scope of the present application is not limited to this, and any person skilled in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements should be covered within the protection scope of the present application.
Claims
1. An information processing method of human-computer interaction, applied to a human-computer interaction device, characterized in that, The method comprises: Collecting fingerprint data, voiceprint data and dynamic images of a target user; Preprocessing the fingerprint data, voiceprint data and dynamic images to obtain corresponding fingerprint feature sequences, voiceprint feature sequences and dynamic feature sequences; When the data quantity of each of the fingerprint feature sequences, voiceprint feature sequences and dynamic feature sequences meets a preset calculation condition, performing similarity calculation on the fingerprint feature sequences, voiceprint feature sequences and dynamic feature sequences and pre-stored target user feature data respectively; when the similarity between any of the fingerprint feature sequences, voiceprint feature sequences and dynamic feature sequences and the pre-stored target user feature data is lower than a first threshold value, comparing the corresponding similarity with a second threshold value, if the similarity is greater than the second threshold value, allowing the target user to perform corresponding operations based on a first execution authority, and sending a prompt to the target user to collect a third feature; when the third feature matches the pre-stored target user feature data, allowing the target user to perform corresponding operations based on a second authority, When the data quantity of at least one of the fingerprint feature sequences, voiceprint feature sequences and dynamic feature sequences does not meet the preset calculation condition, combining the fingerprint feature sequences, voiceprint feature sequences and dynamic feature sequences, and performing similarity calculation on the combined features and the pre-stored target user feature data, when the similarity is greater than a fourth threshold value, allowing the target user to perform corresponding operations based on the second authority.
2. The method of claim 1, wherein, The combination of the fingerprint feature sequences, voiceprint feature sequences and dynamic feature sequences, and the similarity calculation on the combined features and the pre-stored target user feature data, comprises: Extracting a target area in the fingerprint feature sequence multiple times, and taking the average of the multiple extracted results to obtain a to-be-recognized fingerprint feature sequence a; Training a GMM Gaussian mixture model based on the voiceprint feature sequence, and taking the feature parameter sequence of the GMM Gaussian mixture model as a to-be-recognized voiceprint feature sequence b; Dividing the dynamic feature sequence into multiple sub-images of the same size through different block methods, respectively processing each sub-image of the same size under the same block method to obtain the gradient value of each sub-image and the pixel mean value of each sub-image, and obtaining the binary sequence of the to-be-processed image under the same block method by using local differential binary according to the gradient value and the pixel mean value of each sub-image under the same block method; and arranging the binary sequences of the to-be-processed images under different block methods in the same order to obtain a to-be-recognized image feature sequence c; Based on the to-be-recognized fingerprint feature sequence a, the to-be-recognized voiceprint feature sequence b and the to-be-recognized image feature sequence c, a to-be-recognized set {to-be-recognized fingerprint feature sequence a, to-be-recognized voiceprint feature sequence b, to-be-recognized image feature sequence c} is formed; Performing similarity calculation on the to-be-recognized set {to-be-recognized fingerprint feature sequence a, to-be-recognized voiceprint feature sequence b, to-be-recognized image feature sequence c} and the pre-stored target user feature data.
3. The method of claim 1, wherein, The method further comprises: when the similarity between the fingerprint feature sequence and the pre-stored target user feature data is lower than the first threshold, increasing the similarity threshold corresponding to the voiceprint feature sequence and the dynamic feature sequence to a third threshold, and allowing the user to perform corresponding operations based on the second authority when the similarity between the voiceprint feature sequence and the dynamic feature sequence and the pre-stored target user feature data is greater than or equal to the third threshold; when the similarity between the voiceprint feature sequence and the pre-stored target user feature data is lower than the first threshold, increasing the similarity threshold corresponding to the fingerprint feature sequence and the dynamic feature sequence to a third threshold, and allowing the user to perform corresponding operations based on the second authority when the similarity between the fingerprint feature sequence and the dynamic feature sequence and the pre-stored target user feature data is greater than or equal to the third threshold; when the similarity between the dynamic feature sequence and the pre-stored target user feature data is lower than the first threshold, increasing the similarity threshold corresponding to the fingerprint feature sequence and the voiceprint feature sequence to a third threshold, and allowing the user to perform corresponding operations based on the second authority when the similarity between the fingerprint feature sequence and the voiceprint feature sequence and the pre-stored target user feature data is greater than or equal to the third threshold.
4. The method of claim 1, wherein, The fingerprint data of the target user is a plurality of frames of fingerprint images collected continuously, wherein each frame of the fingerprint image contains a plurality of fingerprint pixels; The similarity calculation of the fingerprint feature sequence and the pre-stored target user feature data respectively comprises: detecting the offset vector of each fingerprint pixel in the position of adjacent two frames of fingerprint images, and constructing an offset matrix according to the offset vector; normalizing the offset matrix to obtain a feature matrix, and comparing the feature matrix with the pre-stored target user feature data to determine the similarity between the fingerprint feature sequence and the pre-stored target user feature data.
5. The method of claim 1, wherein, The voiceprint data of the target user is generated by the user based on a random defined character set a={a1, a2, a3…, an} displayed by the interpersonal interaction device, and the pre-stored target user feature data stores feature data corresponding to the character set a; The similarity calculation of the voiceprint feature sequence and the pre-stored target user feature data respectively comprises: extracting a plurality of sound signals from the pre-stored target user feature data, wherein the sound signals are sound signals of the target user corresponding to the displayed random defined character set; composing a standard phoneme sequence from a plurality of phonemes contained in a plurality of target user sound signals; using a pre-set speech recognition algorithm to identify a target phoneme sequence in the voiceprint data, and comparing the target phoneme sequence with the standard phoneme sequence to determine the similarity between the voiceprint feature sequence and the pre-stored target user feature data.
6. The method of claim 1, wherein, The dynamic feature sequence comprises a plurality of image features of the target user continuously; The similarity calculation of the dynamic feature sequence and the pre-stored target user feature data respectively comprises: The preset recognition model is used to recognize human skeleton features, body shape features and face features in the multi-frame image features, similarity calculation is performed on the recognition results and the pre-stored target user feature data respectively, and similarity is determined.
7. The method of claim 1, wherein, After allowing the target user to perform corresponding operations based on the second permission, the method further includes: The management personnel corresponding to the human-computer interaction device sends an artificial review request to the target user for updating data of the target user in a database corresponding to the human-computer interaction device.
8. The method of claim 1, wherein, The third feature includes at least one of a certificate of the target user and a signature. When the third feature matches the pre-stored target user feature data, the target user is allowed to perform corresponding operations based on the second permission, including: When the information corresponding to the certificate of the target user read matches the pre-stored target user feature data, and / or when the generated track and image corresponding to the signature read match the pre-stored target user feature data, the target user is allowed to perform corresponding operations based on the second permission.
9. A human-machine interaction device, characterized in that The device includes: The acquisition module is configured to acquire fingerprint data, voiceprint data and dynamic images of the target user. The processing module is configured to pre-process the fingerprint data, voiceprint data and dynamic images to obtain corresponding fingerprint feature sequences, voiceprint feature sequences and dynamic feature sequences. The processing module is further configured to perform similarity calculation on the fingerprint feature sequences, voiceprint feature sequences and dynamic feature sequences and the pre-stored target user feature data respectively when the data amount of each sequence in the fingerprint feature sequences, voiceprint feature sequences and dynamic feature sequences meets a preset calculation condition; when the similarity between any feature in the fingerprint feature sequences, voiceprint feature sequences and dynamic feature sequences and the pre-stored target user feature data is lower than a first threshold value, the corresponding similarity is compared with a second threshold value, if the similarity is greater than the second threshold value, the target user is allowed to perform corresponding operations based on a first execution permission, and a prompt for acquiring a third feature is sent to the target user; when the third feature matches the pre-stored target user feature data, the target user is allowed to perform corresponding operations based on a second permission. The processing module is further configured to combine the fingerprint feature sequences, voiceprint feature sequences and dynamic feature sequences when the data amount of at least one sequence in the fingerprint feature sequences, voiceprint feature sequences and dynamic feature sequences does not meet the preset calculation condition, perform similarity calculation on the combined features and the pre-stored target user feature data, and allow the target user to perform corresponding operations based on the second permission when the similarity is greater than a fourth threshold value.
10. A computer storage medium, characterized in that The computer storage medium stores computer program instructions, and the computer program instructions are executed by a processor to implement the method in any one of claims 1-8.
Citation Information
Patent Citations
Combined authentication method and intelligent interactive system
CN105224850A
Multidimensional user identity identification method
CN106599866A