User identity authenticity verification method and system based on multi-modal AI

Through the comprehensive analysis of multimodal AI technology, the problems of misjudgment, privacy leakage and high device performance requirements in existing identity verification technologies are solved, real-time, lightweight and scalable user identity authenticity verification are achieved, and the accuracy and security of verification are improved.

CN120068037APending Publication Date: 2025-05-30SHANGHAI CHANGZHI CULTURE COMM CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411901372.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing authentication technologies have problems with misjudgment, privacy leaks and high device performance requirements, especially when faced with deep forgery and large-scale data storage.

Method used

Multimodal AI technology is adopted to achieve real-time, lightweight and scalable user identity authenticity verification through comprehensive analysis of video, voice, action and environment. Specific steps include video acquisition and preprocessing, facial feature points and timing analysis, speech feature analysis, background environment analysis, and user identity authenticity scores are generated through weighting algorithms.

Benefits of technology

Improve the accuracy and security of identity verification, avoid privacy leakage, reduce requirements for device performance, enhance user trust, and comply with data protection regulations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068037A_ABST
    Figure CN120068037A_ABST
Patent Text Reader

Abstract

The invention provides a multi-mode AI-based user identity authenticity verification method and system, and relates to the technical field of identity verification, and the method comprises the steps: obtaining a dynamic video of a user through a video collection module, and carrying out the preprocessing of the dynamic video, so as to obtain a preprocessed dynamic video; according to the preprocessed dynamic video, a dynamic analysis module is used for analyzing face feature points and a time sequence to obtain a dynamic analysis result; according to the preprocessed dynamic video, performing voice feature analysis by using a voice analysis module to obtain a voice analysis result; performing background environment analysis according to the preprocessed dynamic video to obtain an environment authenticity analysis result; according to the dynamic analysis result, the voice analysis result and the environment authenticity analysis result, a verification scoring module performs scoring by adopting a weighting algorithm so as to obtain a user identity authenticity score; and according to the user identity authenticity score, the background management system is used for storage, and a user identity authenticity verification report is output. The technical problems of misjudgment and privacy disclosure are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of identity authentication, and particularly to a method and system for authenticating the authenticity of a user's identity based on multi-modal AI. Background Art

[0002] At present, in the market, there is a trust crisis caused by a large difference between a photo and a real person. Conventional real-name authentication technical solutions usually rely on photos or videos uploaded by users for authenticity verification, and introduce face recognition technology or video dynamic verification technology to identify the authenticity of the content uploaded by users.

[0003] The disadvantages of face recognition technology are limited to the verification of photos or videos uploaded by users, and it is impossible to avoid the misappropriation of Deepfake videos and static photos, and it is impossible to prove that the current user is the real operator, and there is a possibility of misjudgment and bypassing verification.

[0004] The disadvantages of dynamic verification technology are that the existing technology is mainly based on fixed instructions or simple dynamic behaviors, such as "blinking" or "smiling", which are easily cracked by pre-recorded forged videos, and fail to effectively balance the needs of user authenticity verification and privacy protection.

[0005] Most existing verification systems will save users' biometric data, such as face recognition data or fingerprint data, which may lead to the risk of user privacy leakage. Especially when facing large-scale data storage, strict compliance and data security issues need to be considered. In addition, the existing verification process often relies on complex hardware and system requirements, and has high performance requirements for terminal devices, which may lead to poor user experience.

[0006] The present invention realizes a real-time, lightweight, and scalable user authenticity verification system through multi-modal AI analysis (video, voice, action, environment). Summary of the Invention

[0007] The present invention provides a method and system for authenticating the authenticity of a user's identity based on multi-modal AI, which solves the technical problems of misjudgment and privacy leakage.

[0008] To solve the above technical problems, the technical solution of the present invention is as follows: In a first aspect, a method for authenticating the authenticity of a user's identity based on multi-modal AI, the method includes: According to a voice command, use a video acquisition module to acquire a dynamic video of the user and perform preprocessing to obtain a preprocessed dynamic video; According to the preprocessed dynamic video, use a dynamic analysis module to analyze face feature points and time series to obtain a dynamic analysis result; Based on the pre - processed dynamic video, use the speech analysis module to perform speech feature analysis to obtain the speech analysis result; Based on the pre - processed dynamic video, perform background environment analysis to obtain the environmental authenticity analysis result; Based on the dynamic analysis result, speech analysis result and environmental authenticity analysis result, the verification scoring module uses a weighted algorithm for scoring to obtain the user identity authenticity score; Based on the user identity authenticity score, use the background management system for storage and output the user identity authenticity verification report.

[0009] Furthermore, use the video capture module to obtain the user's dynamic video and perform pre - processing to obtain the pre - processed dynamic video, including: Use the camera to capture the user's dynamic video in real - time. The dynamic video includes the actions, expressions and environment when the user completes various instructions; In the dynamic video, use the face detection algorithm to identify the face position of the user and perform alignment processing on the detected face; In the dynamic video, use the denoising algorithm to reduce the noise and interference in the video to obtain the denoised dynamic video; In the dynamic video, use video enhancement technology for correction to obtain the corrected dynamic video, that is, the pre - processed dynamic video.

[0010] Furthermore, based on the pre - processed dynamic video, use the dynamic analysis module to perform analysis of face feature points and timing to obtain the dynamic analysis result, including: Based on the pre - processed dynamic video, the dynamic analysis module uses OpenCV for face detection to obtain the coordinates of the face area; Based on the coordinates of the face area, extract the face feature points. The face feature points include key points such as the corners of the eyes, the corners of the mouth, and the nose; Based on the face feature points, calculate the behavioral characteristics of the face in the dynamic behavior to obtain the behavioral characteristics. The behavioral characteristics include the degree of eye opening and closing, the change of the corners of the mouth, and the rotation angle of the head; Based on the behavioral characteristics, identify the dynamic behavior in the video; Perform timing analysis on the identified dynamic behavior to obtain the dynamic behavior timing analysis result; Based on the dynamic behavior and timing analysis result, by comparing the changes in behavioral characteristics in consecutive frames, obtain the dynamic analysis result.

[0011] Furthermore, based on the pre - processed dynamic video, use the speech analysis module to perform speech feature analysis to obtain the speech analysis result, including: Based on the pre - processed dynamic video, separate the audio data and perform pre - processing to obtain the pre - processed audio data; Based on the pre - processed audio data, analyze the frequency and time - domain characteristics of the sound, identify each phoneme in the audio, and determine the pronunciation characteristics of each phoneme; Based on the pronunciation characteristics of the phonemes, analyze the emotional state of the speaker in the audio and the tone of the sentence to obtain intonation characteristics, where the intonation characteristics include the rhythm of the sentence, pitch change, vowel length, and intensity; Based on the pre - processed audio data, convert the audio data into a spectrogram and perform frequency component and amplitude distribution analysis to obtain the audio spectrum analysis result; Align and analyze the dynamic behavior time - series analysis result with the time - domain characteristics of the audio data to obtain the synchronization analysis result; Based on the intonation characteristics, audio spectrum result, and synchronization analysis result, generate a result report, i.e., the speech analysis result.

[0012] Furthermore, based on the pre - processed dynamic video, perform background environment analysis to obtain the environmental authenticity analysis result, including: Based on the pre - processed dynamic video, split the video into a series of individual image frames and perform enhancement processing on the image frames to obtain the enhanced image frames; Based on the enhanced image frames, use an object detection algorithm to identify the image frames to obtain the object detection result; Based on the object detection result, extract the background region in the image frame and perform feature analysis on the extracted background region to obtain the feature analysis result; Based on the feature analysis result, compare it with the preset real - background feature library to obtain the environmental authenticity analysis result.

[0013] Furthermore, based on the dynamic analysis result, speech analysis result, and environmental authenticity analysis result, the verification scoring module uses a weighted algorithm for scoring to obtain the user identity authenticity score, including: Based on the dynamic analysis result, speech analysis result, and environmental authenticity analysis result, perform normalization processing to obtain the normalized analysis result; Based on the normalized analysis result, determine the weights of the dynamic analysis result, speech analysis result, and environmental authenticity analysis result; Based on the weights of the dynamic analysis result, speech analysis result, and environmental authenticity analysis result, use to perform scoring to obtain the user identity authenticity score, where is the final user identity authenticity score, is the dynamic analysis result, is the speech analysis result, is the result of environmental authenticity analysis, is the weight of the dynamic analysis result, is the weight of the voice analysis result, is the weight of the environmental authenticity analysis result.

[0014] Furthermore, according to the user identity authenticity score, use the back-end management system for storage and output a user identity authenticity verification report, including: Construct a database for storing the user identity authenticity score and encrypt the database to obtain an encrypted database; Store the user identity authenticity score in the encrypted database and record the timing information; Generate and output a user identity authenticity verification report based on the user identity authenticity score; Conduct data analysis and mining on the stored score results to obtain analysis results; Update and optimize the verification rules according to the analysis results.

[0015] In a second aspect, a user identity authenticity verification system based on multi-modal AI includes: An acquisition module for using a video acquisition module to acquire a user's dynamic video and perform preprocessing to obtain a preprocessed dynamic video; A processing module for analyzing facial feature points and timing using a dynamic analysis module based on the preprocessed dynamic video to obtain a dynamic analysis result, analyzing voice features using a voice analysis module based on the preprocessed dynamic video to obtain a voice analysis result, performing background environment analysis based on the preprocessed dynamic video to obtain an environmental authenticity analysis result, and using a weighted algorithm for scoring by a verification scoring module based on the dynamic analysis result, voice analysis result, and environmental authenticity analysis result to obtain a user identity authenticity score, and storing and outputting a user identity authenticity verification report using the back-end management system according to the user identity authenticity score.

[0016] In a third aspect, a computing device includes: One or more processors; A storage system for storing one or more programs, which when executed by the one or more processors cause the one or more processors to implement the above method.

[0017] In a fourth aspect, a computer-readable storage medium stores a program that, when executed by a processor, implements the above method.

[0018] The above solution of the present invention has at least the following beneficial effects: The above solution of the present invention detects the temporal consistency of actions, voices, and environments through AI to ensure that the content is completed in real time for the user; the present invention adopts a multi-modal verification system, which comprehensively analyzes the multi-dimensional behavior data of the user to ensure that the verification result is more accurate and reliable; through multi-dimensional data fusion, the present invention can make up for the deficiencies of single verification means, improve the accuracy of verification, and prevent various common forgery means. The present invention uses a weighted scoring algorithm to generate the authenticity score of the user without storing the biometric data of the user, and the personal data of the user can be effectively protected, avoiding the risk of privacy leakage. The privacy protection measures ensure that the user data is not stored or misused, greatly enhancing the user's trust, and also helping to comply with various data protection regulations. In the technical implementation, the present invention optimizes the AI model to avoid overly complex calculations and harsh requirements for hardware performance, enabling the system to operate efficiently on ordinary mobile phones or mobile devices, and ensuring a relatively smooth user experience. The calculation process is optimized to make the system operation more efficient and convenient, improving the user's operation experience. Especially on the mobile device side, it can provide fast response and low-latency services. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 FIG. is a schematic flowchart of a method for authenticating the user identity based on multi-modal AI provided by an embodiment of the present invention.

[0020] Figure 2 FIG. is a schematic diagram of a system for authenticating the user identity based on multi-modal AI provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0022] As Figure 1 shown, an embodiment of the present invention proposes a method for authenticating the user identity based on multi-modal AI, and the method includes: 11. According to the voice command, use the video acquisition module to obtain the dynamic video of the user and perform preprocessing to obtain the preprocessed dynamic video; 12. According to the preprocessed dynamic video, use the dynamic analysis module to analyze the facial feature points and time series to obtain the dynamic analysis result; 13. According to the preprocessed dynamic video, use the voice analysis module to perform voice feature analysis to obtain the voice analysis result; 14. Analyze the background environment based on the preprocessed dynamic video to obtain the analysis result of environmental authenticity; 15. According to the dynamic analysis result, voice analysis result and environmental authenticity analysis result, the verification scoring module uses a weighted algorithm for scoring to obtain the user identity authenticity score; 16. Store according to the user identity authenticity score using the background management system and output the user identity authenticity verification report.

[0023] In the embodiment of the present invention, video analysis, voice analysis and environmental background analysis are combined to verify the authenticity of the user identity through multi-dimensional data, improving the accuracy and security of the verification; Step 11 uses the video acquisition module to receive the actions or expressions made by the user according to the voice command to generate the original dynamic video, and preprocess this video, including denoising, enhancing clarity, frame rate adjustment, etc., to obtain the optimized dynamic video. The preprocessing ensures the quality of the video data and provides a clear image basis for subsequent face feature point and timing analysis; Step 12 uses the dynamic analysis module to detect the face feature points of the preprocessed dynamic video, track the changes of these feature points over time, and analyze the facial dynamic features of the user. By analyzing the timing changes of the face feature points, it can be identified whether the user is a real live body, effectively preventing photo or video fraud; Step 13 uses the voice analysis module to extract the features of the voice part in the preprocessed dynamic video, including the analysis of features such as voice frequency, timbre, and speech rate. Voice feature analysis can further confirm the user's identity because everyone's voice features are unique, which increases the reliability of the verification; Step 14 analyzes the background environment in the preprocessed dynamic video, including detecting whether there are abnormal objects in the environment, whether the light conditions are natural, whether the background sounds match, etc. Background environment analysis helps to identify whether the user is verifying in a real environment and prevent fraud using pre-recorded videos or images; Step 15, the verification scoring module uses a weighted algorithm for comprehensive scoring according to the face feature point analysis result, voice analysis result and environmental authenticity analysis result, calculates the user identity authenticity score, and by integrating the analysis results of multiple dimensions, can more accurately evaluate the authenticity of the user identity, improving the accuracy and robustness of the verification; Step 16, the background management system stores the user identity authenticity score and generates a user identity authenticity verification report, including information such as verification results and scoring details. The stored verification data can be used for subsequent auditing or analysis, and the verification report provides an intuitive verification result for the user or relevant institutions, facilitating decision-making and record-keeping.

[0024] As Figure 1 shown, 11. According to the voice command, use the video acquisition module to obtain the dynamic video of the user and perform preprocessing to obtain the preprocessed dynamic video, including: Use a camera to capture the user's dynamic video in real time. The dynamic video includes the user's actions, expressions, and environment when completing various instructions. In the dynamic video, use a face detection algorithm to identify the position of the user's face and perform alignment processing on the detected face. In the dynamic video, use a denoising algorithm to reduce the noise and interference in the video to obtain a denoised dynamic video. In the dynamic video, use video enhancement technology for correction to obtain a corrected dynamic video, that is, the preprocessed dynamic video.

[0025] In the embodiments of the present invention, a camera device is used to capture in real time the actions, expressions, and environment of the user when completing various voice instructions, generating an original dynamic video containing this information. Real-time capture ensures the immediacy and authenticity of the video content, providing basic data for subsequent analysis. In the captured dynamic video, a face detection algorithm is applied to accurately identify the position of the user's face and perform alignment processing on the detected face, that is, adjusting the pose and angle of the face to conform to a certain standard or template, so that subsequent analysis can be carried out more accurately. Face detection and alignment are key steps in video preprocessing, ensuring the accurate extraction of face features in the video and providing high-quality data for subsequent face feature analysis. A denoising algorithm is used in the dynamic video to identify and reduce the noise and interference elements in the video, such as spots, flickers, jitters, etc. in the image. The denoising algorithm may include spatial domain denoising, temporal domain denoising, or a combination of both methods to effectively remove the noise in the video. Video denoising improves the clarity and quality of the video, reduces the interference of noise on subsequent analysis, and makes face features, action details, etc. more clearly distinguishable. Further enhancement processing is performed on the denoised dynamic video, such as brightness adjustment, contrast enhancement, color correction, etc. The enhancement technology aims to improve the visual effect of the video, making it more suitable for subsequent face feature analysis, action recognition, and other tasks. Video enhancement and correction not only improve the visual effect of the video but also ensure the accuracy and reliability of the video data, providing a high-quality data basis for subsequent multi-modal AI analysis.

[0026] As Figure 1 shown in FIG. 12, according to the preprocessed dynamic video, use a dynamic analysis module to analyze the face feature points and timing to obtain a dynamic analysis result, including: According to the preprocessed dynamic video, the dynamic analysis module uses OpenCV for face detection to obtain the coordinates of the face region. According to the coordinates of the face region, extract face feature points. The face feature points include key points such as the corners of the eyes, the corners of the mouth, and the nose. Calculate the behavioral features of the face in dynamic behavior based on facial feature points to obtain behavioral features, where the behavioral features include the opening and closing degree of the eyes, the changes of the corners of the mouth, and the rotation angle of the head. Identify the dynamic behavior in the video according to the behavioral features. Conduct a temporal analysis of the identified dynamic behavior to obtain the temporal analysis result of the dynamic behavior. According to the dynamic behavior and the temporal analysis result, obtain the dynamic analysis result by comparing the changes in behavioral features in consecutive frames.

[0027] In the embodiments of the present invention, a camera device is used to capture in real time the actions, expressions, and the environment where the user shows when completing various voice commands, generating an original dynamic video containing this information. The real-time capture ensures the immediacy and authenticity of the video content, providing basic data for subsequent analysis; in the captured dynamic video, a face detection algorithm is applied to accurately identify the position of the user's face, and the detected face is aligned, that is, the pose and angle of the face are adjusted to conform to a certain standard or template, so that subsequent analysis can be carried out more accurately. Face detection and alignment are key steps in video preprocessing, which ensure the accurate extraction of facial features in the video and provide high-quality data for subsequent facial feature analysis; a denoising algorithm is used in the dynamic video to identify and reduce the noise and interference elements in the video, such as spots, flickers, jitters, etc. in the image. The denoising algorithm may include spatial domain denoising, temporal domain denoising, or a method combining the two to effectively remove the noise in the video. Video denoising improves the clarity and quality of the video, reduces the interference of noise on subsequent analysis, and makes facial features, action details, etc. more clearly distinguishable; further enhancement processing is performed on the denoised dynamic video, such as brightness adjustment, contrast enhancement, color correction, etc., to improve the visual effect of the video and make it more suitable for subsequent facial feature analysis, action recognition, etc. tasks. Video enhancement and correction not only improve the visual effect of the video, but also ensure the accuracy and reliability of the video data, providing a high-quality data basis for subsequent multi-modal AI analysis.

[0028] As Figure 1 shown in 13, according to the preprocessed dynamic video, use a voice analysis module to perform voice feature analysis to obtain a voice analysis result, including: Separate the audio data from the preprocessed dynamic video and perform preprocessing to obtain the preprocessed audio data. According to the preprocessed audio data, analyze the frequency and temporal domain features of the sound, identify each phoneme in the audio, and determine the pronunciation features of each phoneme. Analyze according to the pronunciation characteristics of phonemes, the emotional state of the speaker in the audio, and the tone of the sentence to obtain intonation characteristics, which include the rhythm, pitch change, duration, and intensity of the sentence; Convert the audio data into a spectrogram based on the preprocessed audio data, and perform frequency component and amplitude distribution analysis to obtain the audio spectrum analysis result; Align and analyze the dynamic behavior time series analysis result with the time domain characteristics of the audio data to obtain the synchronization analysis result; Generate a result report, that is, the speech analysis result, according to the intonation characteristics, audio spectrum result, and synchronization analysis result.

[0029] In the embodiments of the present invention, audio data is separated from the preprocessed dynamic video, and these audio data are preprocessed, such as denoising, normalization, volume adjustment, etc., to obtain preprocessed audio data. The preprocessed audio data is clearer and purer, providing a high-quality basis for subsequent speech feature analysis; for the preprocessed audio data, using a phoneme recognition algorithm to analyze the frequency and time-domain characteristics of the sound, identify each phoneme in the audio, and determine the pronunciation characteristics of each phoneme, including the duration, frequency distribution, energy distribution, etc. of the phoneme. Phoneme recognition and pronunciation feature analysis provide important speech information for subsequent emotional state analysis and intonation feature extraction; according to the pronunciation characteristics of the phonemes, combined with the speaking speed, intonation, pauses, etc. of the speaker in the audio, analyze the emotional state of the speaker, such as joy, sadness, anger, etc., and analyze the tone of the sentence, such as statement, question, exclamation, etc. Emotional state and tone analysis help to understand the true intention and emotional state of the speaker, providing important reference for subsequent synchronization analysis and result report generation; convert the preprocessed audio data into a spectrogram to display the distribution of the audio data in frequency and time, analyze the frequency components and amplitude distribution of the spectrogram, and identify the main frequency components and amplitude distribution characteristics in the audio. Audio spectrum analysis provides an intuitive visualization tool for understanding the frequency characteristics of the audio data and provides important spectrum information for subsequent synchronization analysis and result report generation; align the results of dynamic behavior time series analysis with the time-domain characteristics of the audio data to analyze the synchronization relationship between the dynamic behavior and the audio data, such as the synchronization of actions and speech, the synchronization of expressions and intonation, etc. Dynamic behavior time series analysis and audio data alignment help to understand the association between the actions and expressions of the speaker during speech and the speech, providing important synchronization information for subsequent synchronization analysis and result report generation; according to the intonation characteristics, audio spectrum results and synchronization analysis results, comprehensively generate a result report. The result report should include the main findings, conclusions and suggestions of the speech analysis results, as well as relevant visualization charts and data. The result report provides a comprehensive and intuitive speech analysis result for the end user, helping the user to understand the true intention, emotional state and speech characteristics of the speaker, and providing important reference for subsequent decision-making and actions.

[0030] As Figure 1 shown in FIG. 14, according to the preprocessed dynamic video, background environment analysis is performed to obtain the environmental authenticity analysis result, including: According to the preprocessed dynamic video, the video is split into a series of individual image frames, and the image frames are enhanced to obtain enhanced image frames; According to the enhanced image frames, a target detection algorithm is used to identify the image frames to obtain the target detection result; According to the object detection results, extract the background region in the image frame, and perform feature analysis on the extracted background region to obtain the feature analysis results; According to the feature analysis results, compare with the preset real background feature library to obtain the environmental authenticity analysis results.

[0031] In the embodiments of the present invention, the preprocessed dynamic video is split into a series of individual image frames. These image frames are static representations of the video content, facilitating subsequent processing. Perform enhancement processing on the split image frames, such as contrast enhancement, brightness adjustment, denoising, etc., to improve the image quality and make subsequent object detection and feature analysis more accurate, obtaining the enhanced image frames, providing high-quality input for subsequent object detection and background analysis; Use object detection algorithms to identify the enhanced image frames. These algorithms can automatically detect objects or targets in the image, such as people, animals, buildings, etc. The effect is: Obtain the object detection results, including information such as the position, size, and category of the detected targets. This information helps with subsequent background region extraction; According to the object detection results, extract the background area in the image frame, achieved by excluding the detected targets, perform feature analysis on the extracted background region, extract features such as the color, texture, and shape of the background, obtaining the feature analysis results. These features describe the uniqueness and authenticity of the background environment; Compare the feature analysis results with the preset real background feature library. The feature library contains feature data of various real environments and is used to evaluate the authenticity of the background. Through comparison, obtain the environmental authenticity analysis results, which are used to indicate the authenticity degree of the background environment.

[0032] As Figure 1 shown in Figure 15, according to the dynamic analysis results, speech analysis results, and environmental authenticity analysis results, the verification scoring module uses a weighted algorithm for scoring to obtain the user identity authenticity score, including: According to the dynamic analysis results, speech analysis results, and environmental authenticity analysis results, perform normalization processing to obtain the normalized analysis results; According to the normalized analysis results, determine the weights of the dynamic analysis results, speech analysis results, and environmental authenticity analysis results; According to the weights of the dynamic analysis results, speech analysis results, and environmental authenticity analysis results, use for scoring to obtain the user identity authenticity score, where is the final user identity authenticity score, is the dynamic analysis result, is the speech analysis result, is the environmental authenticity analysis result, is the weight of the dynamic analysis result, is the weight of the speech analysis result, is the weight of the environmental authenticity analysis result.

[0033] In the embodiments of the present invention, the dynamic analysis result, the voice analysis result, and the environmental authenticity analysis result are normalized. Normalization is a process of converting numerical values of different magnitudes or ranges into the same magnitude or range for subsequent comparison and calculation. The analysis results after normalization have a unified numerical range and magnitude, facilitating subsequent weighted calculation; according to the analysis results after normalization, as well as the requirements and background of the actual application scenario, the weights of the dynamic analysis result, the voice analysis result, and the environmental authenticity analysis result are determined. These weights reflect the importance of each analysis result in the final score. Determining reasonable weights is a key step to ensure the accuracy of the final score. Different weight assignments will directly affect the final user identity authenticity score; a weighted algorithm is used for score calculation, multiplying the dynamic analysis result, the voice analysis result, and the environmental authenticity analysis result by their respective weights, and then adding these three weighted results to obtain the final user identity authenticity score: through weighted score calculation, a final score that combines the dynamic analysis result, the voice analysis result, and the environmental authenticity analysis result can be obtained, and this score can more comprehensively reflect the authenticity of the user identity; in actual applications, it may be necessary to verify and adjust the weights. This can be achieved by comparing the consistency between the actual verification result and the score result. If there is a large deviation between the score result and the actual situation, the weights may need to be readjusted. Through verification and adjustment, the accuracy and reliability of the weighted scoring algorithm in actual applications can be ensured.

[0034] As Figure 1 shown in 16, according to the user identity authenticity score, use the back-end management system for storage and output a user identity authenticity verification report, including: Construct a database for storing the user identity authenticity score and encrypt the database to obtain an encrypted database; Store the user identity authenticity score in the encrypted database and record the timing information; Generate and output a user identity authenticity verification report according to the user identity authenticity score; Conduct data analysis and mining on the stored score results to obtain analysis results; Update and optimize the verification rules according to the analysis results.

[0035] In an embodiment of the present invention, a database dedicated to storing the user identity authenticity score is constructed. The database has high security and stability to ensure the integrity and security of the score data. The database is encrypted to prevent unauthorized access and data leakage. The encryption technology may include, but is not limited to, encryption algorithms such as AES and RSA to ensure the security of the database content during transmission and storage, resulting in an encrypted database, which provides security for the storage of score data. The user identity authenticity score is stored in the encrypted database, and each score should be associated with a specific user or verification event for subsequent query and analysis. The timing information of the score data, such as the scoring time and scoring source, is recorded. The timing information helps to track the changes in the score data and analyze the stability and reliability of the score data. The score data and its timing information are stored securely and orderly in the database, providing a basis for subsequent analysis and report generation. According to the user identity authenticity score, a user identity authenticity verification report is automatically generated. The report should include information such as the scoring result, scoring grade, and possible verification conclusions. The generated verification report is output to a specified location or device, such as the user's email or the interface of the back-end management system. The output method should be convenient for the user to view and download, enabling the user to easily obtain the identity verification result and understand their own identity authenticity score and verification conclusion. In-depth data analysis and mining are performed on the stored score results. The analysis content may include the distribution characteristics, change trends, and outlier detection of the score data. Through data analysis and mining, potential rules and trends in the score data can be discovered, providing data support for the subsequent update and optimization of verification rules. According to the results of data analysis and mining, the existing verification rules are updated and optimized. The update content may include adjusting the scoring threshold, adding new verification metrics, and optimizing the verification process. By continuously updating and optimizing the verification rules, the accuracy and efficiency of identity verification can be improved, and the risks of false positives and false negatives can be reduced.

[0036] As Figure 2 shown, an embodiment of the present invention further provides a user identity authenticity verification system 20 based on multi-modal AI, including: An acquisition module 21, configured to use a video acquisition module to acquire a dynamic video of a user and perform preprocessing to obtain a preprocessed dynamic video; A processing module 22, configured to analyze facial feature points and timing using a dynamic analysis module based on the preprocessed dynamic video to obtain a dynamic analysis result, analyze voice features using a voice analysis module based on the preprocessed dynamic video to obtain a voice analysis result, perform background environment analysis based on the preprocessed dynamic video to obtain an environmental authenticity analysis result, and verify that a scoring module uses a weighted algorithm for scoring based on the dynamic analysis result, the voice analysis result, and the environmental authenticity analysis result to obtain a user identity authenticity score, and store the score using a background management system and output a user identity authenticity verification report.

[0037] It should be noted that this system corresponds to the above method. All implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0038] An embodiment of the present invention further provides a computing device, including: a processor and a memory storing a computer program. When the computer program is run by the processor, it executes the method as described above. All implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0039] An embodiment of the present invention further provides a computer-readable storage medium storing instructions that, when run on a computer, cause the computer to execute the method as described above. All implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0040] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0041] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, systems, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0042] In the embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the system or unit can be in electrical, mechanical or other forms.

[0043] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0044] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0045] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0046] In addition, it should be noted that in the systems and methods of the present invention, obviously, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations shall be regarded as equivalent solutions of the present invention. Moreover, the steps of performing the above series of processes can naturally be executed chronologically in the order described, but it is not necessary to be executed chronologically. Some steps can be executed in parallel or independently of each other. For those of ordinary skill in the art, it is understandable that all or any steps or components of the method and system of the present invention can be implemented in any computing system (including processors, storage media, etc.) or a network of computing systems in the form of hardware, firmware, software, or a combination thereof, which can be achieved by those of ordinary skill in the art by applying their basic programming skills after reading the description of the present invention.

[0047] Therefore, the object of the present invention can also be achieved by running a program or a set of programs on any computing system. The computing system can be a well-known general system. Therefore, the object of the present invention can also be achieved only by providing a program product containing program code for implementing the method or system. That is to say, such a program product also constitutes the present invention, and the storage medium storing such a program product also constitutes the present invention. Obviously, the storage medium can be any well-known storage medium or any storage medium developed in the future. It should also be noted that in the systems and methods of the present invention, obviously, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations shall be regarded as equivalent solutions of the present invention. Moreover, the steps of performing the above series of processes can naturally be executed chronologically in the order described, but it is not necessary to be executed chronologically. Some steps can be executed in parallel or independently of each other.

[0048] The above is the preferred embodiment of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A user identity authenticity verification method based on multimodal AI, characterized in that: The method comprises: According to the voice command, the video acquisition module is used to obtain the user's dynamic video and pre-process it to obtain the pre-processed dynamic video; According to the pre-processed dynamic video, the dynamic analysis module is used to analyze the facial feature points and timing to obtain the dynamic analysis results; According to the pre-processed dynamic video, a speech analysis module is used to perform speech feature analysis to obtain a speech analysis result; According to the pre-processed dynamic video, background environment analysis is performed to obtain the environmental authenticity analysis results; Based on the results of dynamic analysis, voice analysis and environmental authenticity analysis, the verification scoring module uses a weighted algorithm to score to obtain the user identity authenticity score; Based on the user identity authenticity score, the background management system is used to store and output the user identity authenticity verification report.

2. The user identity authenticity verification method based on multimodal AI according to claim 1 is characterized in that: According to the voice command, the video acquisition module is used to obtain the user's dynamic video and preprocess it to obtain the preprocessed dynamic video, including: Use a camera to capture the user's dynamic video in real time. The dynamic video includes the user's movements, expressions, and environment when completing various instructions. Use face detection algorithms to identify the user's face position in dynamic videos and align the detected faces; Use denoising algorithms in dynamic videos to reduce noise and interference in the video to obtain denoised dynamic videos; The video enhancement technology is used to correct the dynamic video to obtain a corrected dynamic video, that is, a preprocessed dynamic video.

3. The user identity authenticity verification method based on multimodal AI according to claim 2 is characterized in that: According to the pre-processed dynamic video, the dynamic analysis module is used to analyze the facial feature points and timing to obtain dynamic analysis results, including: According to the preprocessed dynamic video, the dynamic analysis module uses OpenCV to perform face detection and obtain the coordinates of the face area; According to the coordinates of the face area, facial feature points are extracted, including key points such as the corners of the eyes, the corners of the mouth, and the nose; According to the facial feature points, the behavioral features of the face in dynamic behavior are calculated to obtain the behavioral features, which include the degree of opening and closing of the eyes, the changes in the corners of the mouth, and the rotation angle of the head; Identify dynamic behaviors in the video based on behavioral characteristics; Performing timing analysis on the identified dynamic behavior to obtain a dynamic behavior timing analysis result; According to the dynamic behavior and timing analysis results, the dynamic analysis results are obtained by comparing the changes in behavioral characteristics in consecutive frames.

4. The user identity authenticity verification method based on multimodal AI according to claim 3 is characterized in that: According to the pre-processed dynamic video, the speech analysis module is used to perform speech feature analysis to obtain speech analysis results, including: Separating audio data from the preprocessed dynamic video and performing preprocessing to obtain preprocessed audio data; According to the pre-processed audio data, the frequency and time domain characteristics of the sound are analyzed to identify each phoneme in the audio and determine the pronunciation characteristics of each phoneme; Analyze the pronunciation characteristics of phonemes, the speaker's emotional state in the audio, and the tone of the sentence to obtain intonation features, which include the rhythm of the sentence, pitch change, length, and intensity. According to the preprocessed audio data, the audio data is converted into a spectrum diagram, and the frequency component and amplitude distribution analysis is performed to obtain the audio spectrum analysis result; Align and analyze the dynamic behavior timing analysis results with the time domain characteristics of the audio data to obtain synchronized analysis results; According to the intonation features, audio spectrum results and synchronous analysis results, a result report, namely the speech analysis result, is generated.

5. The user identity authenticity verification method based on multimodal AI according to claim 4 is characterized in that: Based on the pre-processed dynamic video, background environment analysis is performed to obtain environmental authenticity analysis results, including: According to the pre-processed dynamic video, the video is split into a series of separate image frames, and the image frames are enhanced to obtain enhanced image frames; According to the enhanced image frame, the target detection algorithm is used to identify the image frame to obtain the target detection result; According to the target detection result, a background area in the image frame is extracted, and feature analysis is performed on the extracted background area to obtain a feature analysis result; According to the feature analysis results, it is compared with the preset real background feature library to obtain the environmental authenticity analysis results.

6. The user identity authenticity verification method based on multimodal AI according to claim 5 is characterized in that: Based on the results of dynamic analysis, voice analysis, and environmental authenticity analysis, the verification scoring module uses a weighted algorithm to score to obtain a user identity authenticity score, including: Performing normalization processing according to the dynamic analysis results, the speech analysis results and the environmental authenticity analysis results to obtain normalized analysis results; Determine the weights of the dynamic analysis results, the speech analysis results, and the environmental authenticity analysis results according to the normalized analysis results; According to the weights of dynamic analysis results, speech analysis results and environmental authenticity analysis results, use Score to get the user identity authenticity score, where: Score the authenticity of the final user identity, For dynamic analysis results, is the result of speech analysis. The results of the environmental authenticity analysis are: is the weight of the dynamic analysis result, is the weight of the speech analysis result, The weight of the environmental authenticity analysis results.

7. The user identity authenticity verification method based on multimodal AI according to claim 6 is characterized in that: Based on the user identity authenticity score, the background management system is used to store and output the user identity authenticity verification report, including: Constructing a database for storing user identity authenticity scores, and encrypting the database to obtain an encrypted database; Store the user identity authenticity score in an encrypted database and record the time series information; Generate and output a user identity authenticity verification report based on the user identity authenticity score; Perform data analysis and mining on the stored scoring results to obtain analysis results; Based on the analysis results, the validation rules are updated and optimized.

8. A user identity authenticity verification system based on multimodal AI, characterized in that: include: The acquisition module is used to use the video acquisition module to acquire the user's dynamic video and pre-process it to obtain the pre-processed dynamic video; The processing module is used to analyze facial feature points and timing using the dynamic analysis module according to the preprocessed dynamic video to obtain dynamic analysis results; to use the voice analysis module to analyze voice features according to the preprocessed dynamic video to obtain voice analysis results; to perform background environment analysis according to the preprocessed dynamic video to obtain environmental authenticity analysis results; based on the dynamic analysis results, the voice analysis results and the environmental authenticity analysis results, the verification scoring module uses a weighted algorithm to score to obtain a user identity authenticity score; based on the user identity authenticity score, the background management system is used to store it and output a user identity authenticity verification report.

9. A computing device, characterized in that include: one or more processors; A storage system for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Middle-aged and elderly tourism ERP business whole-process management method and system

    CN120410783A

  • Artificial intelligence psychological assessment method and device based on multiple modes

    CN120600318A