A video identification method and system based on human posture determination model

Through multimodal analysis based on the human posture judgment model, combined with human voice emotions and posture change data, the problems of simple judgment steps and low accuracy in existing video authentication technology are solved, and more accurate and intuitive video authentication results are achieved.

CN117152594BActive Publication Date: 2025-09-26CHINA ACADEMY OF INFORMATION & COMM
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310865307.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-13
Publication Date
2025-09-26
Estimated Expiration
2043-07-13

AI Technical Summary

Technical Problem

The judgment steps in existing video authentication technology are too simple and lack credibility, and the single data comparison leads to low accuracy of the authentication results.

Method used

A video authentication method based on the human posture judgment model is adopted. By extracting human voice emotion data and human posture change data, an emotion judgment model and a human motion judgment model are established, multimodal analysis is performed, and emotion authenticity marks and posture change degree marks are generated. Finally, the authenticity of the video is analyzed.

Benefits of technology

It improves the accuracy and explainability of video authentication, provides direct judgment basis through human body movement trajectory and voice changes, and enhances the ability to judge the authenticity of videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117152594B_ABST
    Figure CN117152594B_ABST
Patent Text Reader

Abstract

The present invention discloses a video identification method and system based on a human posture determination model, relates to the field of video identification technology, and is used to solve the problem of the singleness and large error of current video identification methods. The specific steps include: S100, preprocessing the identification video and extracting video feature data, wherein the video feature data includes human voice emotion data and human posture change data; S200, establishing an emotion determination model, and generating emotion determination information through emotion determination of the extracted human voice emotion data; S300, inputting real emotion information, integrating the real emotion information with the emotion determination information, and generating an emotion authenticity mark for the identification video; S400, establishing a human motion determination model, comparing and analyzing the extracted human posture change data, generating a motion change factor, substituting the motion change factor into a change factor threshold for comparison, and generating a posture change degree mark for the identification video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video identification, and in particular to a video identification method and system based on a human body posture determination model. Background Art

[0002] Traditional video authentication techniques primarily rely on comparative analysis of images, video bitrates, or audio data. Publication No. CN103327320A, for example, proposes a method for identifying pseudo-high-bitrate videos. This method analyzes how video quality changes when the bitrate is reduced to determine whether the video has undergone a bitrate increase, thereby identifying whether the video is at its original bitrate. Furthermore, a support vector machine classifier is used to determine the original bitrate of the video. This invention can effectively detect the original bitrate of a video that has undergone a bitrate increase, providing an effective and simple method for identifying the original bitrate of a video.

[0003] However, after comparison, it was found that the following deficiencies still existed in the comparison documents:

[0004] 1. The judgment steps are too simple, and the identification results lack credibility when they come out.

[0005] 2. Since only single data was used for data comparison in a single model without repeated verification, there were errors in the judgment process and the accuracy of the identification results was low.

[0006] In order to solve the above-mentioned problems, a video identification method and system based on a human posture determination model are proposed. Summary of the Invention

[0007] The purpose of the present invention is to provide a video identification method and system based on a human posture determination model to overcome the shortcomings in the background technology.

[0008] In order to achieve the above object, the present invention provides the following technical solution: the video identification method based on the human body posture determination model comprises the following steps:

[0009] S100, pre-processing the identification video and extracting video feature data, wherein the video feature data includes human voice emotion data and human body posture change data;

[0010] S200, establishing an emotion determination model, and generating emotion determination information by performing emotion determination on the extracted human voice emotion data;

[0011] S300, inputting real emotion information, integrating the real emotion information with emotion determination information, and generating an emotion authenticity mark for the identification video;

[0012] S400, establishing a human motion determination model, performing comparative analysis on the extracted human posture change data, generating a motion change factor, substituting the motion change factor into a change factor threshold for comparison, and generating a posture change degree mark for the identification video;

[0013] S500, establishing a re-analysis module, performing video authenticity analysis by identifying the emotional authenticity mark and the posture change degree mark in the video, and marking the authenticity target of the identified video according to the analysis result;

[0014] S600: Feedback the result of identifying the true and false target markings in the video to the display panel according to the true and false target markings.

[0015] In a preferred embodiment, the human voice emotion data includes sound intensity Int, pitch variation amplitude Pv and speech speed Sp, wherein the sound intensity Int represents the energy or volume level of the sound, the pitch variation amplitude Pv represents the pitch variation amplitude of the sound, and the speech speed Sp represents the speaking speed;

[0016] The human body posture change data includes the elbow movement value Emv, the human body lateral displacement path Lbd and the portrait proportion change value Pic, wherein the elbow movement value Emv is the value of the elbow point moving within the detection interval, the human body lateral displacement path Lbd is the displacement value of the human body within the detection interval, and the portrait proportion change value Pic is the proportion change value of the presented portrait within the detection interval.

[0017] In a preferred embodiment, the emotion determination step includes:

[0018] The identification video is divided into n identification intervals according to the video frame, where n is an integer greater than 1. The sound intensity Int, tone variation amplitude Pv and speech speed Sp of the n identification intervals are obtained respectively, and the emotion influence factor α of the n identification intervals is obtained by formulating and analyzing them. q ;

[0019] The emotion determination information includes depressed emotion information, calm emotion information and excited emotion information, and sets the emotion influence factor reference thresholds Ee1 and Ee2, where Ee1<Ee2, and α q Substitute the emotional influence factor reference thresholds Ee1 and Ee2 for comparison. When 0<α q <Ee1, generate low emotion information for the detection interval; when Ee1<α q When α < Ee2, the smooth emotion information is generated for the detection interval; when α q When >Ee2, excitement emotion information is generated for the detection interval.

[0020] In a preferred embodiment, the steps of integrating the real emotion information and the emotion determination information are as follows:

[0021] According to the emotional reflection observed in the image, real emotional information is input into the emotional judgment module, and the real emotional information includes actual low emotional information, actual medium emotional information and actual high emotional information;

[0022] The emotional authenticity mark includes an emotional authenticity mismatch mark, a low emotional authenticity mark, a flat emotional authenticity mark and an excited emotional authenticity mark. If there is a difference between the input real emotional information and the emotional judgment information, an emotional authenticity mismatch mark is generated for the detection interval; if there is no difference between the input real emotional information and the emotional judgment information, when the corresponding actual low emotional information and the low emotional information are in the same interval, a low emotional authenticity mark is generated; when the corresponding actual medium emotional information and the flat emotional information are in the same interval, a flat emotional authenticity mark is generated; when the corresponding actual high emotional information and the excited emotional information are in the same interval, an excited emotional authenticity mark is generated.

[0023] In a preferred embodiment, the step of generating the posture change degree mark is:

[0024] Obtain the elbow movement value Emv, the human body lateral displacement path Lbd, and the portrait proportion change value Pic in n detection intervals, and obtain the motion change factor γ through formula analysis;

[0025] The posture change degree mark includes a small change mark, a medium change mark and a large change mark, and the motion change factor comparison thresholds Sce1 and Sce2 are set, where Sce1>Sce2>0, and γ q Substitute the motion change factor comparison threshold Sce1 and Sce2 for comparison. When 0<α q <Sce2, a small change mark is generated for the detection interval; when Sce2<α q When α<Sce1, a medium-amplitude change mark is generated for the detection interval; when α q When >Sce1, a large change flag is generated for the detection interval.

[0026] In a preferred embodiment, the video authenticity analysis step includes:

[0027] The emotion authenticity marks and posture change degree marks in n detection intervals are input into the re-analysis module for integration processing. If a single detection interval has an emotion authenticity mismatch mark and any posture change degree mark, an excitement authenticity mark and a small change mark, or a depression authenticity mark and a large change mark, the detection interval is marked as a forged target;

[0028] When a single detection interval has an exciting true mark and a medium-amplitude change mark, a gentle true mark and a large-amplitude change mark, a gentle true mark and a small-amplitude change mark, or a depressed true mark and a medium-amplitude change mark, the detection interval is marked as a questionable target;

[0029] When a single detection interval has an exciting true mark and a large-amplitude change mark, a flat true mark and a medium-amplitude change mark, or a low true mark and a small-amplitude change mark, the detection interval is marked as a true target.

[0030] The present invention also provides a video identification system based on a human posture determination model, comprising:

[0031] Video preprocessing module: preprocess the identification video to extract human voice emotion data and human posture change data;

[0032] Emotional information processing module: integrates and analyzes the extracted human voice emotion data to generate emotion judgment information;

[0033] Emotional authenticity analysis module: compares the input real emotion information with the emotion judgment information to generate an emotion authenticity mark;

[0034] Posture change degree analysis module: analyzes and processes the extracted human posture change data to generate posture change degree marks;

[0035] Video authenticity analysis and target labeling module: This module combines emotion authenticity labels and gesture change degree labels for analysis, and uses the analysis results to label the authenticity of the identified video.

[0036] Identification result display module: Feedback the video identification results to the display panel based on the true and false target labels.

[0037] In a preferred embodiment, the emotional authenticity mark includes an emotional authenticity mismatch mark, a low-level authenticity mark, a flat-level authenticity mark, and an excited-level authenticity mark. The steps for generating the emotional authenticity mark are:

[0038] The human voice emotion data includes sound intensity Int, pitch variation amplitude Pv and speech speed Sp, wherein the sound intensity Int represents the energy or volume level of the sound, the pitch variation amplitude Pv represents the pitch variation amplitude of the sound, and the speech speed Sp represents the speaking speed;

[0039] The identification video is divided into n identification intervals according to the video frame, where n is an integer greater than 1. The sound intensity Int, tone variation amplitude Pv and speech speed Sp of the n identification intervals are obtained respectively, and the emotion influence factor α of the n identification intervals is obtained by formulating and analyzing them. q , set the reference thresholds of emotion influence factors Ee1 and Ee2, where Ee1<Ee2, and set αq Substitute the emotional influence factor reference thresholds Ee1 and Ee2 for comparison. When 0<α q <Ee1, generate low emotion information for the detection interval; when Ee1<α q When α < Ee2, the smooth emotion information is generated for the detection interval; when α q >Ee2, generate excitement emotion information for the detection interval;

[0040] According to the emotional reflection observed in the image, real emotional information is input into the emotional judgment module, and the real emotional information includes actual low emotional information, actual medium emotional information and actual high emotional information;

[0041] If there is a difference between the input real emotion information and the emotion judgment information, an emotion authenticity mismatch mark is generated for the detection interval; if there is no difference between the input real emotion information and the emotion judgment information, when the corresponding actual low emotion information and the low emotion information are in the same interval, a low true mark is generated; when the corresponding actual medium emotion information and the flat emotion information are in the same interval, a flat true mark is generated; when the corresponding actual high emotion information and the excited emotion information are in the same interval, an excited true mark is generated.

[0042] In a preferred embodiment, the steps of generating a posture change degree mark are:

[0043] The human body posture change data includes the elbow movement value Emv, the human body lateral displacement path Lbd and the portrait proportion change value Pic, wherein the elbow movement value Emv is the value of the elbow point moving within the detection interval, the human body lateral displacement path Lbd is the displacement value of the human body within the detection interval, and the portrait proportion change value Pic is the proportion change value of the presented portrait within the detection interval;

[0044] Obtain the elbow movement value Emv, the human body lateral displacement path Lbd, and the portrait proportion change value Pic in n detection intervals, obtain the motion change factor γ through formula analysis, set the motion change factor comparison thresholds Sce1 and Sce2, where Sce1>Sce2>0, and set γ q Substitute the motion change factor comparison threshold Sce1 and Sce2 for comparison. When 0<α q <Sce2, a small change mark is generated for the detection interval; when Sce2<α q When α<Sce1, a medium-amplitude change mark is generated for the detection interval; when α q When >Sce1, a large change flag is generated for the detection interval.

[0045] In a preferred embodiment, the steps of combining and analyzing the emotion authenticity mark and the posture change degree mark are as follows:

[0046] The emotion authenticity marks and posture change degree marks in n detection intervals are input into the re-analysis module for integration processing. If a single detection interval has an emotion authenticity mismatch mark and any posture change degree mark, an excitement authenticity mark and a small change mark, or a depression authenticity mark and a large change mark, the detection interval is marked as a forged target;

[0047] When a single detection interval has an exciting true mark and a medium-amplitude change mark, a gentle true mark and a large-amplitude change mark, a gentle true mark and a small-amplitude change mark, or a depressed true mark and a medium-amplitude change mark, the detection interval is marked as a questionable target;

[0048] When a single detection interval has an exciting true mark and a large-amplitude change mark, a flat true mark and a medium-amplitude change mark, or a low true mark and a small-amplitude change mark, the detection interval is marked as a true target.

[0049] In the above technical solution, the technical effects and advantages provided by the present invention are:

[0050] The present invention uses the human body's motion trajectory and the changes in the human voice to more intuitively display the dynamic information and emotional changes in the video. By analyzing these dynamic features, it provides a more direct and explainable basis for the authentication results, helping users better understand and accept the judgment results.

[0051] Furthermore, multimodal analysis can be implemented, combining the body's motion trajectory and changes in the human voice to provide more information sources and enhance the ability to determine the authenticity of the video;

[0052] In addition, the accuracy of authentication has been improved. The movement trajectory of the human body and the changes in the human voice can be used as key indicators of the authenticity of the video. By analyzing the image features of human movement and the changes in the intensity of the sound, it is possible to more accurately determine whether the video has been tampered with or forged. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction to the drawings required for use in the embodiments will be given below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0054] Figure 1 The figure is a module diagram of a video identification system based on a human posture determination model of the present invention.

[0055] Figure 2 The present invention is a flow chart of a video identification method based on a human posture determination model. DETAILED DESCRIPTION

[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0057] It should be noted that both the first and second embodiments are based on the video image features of human presence.

[0058] Example 1

[0059] See Figure 1 It can be seen that the video identification system based on the human posture determination model described in this embodiment includes:

[0060] Video pre-processing module, emotion information processing module, emotion authenticity analysis module, posture change degree analysis module, video authenticity analysis and target marking module, and identification result display module;

[0061] The preprocessing steps specifically include video reading, video decoding, inter-frame processing, image noise reduction processing, and extraction of video comparison features. It should be noted that the preprocessing steps are performed in video processing software, and the specific video processing software involved includes Fmpeg and MATLAB.

[0062] The human voice emotion data includes sound intensity Int, pitch variation amplitude Pv and speech speed Sp, wherein the sound intensity Int represents the energy or volume level of the sound, the pitch variation amplitude Pv represents the pitch variation amplitude of the sound, and the speech speed Sp represents the speaking speed;

[0063] It should be noted that the emotion determination model specifically adopts a machine learning algorithm model. Specifically, the machine learning algorithm includes support vector machine, random forest, neural network, etc. The processing steps of the emotion machine learning algorithm model include:

[0064] Extract human voice audio data from videos and perform feature extraction. Common audio features include sound intensity, spectral characteristics, speaking rate, and pitch. Collect a set of annotated video datasets containing emotion-related labels, such as excitement, sadness, and fear. These datasets can be used to train and evaluate emotion detection models. The emotion detection model can then be used to train emotion detection.

[0065] The human body posture change data includes the elbow movement value Emv, the human body lateral displacement path Lbd and the portrait proportion change value Pic, wherein the elbow movement value Emv is the value of the elbow point moving within the detection interval, the human body lateral displacement path Lbd is the displacement value of the human body within the detection interval, and the portrait proportion change value Pic is the proportion change value of the presented portrait within the detection interval.

[0066] It should be noted that the portrait ratio refers to the coverage ratio of a single-frame image corresponding to the portrait coverage. The human motion judgment model can adopt but is not limited to a hidden Markov model: by using the human motion trajectory as sequence data, a hidden Markov model can be used to model the conversion relationship between different emotional states, and the most likely emotional state sequence can be inferred based on the observed motion trajectory. The specific algorithms involved include forward algorithm, backward algorithm, Viterbi algorithm and Baum-Welch algorithm.

[0067] Emotional information processing module: integrates and analyzes the extracted human voice emotion data to generate emotion judgment information;

[0068] The identification video is divided into n identification intervals according to the video frame, where n is an integer greater than 1. The sound intensity Int, tone variation amplitude Pv and speech speed Sp of the n identification intervals are obtained respectively, and the emotion influence factor α of the n identification intervals is obtained by formulating and analyzing them. q , the specific formula is:

[0069] α q = ;

[0070] Wherein, a1, a2 and a3 are weight factors of sound intensity Int, pitch variation Pv and speech speed Sp respectively, and a2>a1>a3, a1+a2+a3=1.135, and q is the number of video frames included in the identification interval, and q= , k is the error correction constant, K>0.

[0071] It should be noted that the image backgrounds displayed by the number of video frames selected by q are all under the same background image, that is, this identification scheme is based on the above α under the condition that the background image remains unchanged. q According to the formula, higher voice intensity may be related to emotions such as excitement and anger, larger tone changes may be related to emotions such as excitement and excitement, and faster speaking speed may be related to emotions such as excitement and tension. The specific analysis is as follows:

[0072] Emotional excitement: The voice has high intensity. Excitement is usually accompanied by high volume and energy. The speech rate is fast. Excited people will speak faster. The voice has large changes in pitch. When excited, the voice has large changes in pitch.

[0073] Angry emotions: Strong voice intensity. Angry emotions are usually accompanied by high volume and energy. Fast speech rate. Angry people speak quickly. Large variations in voice pitch. When angry, the voice has larger pitch variations.

[0074] Calm emotions: The voice intensity is weak. Calm emotions are usually accompanied by low voice volume and energy. The speaking speed is slow. Calm people will speak slowly. The voice has small changes in pitch. When calm, the voice does not have much pitch change.

[0075] Sadness: The voice intensity is weaker. Sadness is usually accompanied by low voice volume and energy. The speech speed is slow. Sad people will speak slowly. The voice pitch changes slightly. When sad, the voice does not have much pitch change.

[0076] Emotional anxiety: The voice intensity is strong. Anxiety is usually accompanied by tension and excitement, and the voice has a certain amount of energy; the speaking speed is fast, and anxious people will speed up their speaking speed; the tone changes greatly. When anxious, the voice has greater tone changes.

[0077] Emotional fear: The voice intensity is strong. Fear is usually accompanied by tension and fear, and the voice has a certain amount of energy; the speaking speed is fast, and fearful people will speak quickly; the tone changes greatly. When people are afraid, the voice has a larger tone change.

[0078] The emotion determination information includes depressed emotion information, calm emotion information and excited emotion information, and sets the emotion influence factor reference thresholds Ee1 and Ee2, where Ee1<Ee2, and α q Substitute the emotional influence factor reference thresholds Ee1 and Ee2 for comparison. When 0<α q <Ee1, generate low emotion information for the detection interval; when Ee1<α q When α < Ee2, the smooth emotion information is generated for the detection interval; when α q When >Ee2, excitement emotion information is generated for the detection interval.

[0079] It should be noted that: compared with the calm emotional information, the excited emotional information has a greater degree of emotional change. If the excited emotional information exists in the detection interval, the person's emotional fluctuation is relatively large, and so on.

[0080] Emotional authenticity analysis module: compares the input real emotion information with the emotion judgment information to generate an emotion authenticity mark;

[0081] According to the emotional reflection observed in the image, real emotional information is input into the emotional judgment module, and the real emotional information includes actual low emotional information, actual medium emotional information and actual high emotional information;

[0082] The emotional authenticity mark includes an emotional authenticity mismatch mark, a low emotional authenticity mark, a flat emotional authenticity mark and an excited emotional authenticity mark. If there is a difference between the input real emotional information and the emotional judgment information, an emotional authenticity mismatch mark is generated for the detection interval; if there is no difference between the input real emotional information and the emotional judgment information, when the corresponding actual low emotional information and the low emotional information are in the same interval, a low emotional authenticity mark is generated; when the corresponding actual medium emotional information and the flat emotional information are in the same interval, a flat emotional authenticity mark is generated; when the corresponding actual high emotional information and the excited emotional information are in the same interval, an excited emotional authenticity mark is generated.

[0083] It should be noted that the detection interval with excitement real mark has more intense emotional fluctuations than the detection interval with calm real mark. Similarly, the detection interval with any of the three types of real marks has authenticity guarantee. When a single detection interval has an emotional authenticity discrepancy mark, the degree of falsification of the detection interval increases.

[0084] Posture change degree analysis module: analyzes and processes the extracted human posture change data to generate posture change degree marks;

[0085] Obtain the elbow movement value Emv, the human body lateral displacement path Lbd, and the portrait proportion change value Pic in n detection intervals, and obtain the motion change factor γ through formula analysis. The acquisition formula of the motion change factor γ is:

[0086] γ q =( +Pic) 0.526

[0087] where γ q > 0, q is the number of video frames included in the identification interval, and q= ,0.526 is the correction ratio value of the motion change factor.

[0088] The posture change degree mark includes a small change mark, a medium change mark and a large change mark, and the motion change factor comparison thresholds Sce1 and Sce2 are set, where Sce1>Sce2>0, and γ q Substitute the motion change factor comparison threshold Sce1 and Sce2 for comparison. When 0<α q <Sce2, a small change mark is generated for the detection interval; when Sce2<α q When α<Sce1, a medium-amplitude change mark is generated for the detection interval; when α q When >Sce1, a large change flag is generated for the detection interval.

[0089] It should be noted that the detection interval with a large change mark has a larger change in motion trajectory than the detection interval with a medium change mark, indicating that the current character state is more excited, and so on.

[0090] Video authenticity analysis and target labeling module: This module combines emotion authenticity labels and gesture change degree labels for analysis, and uses the analysis results to label the authenticity of the identified video.

[0091] The emotion authenticity marks and posture change degree marks in n detection intervals are input into the re-analysis module for integration processing. If a single detection interval has an emotion authenticity mismatch mark and any posture change degree mark, an excitement authenticity mark and a small change mark, or a depression authenticity mark and a large change mark, the detection interval is marked as a forged target;

[0092] When a single detection interval has an exciting true mark and a medium-amplitude change mark, a gentle true mark and a large-amplitude change mark, a gentle true mark and a small-amplitude change mark, or a depressed true mark and a medium-amplitude change mark, the detection interval is marked as a questionable target;

[0093] When a single detection interval has an exciting true mark and a large-amplitude change mark, a flat true mark and a medium-amplitude change mark, or a low true mark and a small-amplitude change mark, the detection interval is marked as a true target.

[0094] It should be noted that the reanalysis module specifically adopts a multimodal deep learning model to combine the features of the acquired inputs through a fusion layer, such as using a fusion layer to perform weighted summation or splicing operations, and finally performs sentiment classification or other tasks.

[0095] Identification result display module: Feedback the video identification results to the display panel based on the true and false target labels.

[0096] Example 2

[0097] See Figure 2 It can be seen that the video identification method based on the human posture determination model described in this embodiment includes the following steps:

[0098] S100, pre-processing the identification video and extracting video feature data, wherein the video feature data includes human voice emotion data and human body posture change data;

[0099] S200, establishing an emotion determination model, and generating emotion determination information by performing emotion determination on the extracted human voice emotion data;

[0100] S300, inputting real emotion information, integrating the real emotion information with emotion determination information, and generating an emotion authenticity mark for the identification video;

[0101] S400, establishing a human motion determination model, performing comparative analysis on the extracted human posture change data, generating a motion change factor, substituting the motion change factor into a change factor threshold for comparison, and generating a posture change degree mark for the identification video;

[0102] S500, establishing a re-analysis module, performing video authenticity analysis by identifying the emotional authenticity mark and the posture change degree mark in the video, and marking the authenticity target of the identified video according to the analysis result;

[0103] S600: Feedback the result of identifying the true and false target markings in the video to the display panel according to the true and false target markings.

[0104] The human voice emotion data includes sound intensity Int, pitch variation amplitude Pv and speech speed Sp, wherein the sound intensity Int represents the energy or volume level of the sound, the pitch variation amplitude Pv represents the pitch variation amplitude of the sound, and the speech speed Sp represents the speaking speed.

[0105] The emotion determination step includes:

[0106] The identification video is divided into n identification intervals according to the video frame, where n is an integer greater than 1. The sound intensity Int, tone variation amplitude Pv and speech speed Sp of the n identification intervals are obtained respectively, and the emotion influence factor α of the n identification intervals is obtained by formulating and analyzing them. q ;

[0107] The emotion determination information includes depressed emotion information, calm emotion information and excited emotion information, and sets the emotion influence factor reference thresholds Ee1 and Ee2, where Ee1<Ee2, and α q Substitute the emotional influence factor reference thresholds Ee1 and Ee2 for comparison. When 0<α q <Ee1, generate low emotion information for the detection interval; when Ee1<α q When α < Ee2, the smooth emotion information is generated for the detection interval; when α q When >Ee2, excitement emotion information is generated for the detection interval.

[0108] The steps for integrating the real emotion information and emotion judgment information in the emotion judgment model are as follows:

[0109] According to the emotional reflection observed in the image, real emotional information is input into the emotional judgment module, and the real emotional information includes actual low emotional information, actual medium emotional information and actual high emotional information;

[0110] If there is a difference between the input real emotion information and the emotion judgment information, an emotion authenticity mismatch mark is generated for the detection interval; if there is no difference between the input real emotion information and the emotion judgment information, when the corresponding actual low emotion information and the low emotion information are in the same interval, a low true mark is generated; when the corresponding actual medium emotion information and the flat emotion information are in the same interval, a flat true mark is generated; when the corresponding actual high emotion information and the excited emotion information are in the same interval, an excited true mark is generated.

[0111] The human body posture change data includes the elbow movement value Emv, the human body lateral displacement path Lbd and the portrait proportion change value Pic, wherein the elbow movement value Emv is the value of the elbow point moving within the detection interval, the human body lateral displacement path Lbd is the displacement value of the human body within the detection interval, and the portrait proportion change value Pic is the proportion change value of the presented portrait within the detection interval.

[0112] The step of generating the posture change degree mark is:

[0113] Obtain the elbow movement value Emv, the human body lateral displacement path Lbd, and the portrait proportion change value Pic in n detection intervals, and obtain the motion change factor γ through formula analysis;

[0114] The posture change degree mark includes a small change mark, a medium change mark and a large change mark, and the motion change factor comparison thresholds Sce1 and Sce2 are set, where Sce1>Sce2>0, and γ q Substitute the motion change factor comparison threshold Sce1 and Sce2 for comparison. When 0<α q <Sce2, a small change mark is generated for the detection interval; when Sce2<α q When α<Sce1, a medium-amplitude change mark is generated for the detection interval; when α q When >Sce1, a large change flag is generated for the detection interval.

[0115] The video authenticity analysis step includes:

[0116] The emotion authenticity marks and posture change degree marks in n detection intervals are input into the re-analysis module for integration processing. If a single detection interval has an emotion authenticity mismatch mark and any posture change degree mark, an excitement authenticity mark and a small change mark, or a depression authenticity mark and a large change mark, the detection interval is marked as a forged target;

[0117] When a single detection interval has an exciting true mark and a medium-amplitude change mark, a gentle true mark and a large-amplitude change mark, a gentle true mark and a small-amplitude change mark, or a depressed true mark and a medium-amplitude change mark, the detection interval is marked as a questionable target;

[0118] When a single detection interval has an exciting true mark and a large-amplitude change mark, a flat true mark and a medium-amplitude change mark, or a low true mark and a small-amplitude change mark, the detection interval is marked as a true target.

[0119] The detection interval target marking results are fed back to the display panel for display.

[0120] The present invention more intuitively displays the dynamic information and emotional changes in the video through the movement trajectory of the human body and the changes in the human voice, and provides a more direct and explainable judgment basis for the authentication result through the analysis of these dynamic features, which helps users better understand and accept the judgment results. Furthermore, multimodal analysis is realized, and the combination of the movement trajectory of the human body and the changes in the human voice can provide more sources of information, thereby enhancing the ability to judge the authenticity of the video. In addition, the accuracy of authentication is improved. The movement trajectory of the human body and the changes in the human voice can be used as key indicators of the authenticity of the video. By analyzing the image characteristics of the human body movement and the changes in the intensity of the sound, it can be more accurately judged whether the video has been tampered with or forged.

[0121] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.

[0122] It should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " as used herein generally indicates an "or" relationship between the associated objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.

[0123] In this application, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, "at least one of a, b, or c" can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.

[0124] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0125] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0126] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, and may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment as needed.

[0127] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0128] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A video identification method based on a human posture determination model, characterized in that: The method comprises the following steps: S100, preprocessing the identification video and extracting video feature data, dividing the identification video into n detection intervals according to the video frame, where n is an integer greater than 1, and the video feature data includes human voice emotion data and human body posture change data; S200, establishing an emotion determination model, and generating emotion determination information by performing emotion determination on the extracted human voice emotion data; S300, inputting real emotion information based on the emotion reflection observed in the image, integrating the real emotion information with the emotion judgment information, and generating an emotion authenticity mark for the identification video; The emotional authenticity mark includes an emotional authenticity mark, a low authenticity mark, a flat authenticity mark and an excited authenticity mark; S400, establishing a human motion determination model, performing comparative analysis on the extracted human posture change data, generating a motion change factor, substituting the motion change factor into a change factor threshold for comparison, and generating a posture change degree mark for the identification video; The posture change degree mark includes a small change mark, a medium change mark and a large change mark; S500, establishing a re-analysis module, performing video authenticity analysis by identifying the emotional authenticity mark and the posture change degree mark in the video, and marking the authenticity target of the identified video according to the analysis result; The authenticity analysis of the video includes: The emotion authenticity marks and posture change degree marks in n detection intervals are input into the re-analysis module for integration processing. If a single detection interval has an emotion authenticity mismatch mark and any posture change degree mark, an excitement authenticity mark and a small change mark, or a depression authenticity mark and a large change mark, the detection interval is marked as a forged target; When a single detection interval has an exciting true mark and a medium-amplitude change mark, a gentle true mark and a large-amplitude change mark, a gentle true mark and a small-amplitude change mark, or a depressed true mark and a medium-amplitude change mark, the detection interval is marked as a questionable target; When a single detection interval has an exciting true mark and a large-amplitude change mark, a flat true mark and a medium-amplitude change mark, or a low true mark and a small-amplitude change mark, the detection interval is marked as a true target; S600: Feedback the result of identifying the true and false target markings in the video to the display panel according to the true and false target markings.

2. The video identification method based on the human posture determination model according to claim 1, characterized in that: The human voice emotion data includes sound intensity Int, pitch variation amplitude Pv and speech speed Sp, wherein the sound intensity Int represents the energy or volume level of the sound, the pitch variation amplitude Pv represents the pitch variation amplitude of the sound, and the speech speed Sp represents the speaking speed; The human body posture change data includes the elbow movement value Emv, the human body lateral displacement path Lbd and the portrait proportion change value Pic, wherein the elbow movement value Emv is the value of the elbow point moving within the detection interval, the human body lateral displacement path Lbd is the displacement value of the human body within the detection interval, and the portrait proportion change value Pic is the proportion change value of the presented portrait within the detection interval.

3. The video identification method based on the human posture determination model according to claim 2, characterized in that: The emotion determination includes: Get the sound intensity Int, tone variation Pv and speech speed Sp of n detection intervals respectively, and perform formula analysis to obtain the emotion influence factor α of n detection intervals q ; The emotion determination information includes depressed emotion information, calm emotion information and excited emotion information, and sets the emotion influence factor reference thresholds Ee1 and Ee2, where Ee1<Ee2, and α q Substitute the emotional influence factor reference thresholds Ee1 and Ee2 for comparison. When 0<α q <Ee1, generate low emotion information for the detection interval; when Ee1<α q When α < Ee2, the smooth emotion information is generated for the detection interval; when α q When >Ee2, excitement emotion information is generated for the detection interval.

4. The video identification method based on the human posture determination model according to claim 3 is characterized in that: The steps for integrating the real emotion information and the emotion judgment information are as follows: The real emotion information includes actual low emotion information, actual medium emotion information and actual high emotion information; If there is a difference between the input real emotion information and the emotion judgment information, an emotion authenticity mismatch mark is generated for the detection interval; if there is no difference between the input real emotion information and the emotion judgment information, when the corresponding actual low emotion information and the low emotion information are in the same interval, a low true mark is generated; when the corresponding actual medium emotion information and the flat emotion information are in the same interval, a flat true mark is generated; when the corresponding actual high emotion information and the excited emotion information are in the same interval, an excited true mark is generated.

5. The video identification method based on the human posture determination model according to claim 4 is characterized in that: The step of generating the posture change degree mark is: Obtain the elbow movement value Emv, the human body lateral displacement path Lbd and the portrait proportion change value Pic in n detection intervals, and obtain the motion change factor γ through formula analysis q ; Set the motion change factor comparison thresholds Sce1 and Sce2, where Sce1>Sce2>0, and set γ q Substitute the motion change factor comparison threshold Sce1 and Sce2 for comparison. When 0<γ q <Sce2, a small change mark is generated for the detection interval; when Sce2<γ q When <Sce1, a medium-amplitude change mark is generated for the detection interval; when γ q When >Sce1, a large change flag is generated for the detection interval.

6. A video identification system based on a human posture determination model, used to implement the method according to any one of claims 1 to 5, characterized in that: include: Video preprocessing module: preprocess the identification video to extract human voice emotion data and human posture change data; Emotional information processing module: integrates and analyzes the extracted human voice emotion data to generate emotion judgment information; Emotional authenticity analysis module: compares the input real emotion information with the emotion judgment information to generate an emotion authenticity mark; Posture change degree analysis module: analyzes and processes the extracted human posture change data to generate posture change degree marks; Video authenticity analysis and target labeling module: This module combines emotion authenticity labels and gesture change degree labels for analysis, and uses the analysis results to label the authenticity of the identified video. Identification result display module: Feedback the video identification results to the display panel based on the true and false target labels.

7. The video identification system based on the human posture determination model according to claim 6, characterized in that: The emotional authenticity mark includes an emotional authenticity mismatch mark, a low-level authenticity mark, a flat-level authenticity mark, and an excited-level authenticity mark. The steps for generating the emotional authenticity mark are: The human voice emotion data includes sound intensity Int, pitch variation amplitude Pv and speech speed Sp, wherein the sound intensity Int represents the energy or volume level of the sound, the pitch variation amplitude Pv represents the pitch variation amplitude of the sound, and the speech speed Sp represents the speaking speed; The identification video is divided into n detection intervals according to the video frame, where n is an integer greater than 1. The sound intensity Int, tone change amplitude Pv and speech speed Sp of the n detection intervals are obtained respectively, and the emotion influence factor α of the n detection intervals is obtained by formula analysis. q , set the reference thresholds of emotion influence factors Ee1 and Ee2, where Ee1<Ee2, and set α q Substitute the emotional influence factor reference thresholds Ee1 and Ee2 for comparison. When 0<α q <Ee1, generate low emotion information for the detection interval; when Ee1<α q When α < Ee2, the smooth emotion information is generated for the detection interval; when α q >Ee2, generate excitement emotion information for the detection interval; According to the emotional reflection observed in the image, real emotional information is input into the emotion judgment module, and the real emotional information includes actual low emotional information, actual medium emotional information and actual high emotional information; If there is a difference between the input real emotion information and the emotion judgment information, an emotion authenticity mismatch mark is generated for the detection interval; if there is no difference between the input real emotion information and the emotion judgment information, when the corresponding actual low emotion information and the low emotion information are in the same interval, a low true mark is generated; when the corresponding actual medium emotion information and the flat emotion information are in the same interval, a flat true mark is generated; when the corresponding actual high emotion information and the excited emotion information are in the same interval, an excited true mark is generated.

8. The video identification system based on the human posture determination model according to claim 7, characterized in that: The steps to generate the posture change degree mark are: The human body posture change data includes the elbow movement value Emv, the human body lateral displacement path Lbd and the portrait proportion change value Pic, wherein the elbow movement value Emv is the value of the elbow point moving within the detection interval, the human body lateral displacement path Lbd is the displacement value of the human body within the detection interval, and the portrait proportion change value Pic is the proportion change value of the presented portrait within the detection interval; Obtain the elbow movement value Emv, the human body lateral displacement path Lbd and the portrait proportion change value Pic in n detection intervals, and obtain the motion change factor γ through formula analysis q , set the motion change factor comparison thresholds Sce1 and Sce2, where Sce1>Sce2>0, and set γ q Substitute the motion change factor comparison threshold Sce1 and Sce2 for comparison. When 0<γ q <Sce2, a small change mark is generated for the detection interval; when Sce2<γ q When <Sce1, a medium-amplitude change mark is generated for the detection interval; when γ q When >Sce1, a large change flag is generated for the detection interval.

9. The video identification system based on the human posture determination model according to claim 8, characterized in that: The steps for combining and analyzing the emotional authenticity marker and the posture change degree marker are as follows: The emotion authenticity marks and posture change degree marks in n detection intervals are input into the re-analysis module for integration processing. If a single detection interval has an emotion authenticity mismatch mark and any posture change degree mark, an excitement authenticity mark and a small change mark, or a depression authenticity mark and a large change mark, the detection interval is marked as a forged target; When a single detection interval has an exciting true mark and a medium-amplitude change mark, a gentle true mark and a large-amplitude change mark, a gentle true mark and a small-amplitude change mark, or a depressed true mark and a medium-amplitude change mark, the detection interval is marked as a questionable target; When a single detection interval has an exciting true mark and a large-amplitude change mark, a flat true mark and a medium-amplitude change mark, or a low true mark and a small-amplitude change mark, the detection interval is marked as a true target.

Citation Information

Patent Citations

  • Identification method used for fake high code rate video

    CN103327320A

  • Intelligent learning system and method based on emotion recognition

    CN111476217A

  • System and Method for Detecting Fabricated Videos

    US20220138472A1