Intelligent detection system and method for multimodal natural language and abnormal signs analysis

By constructing a probability binary tree and calculating multiple probability coefficients, dynamically adjusting the weight of modal data, the problem of insufficient modal weight adjustment in the existing technology is solved, and the accuracy and robustness of lie detection are improved.

CN119564207BActive Publication Date: 2025-05-13SHANGHAI YUFENG ELECTRONIC INFORMATION TECH DEV CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510131008.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-13
Estimated Expiration
2045-02-06

AI Technical Summary

Technical Problem

The prior art cannot dynamically adjust the attention weight between data of different modes in lie detection, resulting in excessive attention or neglect of certain modes in different users or situations, affecting the detection effect and accuracy.

Method used

By constructing a probability binary tree and calculating multiple probability coefficients, the weights of each modal data are dynamically adjusted to realize weighted comprehensive analysis of dialogue text information, physiological detection information and facial image information.

Benefits of technology

It effectively improves the accuracy and robustness of lie detection and ensures the accuracy of detection results in different users and scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119564207B_ABST
    Figure CN119564207B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of lie analysis. The present invention discloses an intelligent detection system and method for multimodal natural language and abnormal physical sign analysis. The method comprises first traversing M groups of sub-text information, mapping the sub-text information into a data space, thereby constructing a probability binary tree, determining a first probability coefficient based on the probability binary tree, determining a second probability coefficient based on sub-detection information corresponding to the sub-text information, and determining a third probability coefficient based on sub-image information corresponding to the sub-text information. The present invention solves the problem that the prior art cannot dynamically adjust the attention weights of different modalities by performing weighted comprehensive analysis on data of different modalities. By constructing a probability binary tree and calculating multiple probability coefficients, the weight of each modality can be dynamically adjusted according to the actual situation, thereby avoiding the situation where certain modalities are over-emphasized or neglected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of lie analysis, and more specifically, to an intelligent detection system and method for multimodal natural language and abnormal physical sign analysis. Background Art

[0002] In the prior art, polygraphing is performed using a multi-channel recorder, which can simultaneously record heart rate, blood pressure, respiratory rate, and skin conductivity. With the advancement of computer technology, the emergence of computer-assisted polygraph systems has made the measurement and analysis of physiological data more precise, significantly improving the efficiency and accuracy of polygraphing. After entering the digital age, data processing has become more intelligent and automated, and the rapid development of machine learning and artificial intelligence technologies has further promoted the innovation of polygraphing technology. For example, the prior art has already appeared in polygraphing by combining multimodal data with machine learning technology.

[0003] For example, the Chinese patent application with publication number CN112329438A provides an automatic lie detection method and system based on domain adversarial training. The patent extracts text, audio and facial feature representations through multimodal feature extraction to comprehensively analyze lying behavior. Then, an adaptive attention mechanism is used to fuse multimodal features, and a bidirectional recurrent neural network is used to capture contextual information in the conversation to improve the accuracy of lie detection. Finally, a domain adversarial training method is used to extract speaker-independent lie features, and the lie level is predicted through the trained classifier.

[0004] Although existing technologies have effectively improved the accuracy of lie detection through multimodal data and machine learning methods, these technologies still have a limitation, namely, they are unable to dynamically adjust the attention weights between different modal data based on the actual user's characteristics or specific scenarios. This results in certain modalities being over-emphasized or ignored under different users or situations, thus affecting the ultimate effect and accuracy of lie detection.

[0005] In view of this, the present invention proposes an intelligent detection system and method for multimodal natural language and abnormal vital signs analysis to solve the above problems. Summary of the invention

[0006] In order to overcome the above-mentioned defects of the prior art, the present invention provides an intelligent detection system and method for multimodal natural language and abnormal signs analysis.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] First, an intelligent detection method for multimodal natural language and abnormal signs analysis is provided, including:

[0009] Acquire the conversation text information, physiological detection information and facial image information of the subject during the detection process, wherein the conversation text information includes M groups of sub-text information, each group of sub-text information has corresponding sub-detection information and sub-image information, the sub-detection information belongs to the physiological detection information, and the sub-image information belongs to the facial image information;

[0010] Traversing M groups of sub-text information, mapping the sub-text information into the data space, thereby constructing a probability binary tree, and determining a first probability coefficient based on the probability binary tree;

[0011] Determine a second probability coefficient based on the sub-detection information corresponding to the sub-text information, determine a third probability coefficient based on the sub-image information corresponding to the sub-text information, and construct a weight allocation sequence according to the first probability coefficient, the second probability coefficient, and the third probability coefficient;

[0012] The detection result is determined based on the conversation text information, physiological detection information, facial image information and weight allocation sequence.

[0013] Furthermore, the method of obtaining the sub-detection information and the sub-image information corresponding to each group of sub-text information includes:

[0014] The start time of the Nth group of sub-text information is taken as the first target time, and the start time of the N+1th group of sub-text information is taken as the second target time. The target time period is constructed with the first target time and the second target time. The physiological detection information is segmented according to the target time period to obtain S groups of sub-detection information. The facial image information is segmented according to the target time period to obtain S groups of sub-image information. The sub-text information, sub-detection information and sub-image information within the target time period correspond to each other, and 0<N<S.

[0015] Furthermore, the method of mapping the subtext information into the data space to construct a probability binary tree includes:

[0016] The sub-text information is used as the root node of the probability binary tree, and whether the preset polygraph keywords appear in each layer of the sub-text information is used as the cutting feature. Based on the cutting feature, the root node is divided into two left and right child nodes, and each child node is recursively cut until each child node cannot be cut any further, thereby constructing a probability binary tree.

[0017] Furthermore, the method for determining the first probability coefficient based on the probability binary tree includes:

[0018] The real-time height value of each sub-binary tree in the probability binary tree is obtained according to the depth-first search algorithm, and a weighted sum is performed based on the real-time height value of each sub-binary tree to obtain a comprehensive height value, which is used as the first probability coefficient.

[0019] Furthermore, the sub-detection information includes real-time respiratory rate value, real-time blood pressure value and real-time pulse rate value. The method for determining the second probability coefficient based on the sub-detection information corresponding to the sub-text information includes: calculating the second probability coefficient based on the real-time respiratory rate value, real-time blood pressure value and real-time pulse rate value.

[0020] Furthermore, the sub-image information is a close-up image of an eye, and the method for determining the third probability coefficient based on the sub-image information corresponding to the sub-text information includes:

[0021] Based on computer vision, first eye information and second eye information in the eye close-up image are obtained, the eye aspect ratio is calculated based on the first eye information, the absolute value of the difference between the eye aspect ratio and a preset standard aspect ratio is calculated, and the pupil area ratio is calculated based on the second eye information. The absolute value of the difference and the pupil area ratio are weightedly summed to obtain a third probability coefficient.

[0022] Furthermore, the first eye information includes the inner coordinates of the upper eyelid, the outer coordinates of the upper eyelid, the inner coordinates of the lower eyelid, the outer coordinates of the lower eyelid, the inner corner coordinates and the outer corner coordinates. The method for calculating the eye aspect ratio based on the first eye information includes: calculating the eye aspect ratio according to the inner coordinates of the upper eyelid, the outer coordinates of the upper eyelid, the inner coordinates of the lower eyelid, the outer coordinates of the lower eyelid, the inner corner coordinates and the outer corner coordinates.

[0023] Furthermore, the method of constructing a weight distribution sequence according to the first probability coefficient, the second probability coefficient and the third probability coefficient includes:

[0024] Traverse M groups of sub-text information and the corresponding sub-detection information and sub-image information, determine the first number of times that the first probability coefficient is greater than a preset first coefficient threshold, determine the second number of times that the second probability coefficient is greater than a preset second coefficient threshold, determine the third number of times that the third probability coefficient is greater than a preset third coefficient threshold, calculate the sum of the first number, the second number and the third number as the total number of times, take the ratio of the first number to the total number of times as the first weight of the conversation text information, take the ratio of the second number to the total number of times as the second weight of the physiological detection information, take the ratio of the third number to the total number of times as the third weight of the facial image information, and take the first weight, the second weight and the third weight as a weight allocation sequence.

[0025] Further, the method for determining the test result includes:

[0026] Input the conversation text information, physiological detection information, facial image information, and weight distribution sequence into the pre-built machine learning model to obtain the detection results;

[0027] Methods for building machine learning models include:

[0028] Acquire a sample data set, wherein the sample data set includes historical conversation text information, historical physiological detection information, historical facial image information, historical weight distribution sequence, and historical detection results;

[0029] Divide the sample data set into a sample training set and a sample test set, and build a regression network;

[0030] The historical conversation text information, historical physiological detection information, historical facial image information and historical weight distribution sequence in the sample training set are used as input data of the regression network, and the historical detection results in the sample training set are used as output data of the regression network. The regression network is trained to obtain an initial regression network for predicting the initial detection results.

[0031] The initial regression network is tested using a sample test set, and the initial regression network with an output smaller than a preset error value is used as a machine learning model.

[0032] In a second aspect, a multimodal natural language and abnormal sign analysis intelligent detection system is provided, which is used to implement the above-mentioned multimodal natural language and abnormal sign analysis intelligent detection method, including:

[0033] Data acquisition module: used to acquire the conversation text information, physiological detection information and facial image information of the subject during the detection process, the conversation text information includes M groups of sub-text information, each group of sub-text information has corresponding sub-detection information and sub-image information, the sub-detection information belongs to the physiological detection information, and the sub-image information belongs to the facial image information;

[0034] The first processing module is used to traverse the M groups of sub-text information, map the sub-text information to the data space, thereby constructing a probability binary tree, and determine the first probability coefficient based on the probability binary tree;

[0035] The second processing module determines a second probability coefficient based on the sub-detection information corresponding to the sub-text information, determines a third probability coefficient based on the sub-image information corresponding to the sub-text information, and constructs a weight allocation sequence according to the first probability coefficient, the second probability coefficient and the third probability coefficient;

[0036] Detection module: Determines the detection result based on the conversation text information, physiological detection information, facial image information and weight allocation sequence.

[0037] Compared with the prior art, the present invention has the following beneficial effects:

[0038] The present invention first traverses M groups of sub-text information, maps the sub-text information into a data space, thereby constructing a probability binary tree, determining a first probability coefficient based on the probability binary tree, determining a second probability coefficient based on sub-detection information corresponding to the sub-text information, and determining a third probability coefficient based on sub-image information corresponding to the sub-text information. A weight allocation sequence is constructed according to the first probability coefficient, the second probability coefficient, and the third probability coefficient. A detection result is determined based on the conversation text information, the physiological detection information, the facial image information, and the weight allocation sequence. The present invention solves the problem that the prior art cannot dynamically adjust the attention weights of different modalities by performing a weighted comprehensive analysis on different modal data (conversation text information, physiological detection information, and facial image information). By constructing a probability binary tree and calculating multiple probability coefficients, the weight of each modality can be dynamically adjusted according to the actual situation, thereby avoiding the situation where certain modalities are over-emphasized or neglected. The present invention effectively improves the accuracy and robustness of lie detection, and ensures the accuracy of detection results under different users and scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is a flow chart of the intelligent detection method of multimodal natural language and abnormal signs analysis in the present invention;

[0040] Figure 2 It is a structural schematic diagram of the intelligent detection system for multimodal natural language and abnormal physical sign analysis in the present invention;

[0041] Figure 3 Schematic diagram of a probabilistic binary tree in the present invention;

[0042] Figure 4 It is a schematic diagram of a flow chart of determining a third probability coefficient based on sub-image information corresponding to sub-text information in the present invention.

[0043] Reference numerals:

[0044] 10, root node; 20, first child node; 30, second child node; 40, third child node; 50, fourth child node; 60, fifth child node. DETAILED DESCRIPTION

[0045] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0046] Example 1

[0047] See also Figure 1As shown, this embodiment discloses an intelligent detection method for multimodal natural language and abnormal physical sign analysis, including:

[0048] S10: Acquire the conversation text information, physiological detection information and facial image information of the subject during the detection process, wherein the conversation text information includes M groups of sub-text information, each group of sub-text information has corresponding sub-detection information and sub-image information, the sub-detection information belongs to the physiological detection information, and the sub-image information belongs to the facial image information;

[0049] In this embodiment, the detection process is usually in the form of a question-and-answer session between the tester and the subject. During this process, the conversation text information refers to the conversation content generated when the tester and the subject interact. The physiological detection information is the physiological data recorded by the subject during the polygraph process, such as heart rate, breathing rate, etc.; and the facial image information includes the facial images taken by the subject during the detection process. These images can be partial images (such as close-ups of the eyes or mouth) or complete images (such as overall facial contour images).

[0050] It should be added that the polygraph process of question and answer between the tester and the subject is completed based on the dialogue information. Therefore, each set of dialogue information corresponds to a set of sub-text information. The above-mentioned sub-detection information can be the physiological detection information of the subject during the process of the tester and the subject completing a set of dialogue information. Similarly, the above-mentioned sub-image information can be a partial image or a complete image of the subject's face during the process of the tester and the subject completing a set of dialogue information.

[0051] The method of obtaining the sub-detection information and the sub-image information corresponding to each set of sub-text information includes:

[0052] The start time of the Nth group of sub-text information is taken as the first target time, and the start time of the N+1th group of sub-text information is taken as the second target time. The target time period is constructed with the first target time and the second target time. The physiological detection information is segmented according to the target time period to obtain S groups of sub-detection information. The facial image information is segmented according to the target time period to obtain S groups of sub-image information. The sub-text information, sub-detection information and sub-image information within the target time period correspond to each other, and 0<N<S.

[0053] It should be noted that the starting time of the Nth group of sub-text information in the above refers to the starting time point of its corresponding dialogue text information, that is, it is the moment when the dialogue begins when the subject and the tester ask and answer questions. Similarly, the starting time of the N+1th group of sub-text information refers to the starting time point of its corresponding dialogue text information.

[0054] It should be added that the detection process is usually in the form of a question-and-answer session between the tester and the subject. Therefore, when the tester and the subject complete a conversation, the subject not only has physiological reactions and micro-expressions during the conversation, but also has physiological reactions or micro-expressions after the conversation. These are all related to the conversation. Therefore, in this embodiment, the start time of the Nth group of sub-text information is used as the first target time, and the start time of the N+1th group of sub-text information is used as the second target time. The target time period is constructed with the first target time and the second target time. This can not only capture the subject's immediate physiological reactions and micro-expressions during the conversation, but also include physiological reactions and facial expression changes caused by thinking, emotional fluctuations or memories after the conversation. This comprehensive capture of the time period helps to more accurately identify lies or inconsistent language. By extending the division of the time period, the physiological detection information and facial image information can fully correspond to the entire conversation process and its subsequent reactions, thereby enhancing the ability to capture subtle physiological changes and improving the sensitivity of detection.

[0055] S20: traverse the M groups of sub-text information, map the sub-text information to the data space, thereby constructing a probability binary tree, and determine a first probability coefficient based on the probability binary tree;

[0056] In this embodiment, the data space refers to a multi-dimensional or multi-level data structure for storing, organizing and processing different types of information. Mapping sub-text information to the data space refers to converting the features of each group of sub-text information (such as keywords or sentences in the text) into a data point in the data space. In this way, the sub-text information can be effectively represented in the data space.

[0057] Methods for mapping subtext information into data space to construct a probabilistic binary tree include:

[0058] The sub-text information is used as the root node of the probability binary tree, and whether the preset polygraph keywords appear in each layer of the sub-text information is used as the cutting feature. Based on the cutting feature, the root node is divided into two left and right child nodes, and each child node is recursively cut until each child node cannot be cut any further, thereby constructing a probability binary tree.

[0059] It should be noted that each layer in the sub-text information refers to the layer-by-layer structure or progressive analysis level of the sub-text information. For example, when the sub-text information is a whole conversation or a long text, each layer can represent each answer of the subject in the conversation. In this case, each layer corresponds to a sentence or phrase in the text, and detects whether it contains preset lie detection keywords. Based on whether these keywords appear in each sentence, they can be used as cutting features to construct a probabilistic binary tree.

[0060] The above implementation method is described below in combination with specific application scenarios:

[0061] Tester: Did you go to the company's warehouse last night?

[0062] Subject: Um…yeah, I went to the warehouse around 7pm, but I only went in to get some documents;

[0063] Tester: When you entered the warehouse, did you see anything unusual?

[0064] Subject: No, I didn't see anything special, everything in the warehouse was normal;

[0065] Tester: Did you take anything out?

[0066] Subject: I'm not sure, but I don't think I took anything out, just some documents;

[0067] Tester: Did you tell anyone else that you had been to the warehouse?

[0068] Subject: Actually, I didn’t tell anyone because it was late and I just wanted to leave quickly;

[0069] Tester: Do you think you have done anything wrong?

[0070] Subject: Well, I'm not sure, I think I was just doing a normal job and did nothing wrong.

[0071] Then each layer in the subtext information can be:

[0072] First child node 20: Um... yes, I went to the warehouse around 7pm, but I only went in to get some documents;

[0073] Second subnode 30: Subject: No, I didn’t see anything special, everything in the warehouse was normal;

[0074] Third subnode 40: Subject: I’m not sure, but I think I didn’t take anything out, just some documents;

[0075] Fourth subnode 50: Subject: Actually, I didn’t tell anyone because it was already late and I just wanted to leave quickly;

[0076] Fifth subnode 60: Subject: Well, I’m not sure. I think I was just doing a normal job and did nothing wrong.

[0077] In the above, the preset lie detection keywords include but are not limited to "not very clear", "probably", "not very sure", etc., so the constructed probability binary tree is as follows Figure 3 As shown, Figure 3 The root node 10, the first child node 20, the second child node 30, the third child node 40, the fourth child node 50 and the fifth child node 60 are shown. It can be understood that Figure 3 The root node 10 is the sub-text information, and the first sub-node 20, the second sub-node 30, the third sub-node 40, the fourth sub-node 50 and the fifth sub-node 60 are each layer in the sub-text information.

[0078] The method for determining the first probability coefficient based on the probability binary tree includes:

[0079] The real-time height value of each sub-binary tree in the probability binary tree is obtained according to the depth-first search algorithm, and a weighted sum is performed based on the real-time height value of each sub-binary tree to obtain a comprehensive height value, which is used as the first probability coefficient.

[0080] It should be noted that the depth-first search algorithm refers to an algorithm for traversing a tree structure, which can be used to recursively obtain the height of each sub-binary tree by starting from the root node, traversing each path of the tree until the leaf node, and then backtracking and recording the depth of each subtree. Figure 3 The number of sub-binary trees is 2, and the weights corresponding to each sub-binary tree can be pre-set based on expert experience, for example Figure 3 In the example, the weight corresponding to one of the sub-binary trees representing that the subject is lying is set to 0.7, and the weight corresponding to the other sub-binary tree representing that the subject is not lying is set to 0.3. Therefore, in this embodiment, the larger the comprehensive height value, the greater the probability that the subject is lying, that is, the larger the first probability coefficient, the greater the probability that the subject is lying.

[0081] In this embodiment, the probability binary tree is recursively traversed through the depth-first search algorithm, and the real-time height value of each sub-binary tree can be accurately obtained. This depth calculation is combined with multi-level dialogue analysis to ensure that the complexity of each sub-binary tree is fully evaluated. A weighted sum is performed based on the real-time height values ​​of the sub-binary trees to obtain a comprehensive height value, which provides an accurate basis for calculating the first probability coefficient. Through this combination, the probability of the subject lying can ultimately be more accurately evaluated.

[0082] S30: determining a second probability coefficient based on the sub-detection information corresponding to the sub-text information, determining a third probability coefficient based on the sub-image information corresponding to the sub-text information, and constructing a weight allocation sequence according to the first probability coefficient, the second probability coefficient, and the third probability coefficient;

[0083] In this embodiment, the physiological detection information includes but is not limited to real-time respiratory rate value, real-time blood pressure value, and real-time pulse rate value, etc., then similarly, the sub-detection information includes but is not limited to real-time respiratory rate value, real-time blood pressure value, and real-time pulse rate value, etc.

[0084] The method for determining the second probability coefficient based on the sub-detection information corresponding to the sub-text information includes:

[0085] ;

[0086] Among them, SPC is the second probability coefficient, is the real-time blood pressure value, is the standard blood pressure value, is the real-time respiratory rate value, is the standard respiratory rate value, is the real-time pulse frequency value, is the standard pulse frequency value, is the inverse cotangent function, So The logarithmic function with base , is a natural constant.

[0087] In this embodiment, taking the real-time respiratory rate value as an example, when the subject is lying during a conversation, the real-time respiratory rate value deviates more from the standard respiratory rate value. The same is true for the real-time blood pressure value and the real-time pulse rate value. Therefore, it can be seen from the above content that the larger the second probability coefficient is, the greater the probability that the subject is lying.

[0088] In this embodiment, the facial image information is taken as an example of a close-up image of an eye. Similarly, the sub-image information is also a close-up image of an eye. The method for determining the third probability coefficient based on the sub-image information corresponding to the sub-text information includes:

[0089] Based on computer vision, first eye information and second eye information in the eye close-up image are obtained, the eye aspect ratio is calculated based on the first eye information, the absolute value of the difference between the eye aspect ratio and a preset standard aspect ratio is calculated, and the pupil area ratio is calculated based on the second eye information. The absolute value of the difference and the pupil area ratio are weightedly summed to obtain a third probability coefficient.

[0090] Specifically, the first eye information includes the inner coordinates of the upper eyelid, the outer coordinates of the upper eyelid, the inner coordinates of the lower eyelid, the outer coordinates of the lower eyelid, the inner corner of the eye coordinates and the outer corner of the eye coordinates. It can be understood that any point in the eye close-up image is used as the origin of the coordinate system to obtain the inner coordinates of the upper eyelid, the outer coordinates of the upper eyelid, etc. This embodiment does not go into details about this.

[0091] The method for calculating the eye aspect ratio based on the first eye information includes:

[0092] ;

[0093] Where EAR is the aspect ratio of the eye, is the inner eye corner coordinate, is the coordinate of the inner side of the upper eyelid, is the coordinate of the inner side of the lower eyelid, is the outer eye corner coordinate, is the outer coordinate of the lower eyelid, is the outer coordinate of the upper eyelid, is the Euclidean distance function.

[0094] Specifically, the second eye information includes at least a real-time pupil area, and the method for calculating the pupil area ratio based on the second eye information includes taking the ratio of the real-time pupil area to the standard pupil area as the pupil area ratio.

[0095] In this embodiment, the eye aspect ratio represents the degree of openness of the subject's eyes during the lie detection process. When the eye aspect ratio is closer to 1, it means that the eyes are more open. On the contrary, when the eye aspect ratio is closer to 0, it means that the eyes are more closed. When the subject feels nervous or lies, the degree of openness of the eyes changes. In this embodiment, the standard aspect ratio can be set to 0.5. The larger the absolute value of the difference, the more nervous the subject feels. Similarly, the pupil will dilate when the subject is nervous. Therefore, it can be seen from the above that in this embodiment, the third probability coefficient is positively correlated with the degree of nervousness of the subject, that is, the larger the third probability coefficient, the greater the probability that the subject is lying.

[0096] The method for constructing a weight allocation sequence according to the first probability coefficient, the second probability coefficient and the third probability coefficient includes:

[0097] Traverse M groups of sub-text information and the corresponding sub-detection information and sub-image information, determine the first number of times that the first probability coefficient is greater than a preset first coefficient threshold, determine the second number of times that the second probability coefficient is greater than a preset second coefficient threshold, determine the third number of times that the third probability coefficient is greater than a preset third coefficient threshold, calculate the sum of the first number, the second number and the third number as the total number of times, take the ratio of the first number to the total number of times as the first weight of the conversation text information, take the ratio of the second number to the total number of times as the second weight of the physiological detection information, take the ratio of the third number to the total number of times as the third weight of the facial image information, and take the first weight, the second weight and the third weight as a weight allocation sequence.

[0098] It should be noted that each of the M groups of sub-text information has a corresponding sub-detection information and sub-image information. Therefore, in this embodiment, there are M first probability coefficients, M second probability coefficients and M third probability coefficients. Taking the first probability coefficient as an example, the larger the first probability coefficient is, the greater the probability that the subject is lying. Therefore, when the first probability coefficient is greater than the preset first coefficient threshold, a statistics is performed. After the statistics are completed, the first number is obtained. The first number represents the degree of probability of the subject lying based on the conversation text information. Similarly, the second number represents the degree of probability of the subject lying based on the physiological detection information. The third number represents the degree of probability of the subject lying based on the facial image information.

[0099] Therefore, in this embodiment, the first weight, the second weight and the third weight are respectively assigned to the conversation text information, the physiological detection information and the facial image information. The purpose of this is to perform a weighted comprehensive analysis on the probability of the subject lying according to different types of data, so as to achieve more accurate and comprehensive lie detection. Therefore, due to the influence of factors such as cultural background, psychological quality and individual differences of the subjects, the performance and reliability of each type of data in lie detection are different. For example, some people can control their expressions when lying, resulting in the reliability of facial image information cannot be guaranteed. Therefore, the weight of facial image information needs to be appropriately reduced.

[0100] S40: Determine a detection result based on the conversation text information, the physiological detection information, the facial image information, and the weight allocation sequence.

[0101] Methods for determining test results include:

[0102] The conversation text information, physiological detection information, facial image information, and weight assignment sequence are input into the pre-built machine learning model to obtain the detection results.

[0103] Methods for building machine learning models include:

[0104] Acquire a sample data set, wherein the sample data set includes historical conversation text information, historical physiological detection information, historical facial image information, historical weight distribution sequence, and historical detection results;

[0105] Divide the sample data set into a sample training set and a sample test set, and build a regression network;

[0106] The historical conversation text information, historical physiological detection information, historical facial image information and historical weight distribution sequence in the sample training set are used as input data of the regression network, and the historical detection results in the sample training set are used as output data of the regression network. The regression network is trained to obtain an initial regression network for predicting the initial detection results.

[0107] The initial regression network is tested using a sample test set, and the initial regression network with an output smaller than a preset error value is used as a machine learning model. The initial regression network is a deep neural network model.

[0108] It is understandable that the test result may be a probability value of the subject lying during the test. The sample data set mentioned above is obtained in advance by those skilled in the art through experiments, and this embodiment will not go into details.

[0109] In this embodiment, the conversation text information, physiological detection information and facial image information of the subject in the detection process are first obtained, M groups of sub-text information are traversed, and the sub-text information is mapped to the data space, so as to construct a probability binary tree, determine the first probability coefficient based on the probability binary tree, determine the second probability coefficient based on the sub-detection information corresponding to the sub-text information, determine the third probability coefficient based on the sub-image information corresponding to the sub-text information, construct a weight allocation sequence according to the first probability coefficient, the second probability coefficient and the third probability coefficient, and determine the detection result based on the conversation text information, physiological detection information, facial image information and the weight allocation sequence. Then, this embodiment solves the problem that the prior art cannot dynamically adjust the attention weights of different modalities by weighted comprehensive analysis of different modal data (conversation text information, physiological detection information and facial image information). By constructing a probability binary tree and calculating multiple probability coefficients, the weights of each modality can be dynamically adjusted according to the actual situation, thereby avoiding the situation where some modalities are over-emphasized or ignored. The present invention effectively improves the accuracy and robustness of lie detection and ensures the accuracy of detection results under different users and scenarios.

[0110] Example 2

[0111] See also Figure 2 As shown, based on the same inventive concept, this embodiment discloses an intelligent detection system for multimodal natural language and abnormal physical sign analysis. For details not provided in this embodiment, please refer to the description of the relevant parts in Embodiment 1. The system includes:

[0112] Data acquisition module: used to acquire the conversation text information, physiological detection information and facial image information of the subject during the detection process, the conversation text information includes M groups of sub-text information, each group of sub-text information has corresponding sub-detection information and sub-image information, the sub-detection information belongs to the physiological detection information, and the sub-image information belongs to the facial image information;

[0113] In this embodiment, the detection process is usually in the form of a question-and-answer session between the tester and the subject. During this process, the conversation text information refers to the conversation content generated when the tester and the subject interact. The physiological detection information is the physiological data recorded by the subject during the polygraph process, such as heart rate, breathing rate, etc.; and the facial image information includes the facial images taken by the subject during the detection process. These images can be partial images (such as close-ups of the eyes or mouth) or complete images (such as overall facial contour images).

[0114] The method of obtaining the sub-detection information and the sub-image information corresponding to each set of sub-text information includes:

[0115] The start time of the Nth group of sub-text information is taken as the first target time, and the start time of the N+1th group of sub-text information is taken as the second target time. The target time period is constructed with the first target time and the second target time. The physiological detection information is segmented according to the target time period to obtain S groups of sub-detection information. The facial image information is segmented according to the target time period to obtain S groups of sub-image information. The sub-text information, sub-detection information and sub-image information within the target time period correspond to each other, and 0<N<S.

[0116] The first processing module is used to traverse the M groups of sub-text information, map the sub-text information to the data space, thereby constructing a probability binary tree, and determine the first probability coefficient based on the probability binary tree;

[0117] In this embodiment, the data space refers to a multi-dimensional or multi-level data structure for storing, organizing and processing different types of information. Mapping sub-text information to the data space refers to converting the features of each group of sub-text information (such as keywords or sentences in the text) into a data point in the data space. In this way, the sub-text information can be effectively represented in the data space.

[0118] Methods for mapping subtext information into data space to construct a probabilistic binary tree include:

[0119] The sub-text information is used as the root node of the probability binary tree, and whether the preset polygraph keywords appear in each layer of the sub-text information is used as the cutting feature. Based on the cutting feature, the root node is divided into two left and right child nodes, and each child node is recursively cut until each child node cannot be cut any further, thereby constructing a probability binary tree.

[0120] The method for determining the first probability coefficient based on the probability binary tree includes:

[0121] The real-time height value of each sub-binary tree in the probability binary tree is obtained according to the depth-first search algorithm, and a weighted sum is performed based on the real-time height value of each sub-binary tree to obtain a comprehensive height value, which is used as the first probability coefficient.

[0122] The second processing module determines a second probability coefficient based on the sub-detection information corresponding to the sub-text information, determines a third probability coefficient based on the sub-image information corresponding to the sub-text information, and constructs a weight allocation sequence according to the first probability coefficient, the second probability coefficient and the third probability coefficient;

[0123] In this embodiment, the physiological detection information includes but is not limited to real-time respiratory rate value, real-time blood pressure value, and real-time pulse rate value, etc., then similarly, the sub-detection information includes but is not limited to real-time respiratory rate value, real-time blood pressure value, and real-time pulse rate value, etc.

[0124] The method for determining the second probability coefficient based on the sub-detection information corresponding to the sub-text information includes:

[0125] ;

[0126] Among them, SPC is the second probability coefficient, is the real-time blood pressure value, is the standard blood pressure value, is the real-time respiratory rate value, is the standard respiratory rate value, is the real-time pulse frequency value, is the standard pulse frequency value, is the inverse cotangent function, So The logarithmic function with base , is a natural constant.

[0127] like Figure 4 As shown, in this embodiment, the facial image information is taken as an eye close-up image for example. Similarly, the sub-image information is also an eye close-up image. The method for determining the third probability coefficient based on the sub-image information corresponding to the sub-text information includes:

[0128] Based on computer vision, first eye information and second eye information in the eye close-up image are obtained, the eye aspect ratio is calculated based on the first eye information, the absolute value of the difference between the eye aspect ratio and a preset standard aspect ratio is calculated, and the pupil area ratio is calculated based on the second eye information. The absolute value of the difference and the pupil area ratio are weightedly summed to obtain a third probability coefficient.

[0129] The method for calculating the eye aspect ratio based on the first eye information includes:

[0130] ;

[0131] Where EAR is the aspect ratio of the eye, is the inner eye corner coordinate, is the coordinate of the inner side of the upper eyelid, is the coordinate of the inner side of the lower eyelid, is the outer eye corner coordinate, is the outer coordinate of the lower eyelid, is the outer coordinate of the upper eyelid, is the Euclidean distance function.

[0132] The method for constructing a weight allocation sequence according to the first probability coefficient, the second probability coefficient and the third probability coefficient includes:

[0133] Traverse M groups of sub-text information and the corresponding sub-detection information and sub-image information, determine the first number of times that the first probability coefficient is greater than a preset first coefficient threshold, determine the second number of times that the second probability coefficient is greater than a preset second coefficient threshold, determine the third number of times that the third probability coefficient is greater than a preset third coefficient threshold, calculate the sum of the first number, the second number and the third number as the total number of times, take the ratio of the first number to the total number of times as the first weight of the conversation text information, take the ratio of the second number to the total number of times as the second weight of the physiological detection information, take the ratio of the third number to the total number of times as the third weight of the facial image information, and take the first weight, the second weight and the third weight as a weight allocation sequence.

[0134] Detection module: Determines the detection result based on the conversation text information, physiological detection information, facial image information and weight allocation sequence.

[0135] Methods for determining test results include:

[0136] The conversation text information, physiological detection information, facial image information, and weight assignment sequence are input into the pre-built machine learning model to obtain the detection results.

[0137] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters, weights and thresholds in the formula are set by technicians in this field according to actual conditions.

[0138] The above embodiments may be implemented in whole or in part by software, hardware, firmware or any other combination thereof. When implemented by software, the above embodiments may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired network or a wireless network. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media sets. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD) or a semiconductor medium. The semiconductor medium may be a solid-state hard disk.

[0139] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in the present invention can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0140] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0141] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the units is only one, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0142] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0143] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0144] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

[0145] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. An intelligent detection method for multimodal natural language and abnormal signs analysis, characterized in that: include: Acquire the conversation text information, physiological detection information and facial image information of the subject during the detection process, wherein the conversation text information includes M groups of sub-text information, each group of sub-text information has corresponding sub-detection information and sub-image information, the sub-detection information belongs to the physiological detection information, and the sub-image information belongs to the facial image information; Traversing M groups of sub-text information, mapping the sub-text information to the data space, thereby constructing a probability binary tree, and determining a first probability coefficient based on the probability binary tree, including: obtaining a real-time height value of each sub-binary tree in the probability binary tree according to a depth-first search algorithm, performing weighted summation based on the real-time height value of each sub-binary tree to obtain a comprehensive height value, and using the comprehensive height value as the first probability coefficient; Determine a second probability coefficient based on the sub-detection information corresponding to the sub-text information, determine a third probability coefficient based on the sub-image information corresponding to the sub-text information, and construct a weight allocation sequence according to the first probability coefficient, the second probability coefficient, and the third probability coefficient; Determine the detection result based on the conversation text information, physiological detection information, facial image information and weight distribution sequence, and the detection result is used to characterize whether the subject is lying; Among them, the second probability coefficient is: SPC= ;in, , , , , , They are real-time blood pressure value, standard blood pressure value, real-time respiratory rate value, standard respiratory rate value, real-time pulse rate value, standard pulse rate value, is a natural constant; The method for determining a third probability coefficient based on sub-image information corresponding to the sub-text information includes: obtaining first eye information and second eye information in an eye close-up image based on computer vision, calculating an eye aspect ratio based on the first eye information, calculating an absolute value of a difference between the eye aspect ratio and a preset standard aspect ratio, calculating a pupil area ratio based on the second eye information, and performing weighted summation of the absolute value of the difference and the pupil area ratio to obtain a third probability coefficient; The method for constructing a weight allocation sequence according to the first probability coefficient, the second probability coefficient and the third probability coefficient includes: By counting the number of times the first, second, and third probability coefficients exceed their respective thresholds and calculating the total number of times, the proportion of the number of each probability coefficient is used as the weight of the corresponding probability coefficient to generate a weight allocation sequence.

2. The intelligent detection method of multimodal natural language and abnormal physical sign analysis according to claim 1, characterized in that: Methods for mapping subtext information into data space to construct a probabilistic binary tree include: The sub-text information is used as the root node of the probability binary tree, and whether the preset lie detection keywords appear in each layer of the sub-text information is used as the cutting feature. Based on the cutting feature, the root node is divided into two left and right child nodes, and each child node is cut recursively until each child node cannot be cut further, thereby constructing a probability binary tree; Methods for determining test results include: The conversation text information, physiological detection information, facial image information, and weight assignment sequence are input into the pre-built machine learning model to obtain the detection results.

3. The intelligent detection method of multimodal natural language and abnormal physical sign analysis according to claim 1 is characterized in that: The method of obtaining the sub-detection information and the sub-image information corresponding to each set of sub-text information includes: The start time of the Nth group of sub-text information is taken as the first target time, and the start time of the N+1th group of sub-text information is taken as the second target time. The target time period is constructed with the first target time and the second target time. The physiological detection information is segmented according to the target time period to obtain S groups of sub-detection information. The facial image information is segmented according to the target time period to obtain S groups of sub-image information. The sub-text information, sub-detection information and sub-image information within the target time period correspond to each other, and 0<N<S.

4. The intelligent detection method of multimodal natural language and abnormal physical sign analysis according to claim 1 is characterized in that: The first eye information includes the inner coordinates of the upper eyelid, the outer coordinates of the upper eyelid, the inner coordinates of the lower eyelid, the outer coordinates of the lower eyelid, the inner corner coordinates and the outer corner coordinates. The method for calculating the eye aspect ratio based on the first eye information includes: calculating the eye aspect ratio according to the inner coordinates of the upper eyelid, the outer coordinates of the upper eyelid, the inner coordinates of the lower eyelid, the outer coordinates of the lower eyelid, the inner corner coordinates and the outer corner coordinates.

5. The intelligent detection method of multimodal natural language and abnormal signs analysis according to claim 1 is characterized in that: By counting the number of times the first, second, and third probability coefficients exceed their respective thresholds and calculating the total number of times, the proportion of the number of each probability coefficient is used as the weight of the corresponding probability coefficient to generate a weight distribution sequence, specifically including: Traverse M groups of sub-text information and the corresponding sub-detection information and sub-image information, determine the first number of times that the first probability coefficient is greater than a preset first coefficient threshold, determine the second number of times that the second probability coefficient is greater than a preset second coefficient threshold, determine the third number of times that the third probability coefficient is greater than a preset third coefficient threshold, calculate the sum of the first number, the second number and the third number as the total number of times, take the ratio of the first number to the total number of times as the first weight of the conversation text information, take the ratio of the second number to the total number of times as the second weight of the physiological detection information, take the ratio of the third number to the total number of times as the third weight of the facial image information, and take the first weight, the second weight and the third weight as a weight allocation sequence.

6. The intelligent detection method of multimodal natural language and abnormal signs analysis according to claim 1, characterized in that: Methods for building machine learning models include: Obtaining a sample data set, the sample data set including historical conversation text information, historical physiological detection information, historical facial image information, historical weight distribution sequence, and historical detection results; Divide the sample data set into a sample training set and a sample test set, and build a regression network; The historical conversation text information, historical physiological detection information, historical facial image information and historical weight distribution sequence in the sample training set are used as input data of the regression network, and the historical detection results in the sample training set are used as output data of the regression network. The regression network is trained to obtain an initial regression network for predicting the initial detection results. The initial regression network is tested using a sample test set, and the initial regression network with an output smaller than a preset error value is used as a machine learning model.

7. An intelligent detection system for multimodal natural language and abnormal physical sign analysis, which is used to implement the intelligent detection method for multimodal natural language and abnormal physical sign analysis according to any one of claims 1 to 6, characterized in that: include: Data acquisition module: used to acquire the conversation text information, physiological detection information and facial image information of the subject during the detection process. The conversation text information includes M groups of sub-text information. Each group of sub-text information has corresponding sub-detection information and sub-image information. The sub-detection information belongs to the physiological detection information, and the sub-image information belongs to the facial image information. The first processing module is used to traverse the M groups of sub-text information, map the sub-text information to the data space, thereby constructing a probability binary tree, and determine the first probability coefficient based on the probability binary tree; The second processing module determines a second probability coefficient based on the sub-detection information corresponding to the sub-text information, determines a third probability coefficient based on the sub-image information corresponding to the sub-text information, and constructs a weight allocation sequence according to the first probability coefficient, the second probability coefficient and the third probability coefficient; Detection module: Determines the detection result based on the conversation text information, physiological detection information, facial image information and weight allocation sequence.

Citation Information

Patent Citations

  • Automatic lie detection method and system based on domain adversarial training

    CN112329438A

  • Lie detection method and device, computer equipment and storage medium

    CN109793526A

  • Deception detection using oculomotor, cardiovascular, respiratory and electrodermal measures

    US20220087584A1