Multimodal sentiment computing intelligent interrogation system and method based on large model

CN122432305BActive Publication Date: 2026-08-21HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610894487.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-22
Publication Date
2026-08-21
Estimated Expiration
2046-06-22

AI Technical Summary

Technical Problem

(1)缺乏面向旅客问询场景的上下文建模机制,难以将身份、行程、历史行为、业务规则和外部态势等多源先验信息转换为可供大模型推理的上下文表示和上下文关联结构;

Benefits of technology

(1)本发明通过上下文建模模块将与待问询对象相关的多源先验信息转换为机器可理解的上下文表示,并进一步构建上下文关联结构,使大模型智能体能够基于身份状态、行程状态、历史行为状态、风险关联状态、时效性约束状态及其相互关系生成问询策略,避免仅依据孤立字段或固定模板进行问题生成。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122432305B_ABST
    Figure CN122432305B_ABST
Patent Text Reader

Abstract

The application discloses a multi-modal sentiment computing intelligent inquiry system and method based on a large model, and relates to the technical field of intelligent inquiry. The system comprises the following modules: passenger context modeling, initial inquiry strategy generation, multi-modal state shift modeling, inquiry state recursion and strategy updating, inquiry termination discrimination and result generation, etc. The system converts multi-source prior information into passenger context representation and associated structure, generates an initial inquiry strategy in combination with external situation knowledge; time-aligns and jointly encodes multi-modal time series data such as vision, voice and millimeter wave radar, calculates state shift, inter-modal consistency and context matching relationship; dynamically updates the inquiry strategy through inquiry state recursion, triggers a fuse based on the inquiry execution degree and execution boundary, and generates a phased result; and converts the phased result into an evidence association structure, fuses artificial constraint signals to form a traceable collaborative discrimination result. The application is applicable to scenarios such as border defense, customs, airport security, etc., and realizes intelligent inquiry assistance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multimodal affective computing and intelligent inquiry technology, specifically to a multimodal affective computing intelligent inquiry system and method based on a large model, which can be applied to scenarios such as border inspection, customs passenger management, airport security, and port passenger inquiry that require identity verification, itinerary verification, status change analysis, risk clue auxiliary prompts, and manual review.

[0002] The multimodal emotion computing in this invention is not limited to outputting discrete emotion categories. Instead, it generates multimodal state vectors using multi-source time-series data such as visual modalities, speech modalities, text modalities, and millimeter-wave radar modalities. It calculates state offsets, intermodal consistency, and context matching relationships to characterize the state changes of the object to be queried relative to the baseline state during the inquiry process, and provides auxiliary basis for subsequent inquiry strategy updates and manual review. Background Technology

[0003] In scenarios such as border inspection, customs supervision, airport security checks, and port passenger management, staff typically need to verify passenger identity, travel purpose, travel history, consistency of statements, and other risk clues within a limited time. Traditional inquiry methods rely heavily on staff experience, which suffers from problems such as scattered inquiry data, low information retrieval efficiency, inconsistent inquiry strategies among different staff, insufficient utilization of historical records, difficulty in dynamically adjusting the inquiry process, and insufficient basis for verifying inquiry results.

[0004] Existing multimodal emotion recognition or intelligent interrogation technologies typically focus on single scenarios or fixed tasks. For example, some technical solutions utilize facial expressions, voice, text, body language, or physiological signals for emotion analysis, but these are mainly used for dialogue emotion recognition and fail to combine the contextual structure of the person being questioned, business rules, and multi-round questioning states to generate questioning strategies. Some interrogation systems assess credibility based on features such as facial expressions, heart rate, and sound intensity collected from audio and video data, but their application is mainly in criminal interrogation scenarios and is difficult to directly adapt to passenger clearance questioning processes. Other technical solutions use millimeter-wave radar to acquire physiological signals such as heart rate and respiration for emotion recognition, but most remain at the level of single-time recognition and do not combine radar signals with questioning questions, semantic responses, contextual states, and historical interaction processes to form a dynamic questioning loop.

[0005] Therefore, the existing technology has at least the following shortcomings: (1) There is a lack of context modeling mechanism for passenger inquiry scenarios, making it difficult to convert multi-source prior information such as identity, itinerary, historical behavior, business rules and external situation into context representation and context association structure that can be used for large model reasoning; (2) Existing multimodal emotion recognition methods mostly remain at the level of single emotion category recognition or single modality analysis, lacking joint modeling of multimodal state shift, intermodal consistency and context matching relationship of the question segment relative to the baseline state; (3) Existing intelligent inquiry systems are usually unable to perform state recursion on the question representation, answer representation, multimodal state vector and the previous inquiry state in each round, and it is also difficult to dynamically update the subsequent question generation constraints based on the recursed inquiry state. (4) Existing multi-round inquiry processes usually rely on fixed rounds or manual experience to terminate, and lack an automatic circuit breaker mechanism based on inquiry execution degree, information gain and execution boundary conditions; (5) Existing inquiry assistance systems often only output results or prompts, lacking a mechanism to link phased inquiry results, source information, status shifts and manual review results into an evidence structure, making it difficult to meet the needs of subsequent review, auditing and traceability. Summary of the Invention

[0006] 1. The technical problem that the invention aims to solve In view of the problems existing in the prior art, the present invention provides a multimodal emotion computing intelligent inquiry system and method based on a large model. The present invention does not directly utilize a large model to generate inquiry questions based on original fields, nor does it only identify a single multimodal emotional state. Instead, it models the multi-source prior information of the object to be inquired as a context representation and context association structure, models the multimodal temporal data in the inquiry process as a state offset relative to a baseline state, and further recursively calculates the updated inquiry state representation from the question representation, answer representation, multimodal state vector, and the previous inquiry state for each round of inquiry, continuously constraining the generation of subsequent questions. Simultaneously, the system triggers circuit breakers based on inquiry execution degree, information gain, and execution boundary conditions, and forms a traceable and verifiable collaborative discrimination result through evidence association structures and manually constrained signals.

[0007] 2. Technical Solution To achieve the above objectives, the technical solution provided by the present invention is as follows: This invention discloses a multimodal affective computing intelligent inquiry system based on a large model, comprising a context modeling module, an initial inquiry strategy generation module, a multimodal state shift modeling module, an inquiry state recursion and strategy update module, an inquiry termination judgment and result generation module, and a human-computer collaborative judgment constraint and evidence representation module; wherein: The context modeling module is used to obtain multi-source prior information related to the object to be queried, convert the multi-source prior information into a context representation, and construct a context association structure based on the context representation. The initial query strategy generation module is used to form initial question generation constraints based on context representation, context association structure and external situation knowledge, and call the large model agent to generate the initial query strategy. The multimodal state offset modeling module is used to acquire time-series data of at least two modalities during the query process, perform time alignment, quality assessment, normalization and joint encoding on the time-series data of the at least two modalities to obtain a multimodal state vector corresponding to the current query segment, and determine the state offset of the current query segment relative to the baseline state, intermodal consistency and matching relationship with the context representation based on the multimodal state vector. The query state recursion and strategy update module is used to encode and fuse the question representation, answer representation, multimodal state vector and the previous query state representation of each round of query to obtain the updated query state representation, and update the subsequent question generation constraints based on the updated query state representation; if the termination condition is not met, the large model agent is called to generate the next round of query strategy based on the subsequent question generation constraints. The query termination judgment and result generation module is used to calculate the current query execution degree based on the updated query status representation, and determine whether the termination condition is met based on the query execution degree, preset completion conditions and preset execution boundary conditions. When the termination condition is met, a phased query result is generated based on the current query status representation. The human-machine collaborative discrimination constraint and evidence representation module is used to convert the phased inquiry results into evidence association structures and convert the manual discrimination results into manual constraint signals, so that the manual constraint signals and evidence association structures together form collaborative discrimination results.

[0008] Furthermore, the system also includes an inquiry interaction module, which is used to receive multi-source prior information input, carry out inquiry process interaction, and receive manual judgment result input.

[0009] Furthermore, the multi-source prior information includes at least one of the following: identity information, itinerary information, historical behavior information, business rule information, historical interaction information, scenario constraint information, manual input information, and external situational knowledge related to the current inquiry scenario.

[0010] Furthermore, the modalities processed by the multimodal state offset modeling module include at least two of the following: visual modality, speech modality, text modality, millimeter-wave radar modality, infrared modality, depth image modality, thermal imaging modality, and contact or non-contact physiological signal modality.

[0011] Furthermore, before performing joint encoding, the multimodal state offset modeling module performs a quality assessment on the temporal data of each modality and determines the weight of each modality in the joint encoding based on the quality assessment results.

[0012] Furthermore, the baseline state is determined by at least one of the following: initial perception data at the start of the inquiry phase, historical statistical data, a stable time window during the same inquiry process, or a statistical baseline of similar objects.

[0013] Furthermore, the evidence association structure is a graph structure. ,in, Represents the set of evidence nodes. The evidence edge set is represented by evidence nodes, which include at least one of the following: question node, answer semantic node, context state node, multimodal state node, state offset node, external situation node, and artificial constraint node. Evidence edges are used to represent at least one of the following: source relationship, trigger relationship, support relationship, conflict relationship, correction relationship, or continued verification relationship.

[0014] Furthermore, the termination conditions of the query termination judgment and result generation module include at least one of the following: the current query execution degree meets the preset completion condition, the preset maximum number of query rounds is reached, the preset maximum query duration is reached, the maximum number of questions is reached, the additional information gain that can be obtained by continuing to query is lower than a preset threshold, or the manual takeover condition is met.

[0015] Furthermore, the phased query results include at least one of the following: confirmed information, unconfirmed information, unresolved inconsistencies, relevant multimodal state shifts, corresponding contextual basis, external situational basis, and subsequent processing suggestions.

[0016] The present invention provides a multimodal emotion computing intelligent query method based on a large model, comprising the following steps: S1, obtain multi-source prior information related to the object to be queried, and convert the multi-source prior information into a contextual representation; S2, Construct a context association structure based on the context representation; S3, acquire external situational knowledge related to the current query scenario, and map the context representation, context association structure and external situational knowledge to a unified strategy generation space to form initial question generation constraints; S4, based on the initial question generation constraints, call the large model to generate the initial query strategy; S5, during the inquiry process, simultaneously acquire time-series data of at least two modalities, perform time alignment, quality assessment, normalization processing and joint encoding on the time-series data of different modalities, and obtain a multimodal state vector corresponding to the current inquiry segment; S6. Based on the multimodal state vector, determine the state offset of the current query segment relative to the baseline state, the intermodal consistency, and the matching relationship with the context representation, and use the state offset, intermodal consistency, and context matching relationship as inputs for subsequent question generation constraint updates; S7, fuse the question representation, answer representation, multimodal state vector of the current round and the query state representation of the previous round to obtain the updated query state representation, and update the subsequent question generation constraints based on the updated query state representation; S8, calculate the current query execution degree based on the updated query status representation, and determine whether to continue generating the next round of questions based on the query execution degree, preset completion conditions and preset execution boundary conditions; S9, after the termination condition is triggered, generate a phased query result based on the current query status; S10, the phased inquiry results are converted into an evidence association structure, and the manual judgment results are converted into manual constraint signals, so that the manual constraint signals and the evidence association structure together form a collaborative judgment result.

[0017] 3. Beneficial effects Compared with the prior art, the technical solution provided by this invention has the following advantages: (1) The present invention converts multi-source prior information related to the object to be queried into a machine-understandable context representation through a context modeling module, and further constructs a context association structure, so that the large model agent can generate a query strategy based on identity status, travel status, historical behavior status, risk association status, timeliness constraint status and their interrelationships, avoiding the generation of questions based solely on isolated fields or fixed templates.

[0018] (2) This invention introduces external situation knowledge into the initial inquiry strategy generation process, so that the inquiry strategy not only depends on the static context of the object to be inquired, but also can be dynamically constrained by time-sensitive factors such as real-time policy changes, destination situation, emergency information, port operation status, and public health tips, thereby improving the adaptability of the inquiry strategy to the current scenario.

[0019] (3) The present invention converts multiple modalities such as visual modality, speech modality, text modality and millimeter-wave radar modality into a unified multimodal state vector through a multimodal state offset modeling module, and calculates state offset, intermodal consistency and context matching relationship based on the baseline state, so that the system can characterize the state changes in the inquiry process from the perspective of multimodal temporal change, rather than just performing single or single-modal emotion recognition.

[0020] (4) The present invention integrates the question representation, answer representation, multimodal state vector and the previous round of question representation in each round through the question state recursion and strategy update module to form a continuously updated question state representation, so that the subsequent question generation constraints can dynamically change with the confirmed information, unresolved inconsistencies, context gaps and external situation constraints, thereby improving the adaptive capability of multi-round questioning.

[0021] (5) The present invention calculates the current query execution degree through the query termination judgment and result generation module, and determines whether to continue generating follow-up query constraints by combining the maximum query rounds, the longest query duration, the maximum number of questions and the new information gain; after triggering the execution degree circuit breaker or the upper limit circuit breaker, the system can output the phased query results based on the existing query state, avoid unconstrained repeated queries, and ensure that a result that can be reviewed can still be formed when the execution boundary is reached.

[0022] (6) This invention converts the phased inquiry results into an evidence association structure and the manual judgment results into manual constraint signals through a human-machine collaborative judgment constraint and evidence representation module, so that the system output results and the manual review process form a traceable collaborative judgment result. Thus, the state offset, context matching relationship and phased results generated by the system are used to assist the staff in review, rather than directly replacing the final manual judgment. Attached Figure Description

[0023] Figure 1 This is a schematic diagram of the architecture of the multimodal emotion computing intelligent inquiry system based on a large model according to the present invention. Figure 1 It includes a context modeling module, an initial query strategy generation module, a multimodal state offset modeling module, a query state recursion and strategy update module, a query termination judgment and result generation module, a human-machine collaborative judgment constraint and evidence representation module, and a query interaction module. The context modeling module generates context representations and context association structures based on multi-source prior information; the initial query strategy generation module generates an initial query strategy by combining context representations, context association structures, and external situational knowledge; the multimodal state offset modeling module generates multimodal state vectors and state offsets based on multimodal time-series data such as vision, speech, text, and millimeter-wave radar; the query state recursion and strategy update module updates the query state representation and generates constraints for subsequent questions based on question representations, answer representations, and multimodal state vectors; the query termination judgment and result generation module generates phased query results based on query execution degree and termination conditions; the human-machine collaborative judgment constraint and evidence representation module converts phased query results into evidence association structures and combines them with human constraint signals to form collaborative judgment results; and the query interaction module receives multi-source prior information input, carries out query process interaction, and receives human judgment result input.

[0024] Figure 2This is a flowchart illustrating the intelligent question-asking method for multimodal emotion computing based on a large model according to the present invention. Figure 2 The document illustrates the complete process from acquiring multi-source prior information, constructing context representations and context association structures, introducing external situational knowledge and forming initial question generation constraints, to generating an initial query strategy, acquiring multimodal time-series data, calculating state offsets and context matching relationships, updating query state representations, calculating query execution degree, and determining whether the termination condition is met. When the termination condition is not met, the system continues querying and updates the strategy; when the termination condition is met, the system generates interim query results and further generates evidence association structures, combining them with manually constrained signals to form collaborative discrimination results.

[0025] Figure 3 This is a schematic diagram of the multimodal state migration modeling process in this invention. Figure 3 The system comprises input data, a feature processing submodule, an offset and relation modeling submodule, and output results. Input data includes visual modal time-series data, speech modal time-series data, text modal time-series data, and millimeter-wave radar time-series data. The feature processing submodule performs time alignment, normalization, quality assessment, and multimodal joint encoding on the multimodal time-series data to generate multimodal state vectors. The offset and relation modeling submodule determines the baseline state and current state representations, calculates the state offset, and further generates intermodal consistency analysis results and context matching analysis results. Output results include multimodal state vectors, state offsets, intermodal consistency, and context matching relationships.

[0026] Figure 4 This is a schematic diagram of the query state recursion, query execution degree calculation, and circuit breaker mechanism in this invention. Figure 4 The system includes input information, a state recursion and policy update submodule, an execution degree calculation and termination judgment submodule, and output results. Input information includes the previous round of query state representation, the current question representation, the answer representation, and a multimodal state vector. The state recursion and policy update submodule performs multi-source state fusion and query state recursion, generating updated query state representations and subsequent question generation constraints. The execution degree calculation and termination judgment submodule analyzes confirmed information and inconsistencies, calculates query execution degree, and determines whether to continue querying or trigger a circuit breaker based on the termination condition judgment result. When the termination condition is not met, the system returns to the subsequent question generation constraint update process to continue querying; when the termination condition is met, the system triggers a circuit breaker and generates interim query results. Detailed Implementation

[0027] To facilitate understanding of the technical solution of this invention, the relevant concepts are explained below, but this does not constitute a limitation on the scope of protection of this invention.

[0028] The multi-source prior information refers to prior information that can be used to describe the object to be queried and its current scenario before or during the inquiry process. This includes at least one of the following: identity information, travel information, historical behavior information, business rule information, historical interaction information, scenario constraint information, manual input information, and external situational knowledge related to the current inquiry scenario.

[0029] The context representation refers to a machine-understandable representation formed after semantic normalization, entity relationship modeling, temporal relationship modeling, and constraint relationship modeling of multi-source prior information. This representation is used to describe at least one state of the object to be queried, including identity state, travel state, historical behavior state, risk association state, and timeliness constraint state.

[0030] The aforementioned context association structure refers to a structured representation consisting of multiple context state nodes and their associated edges. State nodes represent identity state, travel state, historical behavior state, risk-related state, or time-sensitive constraint state; associated edges represent temporal relationships, entity relationships, causal constraint relationships, consistency relationships, conflict relationships, or missing relationships between different states.

[0031] The external situational knowledge refers to timely knowledge relevant to the current inquiry scenario, including at least one of the following: real-time policy changes, destination situation, emergency information, traffic operation status, port operation status, public health alerts, regional risk alerts, or temporary business rules. External situational knowledge can be obtained from rule bases, business knowledge bases, manually maintained data sources, or authorized external information interfaces, and can be structured according to timestamps, applicable regions, applicable objects, and validity periods.

[0032] The multimodal state vector refers to the state representation obtained by performing time alignment, quality assessment, normalization processing, and joint encoding on at least two modalities among visual modality, speech modality, text modality, millimeter-wave radar modality, and other contact or non-contact physiological signal modalities.

[0033] The state offset refers to the degree of deviation of the multimodal state vector corresponding to the current query segment from the baseline state. The baseline state can be determined by at least one of the following: initial perception data at the beginning of the query, historical statistical data, a stable time window in the same query process, or a statistical baseline of similar objects.

[0034] The query state representation refers to an intermediate state representation formed by context representation, executed query actions, question representation, answer semantic representation, multimodal state vector, state offset, and external situational constraints. It is used to describe confirmed information, unconfirmed information, unresolved inconsistencies, context gaps to be verified, and timeliness constraints in the current query process.

[0035] The unresolved inconsistencies refer to relationships discovered during the inquiry process through matching analysis between contextual association structures, answer semantic representations, multimodal state vectors, state offsets, external situational knowledge, or business rules, and which have not yet been explained, confirmed, corrected, or excluded through subsequent inquiries, supplementary information, manual confirmation, or rule verification. These inconsistencies may include at least one of the following: inconsistencies between identity information and answer content; inconsistencies between travel information and historical behavior; inconsistencies between the semantics of previous and subsequent answers; inconsistencies between answer content and business rules or external situational knowledge; and unexplained relationships between multimodal state offsets and the current answer semantics or contextual state. The term "unresolved" does not indicate a final anomaly judgment, but rather that the relationship is still pending further verification, explanation, or manual review.

[0036] The query execution degree refers to an indicator used to characterize the degree of completion of the current query objective, which may include at least one of the following: identity verification completion degree, itinerary verification completion degree, context consistency resolution degree, abnormal state explanation degree, necessary information coverage degree, and new information gain.

[0037] The interim query result refers to the intermediate result generated by the system based on the current query status after triggering the execution degree circuit breaker or the upper limit circuit breaker. It includes at least one of the following: confirmed information, unconfirmed information, unresolved inconsistencies, related multimodal state offsets, corresponding contextual basis, external situational basis, and subsequent processing suggestions.

[0038] The manual constraint signal refers to the constraint information formed by staff after reviewing the interim inquiry results, used to mark the parts of the interim inquiry results that have been confirmed, denied, corrected, or require further verification.

[0039] The large-scale model agent refers to a policy generation and execution entity built based on a large language model, a multimodal large model, or a large model adapted for the query task. This large-scale model agent is not a parallel system module independent of the context modeling module, the initial query strategy generation module, the multimodal state shift modeling module, the query state recursion and policy update module, the query termination judgment and result generation module, and the human-machine collaborative judgment constraint and evidence representation module. Instead, it is invoked by the initial query strategy generation module and the query state recursion and policy update module to generate the initial query strategy or the next round of query strategy based on at least one of the following: context representation, context association structure, external situational knowledge, initial question generation constraints, updated query state representation, subsequent question generation constraints, state shift, intermodal consistency, and context matching relationship. The large-scale model agent does not directly perform multimodal state shift calculation, query execution degree judgment, stage result generation, or manual final judgment.

[0040] Combination Figure 1 This invention provides a multimodal emotion computing intelligent inquiry system based on a large model, comprising: a context modeling module, an initial inquiry strategy generation module, a multimodal state shift modeling module, an inquiry state recursion and strategy update module, an inquiry termination judgment and result generation module, and a human-computer collaborative judgment constraint and evidence representation module. In an optional embodiment, the system further comprises an inquiry interaction module.

[0041] The inquiry interaction module is used to receive multi-source prior information input, carry out inquiry process interaction, and receive manual judgment result input. The inquiry interaction module can be implemented through a graphical interface, business system interface, voice interaction interface, text interaction interface, or manual input, but is not limited to a specific interaction interface form. The multi-source prior information received by the inquiry interaction module can be input into the context modeling module, and the questions, answers, or manually marked information generated during the inquiry process interaction can be input into the inquiry state recursion and strategy update module or the human-machine collaborative judgment constraint and evidence representation module.

[0042] The context modeling module is used to acquire multi-source prior information related to the object to be queried, and convert the multi-source prior information into a machine-understandable context representation, and construct a context association structure based on the context representation. The context representation is not limited to a single data source or a fixed field format, but is used to represent at least one of the following states of the object to be queried: identity status, travel status, historical behavior status, risk association status, and timeliness constraint status, and their association relationships.

[0043] The context modeling module processes multi-source prior information through at least one of the following methods: semantic normalization, entity relationship modeling, temporal relationship modeling, and constraint relationship modeling. It then extracts state nodes and associated edges to form a contextual association structure. This structure describes the consistency, absence, or abnormal associations of the object to be queried across dimensions such as identity, travel itinerary, historical behavior, risk association, and timeliness constraints. This enables large models to generate query strategies based on the contextual association structure, rather than simply generating questions from isolated fields or fixed templates.

[0044] During the inquiry process, the context modeling module can also update the context representation based on the passenger's answer, the staff's marking results, and the real-time multimodal state analysis results, so that the passenger context is expanded from static prior information to dynamic context that includes the inquiry state.

[0045] The initial query strategy generation module is used to generate an initial query strategy at the beginning of the query stage based on the context representation, context association structure and external situation knowledge output by the context modeling module.

[0046] The initial inquiry strategy generation module maps context representation, context association structure, and external situational knowledge to a unified strategy generation space. It then performs consistency analysis on identity status, travel status, historical behavior status, risk association status, and timeliness constraint status to determine the focus dimensions and question generation constraints for the initial inquiry. These constraints limit the question types, verification objectives, question granularity, semantic direction, follow-up question direction, and information collection priorities generated by the large model agent, ensuring that the first round of inquiry is driven by both the contextual state of the object to be queried and the current situational context.

[0047] When generating the initial inquiry strategy, the initial inquiry strategy generation module outputs at least the inquiry target, a set of candidate questions, the criteria for ranking the questions, and the contextual triggering conditions corresponding to each candidate question. The large model agent generates the first round of verification or clarifying questions based on the above output, and stores these questions in association with their corresponding contextual criteria for subsequent rounds of inquiry, enabling consistency analysis of responses and strategy updates.

[0048] The multimodal state offset modeling module is used to acquire time-series data of at least two modalities during the query process, perform time alignment, quality assessment, normalization and joint encoding on the time-series data of the at least two modalities to obtain a multimodal state vector corresponding to the current query segment, and determine the state offset of the current query segment relative to the baseline state based on the multimodal state vector.

[0049] The modalities may include at least two of the following: visual modalities, speech modalities, text modalities, millimeter-wave radar modalities, infrared modalities, depth imaging modalities, thermal imaging modalities, and contact or non-contact physiological signal modalities. Specifically, visual modalities may include visual behavioral features such as facial region changes, head posture, eye movement, or body movements; speech modalities may include acoustic features such as speech rate, intensity, fundamental frequency, pauses, vibrato, and intonation variations; text modalities may include semantic representations of responses, semantic relationships between keywords, and matching relationships with contextual representations; and millimeter-wave radar modalities may include non-contact physiological or kinematic features such as respiratory rate, heart rate changes, chest micro-movements, head micro-movements, and body movement amplitude.

[0050] The multimodal state offset modeling module aligns and normalizes the temporal features of different modalities, establishes the correspondence between question fragments, answer fragments and multimodal state features; then, based on the signal quality, temporal consistency, semantic relevance or confidence of each modality, it performs weighted fusion, cross-attention fusion, graph structure fusion or sequence model fusion on the features of different modalities to obtain the multimodal state vector corresponding to the current query fragment.

[0051] In one implementation, the state offset can be expressed as:

[0052] in, This represents the multimodal state vector corresponding to the current query segment. Represents the baseline state vector. This represents the distance function or offset metric function. The distance function can be a weighted Euclidean distance, cosine distance, Mahalanobis distance, dynamic time warped distance, or a distance function learned by the model.

[0053] Before joint coding, the multimodal state offset modeling module can perform quality assessment on the time-series data of each modality and determine the weight of each modality in joint coding based on the quality assessment results. When the signal quality of a certain modality is lower than a preset quality threshold, the system reduces the weight of that modality in joint coding or marks the modality as a low-confidence input. When a certain modality is missing, occluded, has excessive noise, or is severely out of sync with other modalities within the current time window, the system can adjust the joint coding results based on its quality label to avoid the unstable impact of a single low-quality modality on the multimodal state vector and state offset calculation.

[0054] The multimodal state vector represents the state offset of the current query fragment relative to the baseline state, intermodal consistency, and the matching relationship with the context representation. Intermodal consistency can be calculated based on the degree of consistency between different modal state offset directions, magnitudes, time positions, or confidence levels; the context matching relationship can be determined based on the degree of matching between the semantic representation of the answer and the corresponding state nodes in the context association structure. The large model agent updates the constraints for subsequent question generation based on the state offset, intermodal consistency, and context matching relationship, enabling the subsequent query strategy to adaptively adjust according to real-time state changes.

[0055] The query state recursion and strategy update module is used to continuously model the query state during multiple rounds of querying and update the subsequent question generation constraints based on the formed query state. The query state includes an intermediate state representation formed by context representation, executed query actions, question representation, answer semantic representation, multimodal state vector, state offset, and external situational knowledge.

[0056] The query state recursion and strategy update module encodes and fuses the question representation, answer representation, multimodal state vector, and previous query state representation for each round of querying to obtain an updated query state representation. This updated query state representation is used to characterize confirmed information, unconfirmed information, unresolved inconsistencies, context gaps requiring further verification, and timeliness constraints related to external situational knowledge during the current query process.

[0057] Based on the updated query state representation, the query state recursion and strategy update module updates the constraints of the subsequent question generation space; the large model agent generates the next round of query strategy under the constraints of the subsequent question generation, so that the generation target, question granularity, semantic direction, verification focus and ranking criteria of the next round of questions change dynamically with the query process.

[0058] When the query termination judgment result indicates that the termination condition is not met, the large model agent generates a next-round query strategy based on the updated query state representation, subsequent question generation constraints, contextual association structure, and external situational knowledge. The next-round query strategy includes at least one of the following: next-round follow-up target, follow-up question, question granularity, verification criteria, expected information gain, and corresponding contextual triggering conditions. Therefore, subsequent queries do not simply continue the previous round of questions, but are driven by both the current query state and the constraint update results.

[0059] The query termination judgment and result generation module is used to determine whether the current query should continue to be executed based on the state recursion results during the multi-round query process, and to generate a stage query result based on the information obtained when the termination condition is met. The termination condition includes at least one of the following: the current query execution degree meets a preset completion condition, a preset maximum number of query rounds is reached, a preset maximum query duration is reached, a maximum number of questions is reached, the additional information gain that can be obtained by continuing the query is less than a preset threshold, or the manual takeover condition is met.

[0060] The query termination judgment and result generation module calculates the current query execution degree based on confirmed information, unresolved inconsistencies, context gaps, multimodal state shifts, and external situational constraints. The query execution degree characterizes the degree to which the current query objective is achieved, including at least one of the following: identity verification completion degree, itinerary verification completion degree, context consistency resolution degree, abnormal state explanation degree, necessary information coverage degree, and added information gain.

[0061] In one implementation, the query execution degree of the current round t can be expressed as:

[0062] in, This indicates the extent to which information coverage has been confirmed. Indicates the degree to which inconsistencies are resolved. Indicates the degree of contextual gap completion. Indicates the degree of interpretation of multimodal state shifts. This indicates the addition of information gain. to This indicates the corresponding weight. The above indicators and weights can be obtained through rule pre-setting, historical sample statistics, manual configuration, or model learning.

[0063] When the query execution rate reaches the preset completion condition, the system triggers an execution rate circuit breaker, stopping the generation of new follow-up constraints. When the query execution rate does not reach the preset completion condition but reaches the maximum number of query rounds, the longest query duration, the maximum number of questions, or the new information gain is below a preset threshold, the system triggers an upper limit circuit breaker. After triggering the execution rate circuit breaker or the upper limit circuit breaker, the query termination judgment and result generation module outputs a phased query result based on the currently formed query state. The phased query result includes at least confirmed information, unconfirmed information, unresolved inconsistencies, relevant multimodal state offsets, corresponding contextual basis, external situational basis, and subsequent processing suggestions.

[0064] The human-machine collaborative discrimination constraint and evidence representation module is used to convert the phased inquiry results output by the inquiry termination discrimination and result generation module into evidence representations that can be manually reviewed, and to constrain and confirm the inquiry status or evidence association structure based on the manual discrimination results.

[0065] The human-machine collaborative discrimination constraint and evidence representation module generates an evidence association structure based on confirmed information, unconfirmed information, context gaps, multimodal state shifts, and external situational constraints in the phased inquiry results. This evidence association structure describes the correspondence between each conclusion to be reviewed and its source information. The source information includes at least one of the following: contextual triggering conditions, inquiry question representation, answer semantic representation, multimodal state vector, state shift degree, external situational constraints, inquiry round information, and time segment information.

[0066] In one implementation, the evidence association structure can be represented as a graph structure. ,in, The set of evidence nodes includes at least one of the following: question nodes, answer semantic nodes, context state nodes, multimodal state nodes, state offset nodes, external situation nodes, and manual constraint nodes. The set of evidence edges represents at least one of the following: source relationship, triggering relationship, supporting relationship, conflict relationship, correction relationship, or continued verification relationship.

[0067] The manual review results are converted into manual constraint signals and written into the inquiry status representation or evidence association structure together with the aforementioned evidence association structure. The manual constraint signals are used to mark the parts of the interim inquiry results that are confirmed, denied, corrected, or require further verification. In this way, the system forms a collaborative judgment result composed of machine state recursion results, evidence association structures, and manual constraint signals, enabling the system output to be reviewed, traced, and corrected, rather than directly replacing the final human judgment.

[0068] The interim inquiry results, state shifts, evidence association structures, and collaborative judgment results output by the system are only used as information for staff review and decision support, and are not the sole basis for making final decisions on the subjects of inquiry.

[0069] Combination Figure 2 Based on the above system, the present invention also provides a multimodal emotion computing intelligent query method based on a large model, comprising the following steps: S1, Obtain multi-source prior information related to the object to be queried, and convert the multi-source prior information into a machine-understandable contextual representation; S2, construct a context association structure based on the context representation, the context association structure is used to characterize at least one state of the object to be queried in identity state, travel state, historical behavior state, risk association state and timeliness constraint state and their association relationship; S3, acquire external situational knowledge related to the current query scenario, and map the context representation, context association structure and external situational knowledge to a unified strategy generation space to form initial question generation constraints; S4. Based on the initial question generation constraints, the large model is invoked to generate an initial query strategy. The initial query strategy includes at least the query target, a set of candidate questions, a question ranking basis, and contextual triggering conditions corresponding to each candidate question. S5. During the inquiry process, at least two modalities of time-series data are acquired simultaneously. The time-series data of different modalities are time-aligned, quality-assessed, normalized, and jointly encoded to obtain a multimodal state vector corresponding to the current inquiry segment. S6. Based on the multimodal state vector, determine the state offset of the current query segment relative to the baseline state, the intermodal consistency, and the matching relationship with the context representation, and use the state offset, intermodal consistency, and context matching relationship as inputs for subsequent question generation constraint updates; S7, the question representation, answer representation, multimodal state vector of the current round, and the query state representation of the previous round are fused to obtain an updated query state representation, and the subsequent question generation constraints are updated based on the updated query state representation; the updated query state representation is used to represent confirmed information, unconfirmed information, unresolved inconsistencies, context gaps to be further verified, and timeliness constraints related to external situational knowledge; the subsequent question generation constraints are used to limit at least one of the following for the next round of query generation: question generation target, question granularity, semantic direction, verification focus, ranking criteria, and corresponding context triggering conditions. S8, calculate the current query execution degree based on the updated query state representation, and determine whether to continue generating the next round of questions based on the query execution degree, preset completion conditions, and preset execution boundary conditions; if the query continues, generate the next round of query strategy based on the updated subsequent question generation constraints; if the termination condition is triggered, stop generating new follow-up question constraints. S9. After the termination condition is triggered, a phased query result is generated based on the current query status. The phased query result includes at least one of the following: confirmed information, unconfirmed information, unresolved inconsistencies, related multimodal state shifts, corresponding contextual basis, external situational basis, and subsequent processing suggestions. S10, the phased inquiry results are converted into evidence association structures that can be manually reviewed, and the manual judgment results are converted into manual constraint signals, so that the manual constraint signals and the evidence association structures together form a collaborative judgment result that can be traced and reviewed.

[0070] To further understand the content of this invention, a detailed description of the invention will be provided in conjunction with the accompanying drawings and embodiments.

[0071] Example 1: Construction of Context Representation like Figure 1 and Figure 2 As shown, this embodiment first acquires multi-source prior information related to the object to be queried and converts it into a contextual representation. The multi-source prior information may include identity information, travel information, historical behavior information, business rule information, and external situational knowledge related to the current scenario. The system does not directly use the original fields as input to the large model; instead, it performs semantic normalization, temporal relationship modeling, entity relationship modeling, and constraint relationship modeling on information from different sources to form a machine-understandable contextual representation.

[0072] In one optional implementation, the system receives multi-source prior information input through an inquiry interaction module, and records the questions, answers, and manual marking information between the staff and the person being inquired about during the inquiry process. The inquiry interaction module can transmit the multi-source prior information to the context modeling module, and transmit the manual judgment results to the human-machine collaborative judgment constraint and evidence representation module.

[0073] The context representation can be expressed as a context association structure consisting of multiple state nodes and their associated edges. State nodes represent at least one state of the object to be queried, including identity state, travel state, historical behavior state, risk-related state, and timeliness constraint state. Associated edges represent temporal relationships, entity relationships, causal constraints, consistency constraints, or conflict relationships between different states. Through this context association structure, the system can transform scattered prior information into inference conditions that can be invoked by large model agents.

[0074] Example 2: Initial Inquiry Strategy Generation under External Situational Constraints In this embodiment, the system maps context representation, context association structure, and external situational knowledge to a unified policy generation space. The external situational knowledge may include real-time policy changes, destination situation, emergency information, flight or port operation status, public health alerts, and other time-sensitive knowledge related to the inquiry scenario.

[0075] The system performs consistency analysis on the contextual structure within the strategy generation space to identify state nodes that require priority verification in the current inquiry phase, insufficiently explained contextual gaps, and question generation constraints influenced by external situational knowledge. These question generation constraints define the inquiry objective, candidate question set, question granularity, question ranking criteria, and contextual triggering conditions for each candidate question in the initial inquiry strategy. Therefore, the initial inquiry strategy is jointly determined by the passenger's contextual state and the current scenario situation, rather than being directly generated from a preset fixed template.

[0076] Example 3: Calculation of Multimodal State Migration like Figure 3 As shown, the system simultaneously acquires temporal data from visual, speech, text, and millimeter-wave radar modalities during the inquiry process, and aligns them temporally according to question and answer segments. For the visual modality, the system can extract visual behavioral features such as facial region changes, head posture, eye movement, or body movements; for the speech modality, the system can extract acoustic features such as speech rate, volume, fundamental frequency, pauses, and intonation variations; for the text modality, the system can extract semantic representations of answers, semantic relationships between keywords, and matching relationships with contextual representations; for the millimeter-wave radar modality, the system can extract non-contact physiological or kinematic features such as respiratory rate, heart rate changes, chest micro-movements, head micro-movements, and body movement amplitude.

[0077] The system normalizes, performs quality assessment, and jointly encodes the temporal features of different modalities to obtain a multimodal state vector corresponding to the current query segment. Subsequently, the system determines a baseline state based on initial perceived data at the start of the query, historical statistical data, a stable time window within the same query process, or a statistical baseline for similar objects, and calculates the state offset of the current multimodal state vector relative to the baseline state. This state offset, along with intermodal consistency and context matching, serves as input for subsequent query strategy updates.

[0078] When a certain modality's quality assessment result falls below a preset quality threshold due to occlusion, noise, low recognition confidence, time asynchrony, or missing data, the system can reduce the weight of that modality in joint encoding or mark it as a low-confidence input to reduce the impact of low-quality modalities on the calculation of multimodal state vectors and state offsets.

[0079] Example 4: Query Status Recursion and Execution Degree Circuit Breaker like Figure 4 As shown, the system fuses the question representation, answer representation, multimodal state vector, and previous round's query state representation for each round of inquiry to obtain the current round's query state representation. This query state representation records confirmed information, unresolved inconsistencies, contextual gaps requiring further verification, and timeliness constraints related to external situational knowledge during the current query process.

[0080] The system updates subsequent question generation constraints based on the current query status, ensuring that the generation target, semantic direction, question granularity, and verification focus of the next round of questions change as the query progresses. Simultaneously, the system calculates the current query execution degree based on at least one of the following factors: the proportion of confirmed information, the degree of context gap resolution, the degree of inconsistency relationship resolution, the degree of multimodal state offset interpretation, and the degree of necessary information coverage. The degree of inconsistency relationship resolution can be determined based on the number, weight, or importance of relationships among the discovered inconsistencies that have been confirmed, explained, corrected, or eliminated through follow-up questioning, supplementary information, contextual verification, external situational verification, or manual review. For inconsistencies that have not yet been confirmed, explained, corrected, or eliminated, the system retains them as unresolved inconsistencies and uses them as input for subsequent question generation constraints, interim query results, or manual review prompts.

[0081] When the system determines that the termination condition is not met, the large model agent generates a next-round inquiry strategy based on the updated inquiry state representation and subsequent question generation constraints. The next-round inquiry strategy may include at least one of the following: inquiry target, inquiry question, verification basis, question granularity, expected information gain, and context triggering condition.

[0082] When the query execution degree reaches the preset completion condition, the system triggers an execution degree circuit breaker, stopping the generation of new follow-up query constraints. When the query execution degree does not reach the preset completion condition but has reached the maximum number of query rounds, the longest query duration, the maximum number of questions, or the new information gain is below a preset threshold, the system triggers an upper limit circuit breaker. After the circuit breaker is triggered, the system generates a phased query result based on the current query status. The phased query result includes at least one of the following: confirmed information, unconfirmed information, unresolved inconsistencies, relevant multimodal state offsets, corresponding contextual basis, external situational basis, and subsequent processing suggestions.

[0083] Example 5: Evidence Association Structure and Artificial Constraint Signal Generation In this embodiment, the system converts the phased query results into an evidence association structure. The evidence association structure is used to describe the correspondence between each phased result and its source information. The source information includes at least one of the following: context triggering conditions, query question representation, answer semantic representation, multimodal state vector, state offset degree, external situational constraints, and query round information.

[0084] The manual review results are converted into manual constraint signals and written into the inquiry status representation or evidence association structure together with the aforementioned evidence association structure. The manual constraint signals are used to mark the parts of the interim inquiry results that are confirmed, denied, corrected, or require further verification. In this way, the system forms a collaborative judgment result composed of machine state recursion results, evidence association structures, and manual constraint signals, enabling the system output to be reviewed, traced, and corrected, rather than directly replacing the final human judgment.

[0085] The manual constraint signals may include confirmation signals, rejection signals, correction signals, continued verification signals, or manual takeover signals. Different types of manual constraint signals can be used to confirm interim inquiry results, reject interim inquiry results, correct evidence association structures, indicate continued verification, or trigger manual takeover, respectively.

[0086] The present invention and its embodiments have been described above illustratively. This description is not restrictive, and the accompanying drawings are only one embodiment of the present invention, and are not actually limited thereto. Therefore, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the present invention, such designs should fall within the protection scope of the present invention.

Claims

1. A multimodal affective computing intelligent inquiry system based on a large model, characterized in that, It includes a context modeling module, an initial query strategy generation module, a multimodal state shift modeling module, a query state recursion and strategy update module, a query termination judgment and result generation module, and a human-machine collaborative judgment constraint and evidence representation module; among which: The context modeling module is used to obtain multi-source prior information related to the object to be queried, convert the multi-source prior information into a context representation, and construct a context association structure based on the context representation. The initial query strategy generation module is used to form initial question generation constraints based on context representation, context association structure and external situation knowledge, and call the large model agent to generate the initial query strategy. The multimodal state offset modeling module is used to acquire time-series data of at least two modalities during the query process, perform time alignment, quality assessment, normalization and joint encoding on the time-series data of the at least two modalities to obtain a multimodal state vector corresponding to the current query segment, and determine the state offset of the current query segment relative to the baseline state, intermodal consistency and matching relationship with the context representation based on the multimodal state vector. The query state recursion and strategy update module is used to encode and fuse the question representation, answer representation, multimodal state vector and the previous query state representation of each round of query to obtain the updated query state representation, and update the subsequent question generation constraints based on the updated query state representation; if the termination condition is not met, the large model agent is called to generate the next round of query strategy based on the subsequent question generation constraints. The query termination judgment and result generation module is used to calculate the current query execution degree based on the updated query status representation, and determine whether the termination condition is met based on the query execution degree, preset completion conditions and preset execution boundary conditions. When the termination condition is met, a phased query result is generated based on the current query status representation. The human-machine collaborative discrimination constraint and evidence representation module is used to convert the phased inquiry results into evidence association structures and convert the manual discrimination results into manual constraint signals, so that the manual constraint signals and evidence association structures together form collaborative discrimination results. The evidence association structure is a graph structure. ,in, Represents the set of evidence nodes. The evidence edge set is represented by evidence nodes, which include at least one of the following: question node, answer semantic node, context state node, multimodal state node, state offset node, external situation node, and artificial constraint node. Evidence edges are used to represent at least one of the following: source relationship, trigger relationship, support relationship, conflict relationship, correction relationship, or continued verification relationship.

2. The intelligent inquiry system based on a large model and multimodal emotion computing according to claim 1, characterized in that, The system also includes an inquiry interaction module, which is used to receive multi-source prior information input, carry out inquiry process interaction, and receive manual judgment result input.

3. The intelligent inquiry system based on a large model and multimodal affective computing according to claim 1, characterized in that, The multi-source prior information includes at least one of the following: identity information, itinerary information, historical behavior information, business rule information, historical interaction information, scenario constraint information, manual input information, and external situational knowledge related to the current inquiry scenario.

4. The intelligent inquiry system based on a large model and multimodal affective computing according to claim 1, characterized in that, The multimodal state migration modeling module processes at least two of the following modalities: visual modality, speech modality, text modality, millimeter-wave radar modality, infrared modality, depth image modality, thermal imaging modality, and contact or non-contact physiological signal modality.

5. The intelligent inquiry system based on a large model and multimodal emotion computing according to claim 1, characterized in that, Before performing joint encoding, the multimodal state offset modeling module performs a quality assessment on the temporal data of each modality and determines the weight of each modality in the joint encoding based on the quality assessment results.

6. The intelligent inquiry system based on a large model and multimodal affective computing according to claim 1, characterized in that, The baseline state is determined by at least one of the following: initial perception data at the start of the inquiry phase, historical statistical data, a stable time window during the same inquiry process, or a statistical baseline of similar objects.

7. The intelligent inquiry system based on a large model and multimodal affective computing according to claim 1, characterized in that, The termination conditions of the inquiry termination judgment and result generation module include at least one of the following: the current inquiry execution degree meets the preset completion condition, the preset maximum number of inquiry rounds is reached, the preset maximum inquiry duration is reached, the maximum number of questions is reached, the additional information gain that can be obtained by continuing the inquiry is lower than the preset threshold, or the conditions for manual takeover are met.

8. The intelligent inquiry system based on a large model and multimodal affective computing according to claim 1, characterized in that, The phased inquiry results include at least one of the following: confirmed information, unconfirmed information, unresolved inconsistencies, relevant multimodal state shifts, corresponding contextual basis, external situational basis, and subsequent processing suggestions.

9. An intelligent query method based on the system according to any one of claims 1-8, characterized in that, Includes the following steps: S1, obtain multi-source prior information related to the object to be queried, and convert the multi-source prior information into a contextual representation; S2, Construct a context association structure based on the context representation; S3, acquire external situational knowledge related to the current query scenario, and map the context representation, context association structure and external situational knowledge to a unified strategy generation space to form initial question generation constraints; S4, based on the initial question generation constraints, call the large model to generate the initial query strategy; S5, during the inquiry process, simultaneously acquire time-series data of at least two modalities, perform time alignment, quality assessment, normalization processing and joint encoding on the time-series data of different modalities, and obtain a multimodal state vector corresponding to the current inquiry segment; S6. Based on the multimodal state vector, determine the state offset of the current query segment relative to the baseline state, the intermodal consistency, and the matching relationship with the context representation, and use the state offset, intermodal consistency, and context matching relationship as inputs for subsequent question generation constraint updates; S7, fuse the question representation, answer representation, multimodal state vector of the current round and the query state representation of the previous round to obtain the updated query state representation, and update the subsequent question generation constraints based on the updated query state representation; S8, calculate the current query execution degree based on the updated query status representation, and determine whether to continue generating the next round of questions based on the query execution degree, preset completion conditions and preset execution boundary conditions; S9, after the termination condition is triggered, generate a phased query result based on the current query status; S10, the phased inquiry results are converted into an evidence association structure, and the manual judgment results are converted into manual constraint signals, so that the manual constraint signals and the evidence association structure together form a collaborative judgment result.

Citation Information

Patent Citations

  • Power plant production data real-time question answering and early warning method based on multi-modal large model

    CN121503667A

  • Multi-agent dynamic arrangement method based on multi-modal analysis and adaptive retrieval

    CN121809477A