Multimodal image-text semantic conflict detection and correction method based on comparative learning

By constructing a semantically faithful fracture propagation chain and modal suppression coupling recognition, the problem of weak modal information being suppressed in multimodal data fusion is solved, and the explicit recognition and correction of latent shifts are realized, restoring the integrity and consistency of semantic expression.

CN122021652AInactive Publication Date: 2026-05-12XIAMEN SHIBAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAMEN SHIBAO NETWORK TECH CO LTD
Filing Date
2026-04-14
Publication Date
2026-05-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the process of multimodal data fusion, existing technologies suppress weak modal information by strong modal information, causing semantic representation to converge toward the dominant distribution. This makes it difficult to identify and correct latent shifts, resulting in the accumulation of self-reinforcing semantic shifts. Furthermore, traditional methods rely on overall similarity discrimination and cannot effectively identify local semantic conflicts.

Method used

By constructing a semantically faithful fracture propagation chain, introducing a modal suppression coupling identification mechanism, performing pseudo-semantic consistency segmentation discrimination and chain strengthening, extracting offset propagation sub-chains, performing feature extraction and event consistency discrimination, identifying editable fields, and performing semantic write-back operations to correct the information content of semantic nodes.

Benefits of technology

It enables the identification of suppressed or absorbed differential information while maintaining overall consistency, avoiding the problem of implicit biases not being apparent. The correction process follows the original semantic propagation structure, maintains the temporal continuity within the chain, and restores the integrity of semantic expression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122021652A_ABST
    Figure CN122021652A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal image-text semantic conflict detection and correction method based on comparative learning, relates to the technical field of data processing, and aims at explicitly modeling a semantic contribution imbalance state originally implied in a multi-modal comparative fusion process by introducing a modal suppression coupling recognition mechanism into a semantic fidelity fracture propagation chain to obtain a semantic contribution imbalance state; semantic analysis does not only depend on expansion and comparison of structure levels, but can describe dominance degrees and attenuation trends of semantic nodes of different sources in a fusion process from the perspective of semantic energy distribution and path conduction relation; on the basis, conflict identification is expanded from structural difference detection to signal dominance degree judgment, so that weak dominance conflict signals which are suppressed and diluted under the condition of relatively high overall matching degree have identifiability, and the dependence path of a traditional method on dominance fine-grained conflicts is broken through.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, specifically to a multimodal graph-text semantic conflict detection and correction method based on contrastive learning. Background Technology

[0002] With the widespread application of multimodal data in intelligent search, content generation, and human-computer interaction systems, multimodal data has gradually become an important form of information expression. In practical applications, multimodal data usually exists in a structured or semi-structured form, containing multi-level semantic relationships and complex contextual dependencies. Therefore, it is necessary to use semantic modeling and structural parsing methods to transform textual and graphical information into computable semantic representations, and to perform association analysis and consistency judgment in a unified semantic space in order to identify and locate semantic conflicts.

[0003] In existing technologies, similarity-based calculations are commonly used to determine multimodal semantics. While this can quickly identify semantics, it lacks characterization of the internal propagation relationships of semantic structures. In particular, during multimodal fusion, strong modal information may suppress weak modal information. As contrastive learning continues to optimize, semantic representations gradually converge toward the dominant distribution, causing ambiguous information to be gradually compressed and become indistinguishable during path propagation.

[0004] During the consistency discrimination stage, since the judgment criteria focus on overall similarity, the compressed information is difficult to form effective discriminative features and is thus implicitly classified into the consistency region, forming a semantic state with internal deviation but external stability. This state is often not explicitly corrected in subsequent correction processes, but is further solidified by simplifying the expression or reducing the information complexity, causing the semantic carrying capacity of the weak modality to continue to decline. As this process iterates, the influence of the dominant modality is continuously strengthened, while the expression of the weak modality gradually degenerates, eventually forming a hidden and self-reinforcing evolutionary path, causing the semantic deviation to accumulate continuously at the structural level and be difficult to perceive or correct in a timely manner. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a multimodal graph-text semantic conflict detection and correction method based on contrastive learning, which solves the problems mentioned in the background technology.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] Multimodal graph-text semantic conflict detection and correction methods based on contrastive learning include:

[0008] Multimodal image and text data is acquired, and data parsing is performed based on contrastive learning-driven natural language processing technology to obtain a semantically faithful break propagation chain.

[0009] Reconstruction of semantically fidelity-preserving broken propagation chains is performed through modal suppression coupling recognition;

[0010] Based on the pseudo-semantic consistency segmentation and chain strengthening of the reconstructed semantic fidelity broken propagation chain, the offset propagation sub-chain is obtained.

[0011] Feature extraction is performed on the offset propagation sub-chain to obtain the time offset chain segment, and event consistency judgment and chain structure update are performed on the semantic nodes to obtain the mapping classification results, including asynchronous chain segments and real conflict chain segments.

[0012] If it is an asynchronous link segment, the information content of the semantic node is extracted and corrected, and after differential processing, the editable field is identified;

[0013] Perform semantic write-back operations based on editable fields.

[0014] The above-described solution of the present invention has at least the following beneficial effects:

[0015] This scheme introduces a modal suppression coupling identification mechanism into the semantic fidelity fracture propagation chain, explicitly modeling the semantic contribution imbalance state that was originally implicit in the multimodal comparison and fusion process. This allows semantic analysis to no longer rely solely on the unfolding and comparison of structural levels, but to characterize the dominance and attenuation trend of semantic nodes from different sources in the fusion process from the perspective of semantic energy distribution and path transmission relationship. On this basis, conflict identification is extended from structural difference detection to signal dominance discrimination, making weakly dominant conflict signals that are suppressed and diluted under the condition of high overall matching degree identifiable, thereby breaking through the traditional method's reliance on the path of dominant fine-grained conflict.

[0016] By constructing a pseudo-semantic consistency segmentation discrimination and offset propagation sub-chain reinforcement mechanism, local semantic differences, global consistency, and path continuity are coupled and analyzed. This allows semantic conflicts to no longer be judged based on single points or local structures, but rather to identify their cumulative and hidden characteristics during chain propagation. This approach can distinguish between true consistency and apparent consistency caused by modal suppression, avoid misclassifying semantic chain segments with offsets into the consistency region, and reconstruct the transmission trajectory of conflict signals at the propagation path level, transforming the identification of semantic conflicts from static structural comparison to dynamic path analysis.

[0017] In the semantic correction stage, by introducing information theory-based information content calculation and differential discrimination mechanism, the information density change of semantic nodes is quantitatively analyzed, and the node probability is corrected by combining modality suppression index and node weight, so that the semantic rewrite process is based on information fidelity constraints. At the same time, through editable domain filtering and semantic rewrite based on boundary anchor points, the local reconstruction of semantic nodes and their stage information is realized, so that the correction process follows the original semantic propagation structure and maintains the temporal continuity within the chain, thereby forming a traceable and locatable semantic update path. Attached Figure Description

[0018] Figure 1 This is a block diagram of the method of the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] like Figure 1 As shown, embodiments of the present invention provide a multimodal graph-text semantic conflict detection and correction method based on contrastive learning, including:

[0021] S1: Acquire multimodal image and text data, and perform data parsing using natural language processing techniques to obtain a semantically faithful fracture propagation chain;

[0022] S2: Reconstruct the semantically faithful broken propagation chain by identifying modal suppression coupling;

[0023] S3: Based on the pseudo-semantic consistency segmentation and chain strengthening of the reconstructed semantic fidelity broken propagation chain, the offset propagation sub-chain is obtained;

[0024] S4: Extract features from the offset propagation subchain to obtain the time offset chain segment, and perform event consistency judgment and chain structure update on the semantic node to obtain the mapping classification result, including asynchronous chain segment and real conflict chain segment;

[0025] S5: If it is an asynchronous chain segment, extract and correct the information content of the semantic node, and identify the editable field after differential processing;

[0026] S6: Perform semantic write-back operations based on editable fields.

[0027] The semantic fidelity break propagation chain is constructed through the above steps. On this basis, a modality suppression coupling recognition and chain-level reconstruction mechanism are introduced, which transforms semantic analysis from a single similarity judgment to a structured propagation analysis along the semantic path. Thus, even with overall consistency, the suppressed or absorbed differential information can still be identified, avoiding the problem that the latent shift cannot be revealed due to the dominance of the dominant modality.

[0028] By using pseudo-semantic consistency segmentation discrimination and offset propagation sub-chain reinforcement, the offset state that was originally masked during local propagation is separated from the chain structure. Combined with the construction of time offset chain segments, the distinction between event stage misalignment and real conflict is realized. This allows asynchronous expressions that were originally easily misjudged as conflicts to be identified separately, thereby simultaneously solving the problem of pseudo-semantic consistency that appears consistent but has internal offsets and the discrimination bias problem of expressions at different stages being misjudged as conflicts.

[0029] Based on this, semantic nodes are constrained and filtered through information content correction and differential processing. Only regions with abnormal changes in information density are located, and semantic write-back is performed in conjunction with editable fields. This makes the correction process no longer dependent on overall reconstruction, but focuses on key nodes in the chain that have shifted. This avoids the information weakening problem caused by reducing the complexity of expression to maintain consistency. For example, in the description of someone just picking up a cup and someone drinking water, the system no longer directly judges it as a conflict or simply weakens it to someone taking an action. Instead, it identifies that they are in different stages of the same event chain and only makes local corrections to the time stage, so that the semantic expression remains complete and the structural consistency is restored. Thus, in the same processing flow, the system simultaneously suppresses the offset accumulation caused by modality suppression, eliminates misjudgments caused by pseudo-semantic consistency, and avoids the continuous loss of information during the correction process.

[0030] In a preferred embodiment of the present invention, S1 includes: semantically decomposing the text descriptions in multimodal image and text data from a database based on natural language processing technology to obtain text nodes; performing structural parsing on non-text descriptions through structured semantic modeling and indexing to obtain structure nodes; and mapping the text nodes and structure nodes to the same semantic vector space to obtain a joint semantic primitive set; this set is a unified semantic node set formed by vectorizing and indexing text nodes and structure nodes, serving as the smallest data unit for subsequent semantic matching, association domain construction, and breakage analysis;

[0031] For text descriptions in image-text data, pre-trained language models (such as BERT, RoBERTa, etc.) are used for processing. The model converts each word or character in the text into a corresponding word vector, and then uses the model's context encoding capabilities to generate the semantic vector of the entire sentence or text segment. For the image portion of the image-text data, pre-trained visual models (such as ResNet, ViT, etc.) are used for processing. The model extracts the visual features of the image and converts them into image semantic vectors.

[0032] In practice, the input image and text data undergoes unified semantic parsing processing. This includes term-level decomposition, syntactic dependency parsing, and semantic role labeling of the text description, transforming the sentence structure into a semantic graph structure in the form of text nodes-relationships. Text nodes are then divided into term nodes, relation nodes, and event nodes. Simultaneously, the non-textual expressions in the image and text are structurally parsed, extracting object units (entities), attribute units (feature descriptions), and interaction relationships (actions or logical connections) as structural nodes (image semantic vectors). These structural nodes are then semantically aligned with the text nodes (text semantic vectors) through a cross-modal encoder to obtain semantic vectors of the same dimension. This process ensures comparability between the two types of nodes within a unified semantic space. Subsequently, all semantic nodes are vectorized and encoded, with three types of indexes appended to each node: a structural hierarchy index representing the node's depth within the semantic structure (e.g., sentence = 0, event = 1); a path index recording the connection paths and lengths between the node and other nodes; and an original position index identifying the node's order or spatial location within the original text and image. Finally, the node and index information are organized into a joint semantic primitive set and its index table to construct cross-semantic relationships. Semantic nodes include text nodes and structural nodes.

[0033] Structured semantic modeling and indexing is a modeling and encoding method that extracts structural information such as objects, attributes, and interaction relationships from non-text data such as images, and adds three types of indexes—hierarchy, path, and location—to each structural unit. It is used to transform non-text data into a structured expression that is compatible with text semantics.

[0034] Based on a joint set of semantic primitives, semantic node pairs are identified through comparative learning of semantic matching. Path expansion analysis is then performed using a connected component algorithm to construct a semantic association domain and generate joint position coordinates. These coordinates are used not only to locate semantic nodes but also to describe the mapping relationship of semantic nodes in different structures, thus supporting subsequent position correction.

[0035] In practice, cosine similarity is used to perform a one-to-one matching operation on all different semantic node pairs in the joint semantic primitive set. Consistency constraints are applied to their structural hierarchy index and path index (the hierarchy and path structure must be consistent) to filter out node pairs with semantic correspondence. The cosine similarity must exceed a preset similarity threshold. Then, nodes with corresponding relationships are aggregated according to their connection paths to form a semantic association domain. Within the association domain, direct path relationships between nodes and indirect path relationships formed through intermediate nodes are recorded, thus constructing a multi-path association structure. Based on this, the hierarchy index, path index, and original position index of each node are fused and encoded to form a joint position coordinate in the form of a triple, used to uniquely identify the composite position of the node in the structural space, path space, and original space. The joint position coordinates are then embedded into the association domain structure to form a unified mapping matrix of "association domain—node—coordinates," used for semantic energy analysis and break detection.

[0036] The semantic association domain is formed by the aggregation of multiple sets of semantic node pairs. It reflects the local structure of semantic equivalence in multimodal graph and text data and serves as the basis for subsequent energy analysis.

[0037] Based on semantic association domains, semantic energy distribution modeling and gradient change analysis are used to identify energy anomaly domains and generate a set of fracture candidate regions. This set is used to exclude isolated noise points, retaining only structurally continuous and path-coherent reliable fracture regions as direct input for subsequent propagation chain construction; specifically including:

[0038] Combining text nodes and structure nodes yields semantic nodes;

[0039] Extract the vector response intensity of semantic nodes and use it as the semantic energy value;

[0040] Based on the semantic association domain, the energy change between adjacent semantic nodes is calculated to obtain the gradient sequence. Each feature in the gradient sequence is compared with a preset gradient threshold to obtain the break nodes and generate an energy anomaly domain. This anomaly domain is used to locate continuous node intervals where semantic transmission has drastic changes, semantic support has been significantly weakened, or even semantic breaks have occurred from the perspective of semantic energy change.

[0041] The continuity of the energy anomaly domain is verified by combining the location coordinates to generate a set of fracture candidate regions.

[0042] The process involves constructing a node energy distribution within the associated domain, then normalizing the energy distribution to make it comparable between different associated domains. Based on this, the energy changes between adjacent nodes within the associated domain are differentially calculated to construct a gradient sequence. If the absolute value of the gradient exceeds a preset gradient threshold, it is marked as a broken node; otherwise, it is not marked. Nodes with consecutive (at least two consecutive sets of nodes) gradient anomalies are then aggregated to obtain an energy anomaly domain.

[0043] By combining location coordinates, only anomaly domains that are continuous in both structure and path are retained to form a set of fracture candidate regions. This set is used to locate the starting region of semantic fracture propagation, and at the same time, the node index range, path span and energy change trajectory of each candidate region are recorded.

[0044] The anomaly domain is the result of the signal layer; the candidate region is the effective fracture region after structural constraints, i.e., a problem segment.

[0045] Vector response strength is the semantic activity of a semantic node in the vector space, i.e., the magnitude of the semantic vector. It reflects the weight and contribution of the semantic node in image-text matching. High energy indicates that the node has a high degree of matching and strong support; low energy or a sharp drop indicates that the node has failed to match and is a breakpoint.

[0046] The node energy distribution is the result of arranging the semantic energy of all nodes in the path order within the semantic association domain. This prepares for the next step of calculating energy difference and is used to locate energy anomaly domains.

[0047] Based on a set of fracture candidate regions, a semantically faithful fracture propagation chain and its indexing system are constructed through path tracing and structural connection. This propagation chain characterizes the contribution distribution, path transmission relationship, and conflict signal attenuation state of semantic nodes of different modalities during image-text fusion. It is used to find the source of conflict and provides a precise operation object and constraint framework for subsequent nonlinear stretching, path rearrangement, and local reconstruction.

[0048] In practice, path-level tracing is performed on the set of fracture candidate regions. First, the connection paths between nodes within the candidate regions are extracted based on the path index, and the search is expanded along the path direction to connect nodes with continuous energy conduction relationships. Then, based on the path connection relationship, the fracture nodes are organized into a chain structure according to the propagation order, and the starting node, intermediate propagation nodes, and ending node of the chain are marked. During the construction process, the joint position coordinates are embedded in the chain structure, so that each node has a unique coordinate identifier in the chain, and its sequence position in the chain is recorded. At the same time, the chain structure is deduplicated and merged to form a set of independent fracture propagation chains. Finally, an index table is constructed, which contains chain number, node sequence, path information, and coordinate mapping relationship.

[0049] By mapping text nodes and structural nodes to the same semantic vector space and constructing a joint semantic primitive set by combining three types of indices—hierarchy, path, and original position—semantic information that was originally scattered in different forms of expression is structurally aligned. At the same time, based on semantic association domain and energy gradient changes, nodes are subjected to continuity constraints and screening, thereby transforming latent offsets that are difficult to show under global similarity into candidate fracture regions that can be traced along the path. Furthermore, a semantically faithful fracture propagation chain is formed through path tracing, so that semantic conflicts no longer remain at point-like discrimination but can be located and restored in the propagation relationship.

[0050] This scheme constructs a semantically fidelity-preserving fracture propagation chain and combines energy distribution, path constraints, and time stage joint modeling to continuously track and make explicit the information of dependency expression and fuzzy expression in the propagation path. Therefore, it no longer relies on single point or local features for judgment, but makes a comprehensive judgment based on structural continuity and evolutionary consistency. Compared with the method of analyzing only local differences, this scheme can reveal the cumulative offset and stage misalignment of semantics in the transmission process, so that anomalies that were originally hidden under the overall consistency can be identified and located.

[0051] In a preferred embodiment of the present invention, S2 includes constructing a chain-level semantic contribution distribution matrix based on semantic fidelity break propagation chains through semantic energy statistics and distribution analysis within the chains; specifically, this matrix is ​​a matrix with each semantic fidelity break propagation chain as a row and the semantic energy value of each semantic node on the chain as a column, used to describe the flow of semantic energy in the break path, and as input for subsequent suppression identification, that is, to find out who is suppressed by whom, where the semantic imbalance is, and where the conflict is transmitted from.

[0052] In practice, the semantic fidelity break propagation chain is processed one by one. For each propagation chain, the semantic energy value is extracted according to the node sequence, and an energy sequence is constructed based on the order of the nodes in the chain. Then, the energy sequence is statistically analyzed to calculate the average energy, energy variance and local fluctuations in the chain, and these statistics are mapped to the node positions to form a node-level energy distribution vector. In this way, each node is a vector with contextual information, rather than just an isolated number.

[0053] Based on this, the semantic energy value is decomposed according to the node source type to obtain the contribution ratio of different source semantics in the chain (the ratio of semantic energy values ​​of different modalities); at the same time, combined with path information, the trend of energy change in the propagation direction is recorded to form a chain-level semantic contribution distribution matrix; this matrix has "chain number - node sequence number - energy value" as its basic structure and is used for suppression identification.

[0054] Specifically, the break does not appear suddenly. It is often caused by an abnormal energy at one node, which is transmitted to the next node, causing the entire chain to collapse. Therefore, we need to observe the energy flow to understand how the break spreads, observe the comparison of energy strength to understand whether text is suppressing the image or the image is suppressing the text, and observe the trend of energy change to understand where the break starts and where the secondary effects are. The ultimate goal is to locate the position and direction of modal suppression and find the propagation path of semantic conflict, so as to provide a basis for subsequent stretching correction, path rearrangement, and conflict repair.

[0055] Based on the chain-level semantic contribution distribution matrix, a modal suppression index is obtained by calculating the degree of contribution deviation. This value is used to correct the explicitness of local semantic differences and represents the overall degree of semantic suppression of the entire semantic fidelity break propagation chain. The more negative the value, the more severe the overall suppression of the chain; the closer to 0, the more balanced the semantics; and the greater than 0, the stronger the overall contribution of the chain, indicating that it is suppressing others. Combined with path constraint analysis, suppression propagation intervals with consecutive contribution offset directions are identified, resulting in a suppression propagation sub-chain. This sub-chain is a small segment of continuous nodes extracted from the original semantic fidelity break propagation chain.

[0056] The essence of contribution deviation is the inconsistency or imbalance of multimodal semantic information, mainly caused by missing or ambiguous modal information, model bias, information conflict between modalities, and data quality issues. Therefore, contribution deviation is a direct manifestation of multimodality in identifying strong and weak modalities.

[0057] Specifically, the difference calculation is performed on the chain-level semantic contribution distribution matrix. For each node, the deviation from the chain average contribution is calculated to obtain the deviation value. A positive value indicates that the node's contribution is relatively strong, and a negative value indicates that the node's contribution is relatively weak (there is a risk of being suppressed). The deviation values ​​are accumulated to form the modal suppression index of the corresponding node.

[0058] Since single-point offsets may be noise, only when a continuous series of nodes offset in the same direction can it be considered true suppression propagation. Therefore, the node sequence is traversed to identify the node intervals where the offset values ​​are all in the same negative direction, and these intervals are marked as suppression propagation intervals. During the identification process, a path index is introduced as a constraint to retain only intervals where the path is continuous and the energy change is monotonic. These intervals are then extracted to form suppression propagation subchains, and their node positions and joint position coordinates in the semantically faithful broken propagation chain are recorded.

[0059] The suppression propagation range is used to locate the specific location where modal suppression occurs, to know whether it is text suppressing an image or an image suppressing text, and to provide a precise operating range for subsequent stretching, rearrangement and correction.

[0060] Based on the suppression propagation subchain and modal suppression index, a reconstructed semantically faithful break propagation chain is obtained through nonlinear semantic stretching and path structure rearrangement. This reconstructed propagation chain is used to eliminate the modal suppression imbalance and semantic energy anomaly in the original propagation chain, restore the normal energy transmission and path connection relationship between semantic nodes, so as to achieve semantic balance between modalities, avoid break misjudgment, conflict omission detection and repair deviation caused by modal strength imbalance, and provide a reliable, faithful and traceable processing carrier for all subsequent refined analysis and correction operations.

[0061] Specifically, the node weights of the suppressed propagation subchain are adjusted, and the semantic vectors of the nodes are nonlinearly transformed according to the node suppression index, so that the semantic responses of the suppressed nodes are redistributed; for example, an increasing nonlinear function, such as an exponential function, can be used. Enlarging the magnitude of the semantic vector is intended to enhance the semantic contribution of the suppressed node and strengthen its voice. For the new semantic vector, This is the original semantic vector of the node. is the base of the exponential function. The absolute value of the modal suppression index; The amplification factor is a hyperparameter greater than 0, which semantically controls the intensity of nonlinear amplification. The larger the value, the more significant the amplification effect on the suppressed nodes. The optimal value can be efficiently found using automated methods such as grid search or random search. value.

[0062] Subsequently, the chain structure is partially rearranged. Without disrupting the overall path structure, some suppression paths are broken (e.g., a path where a strong node directly connects to and suppresses multiple weak nodes), and reconnected to other nodes within or adjacent suppression propagation intervals to reconstruct inter-node paths. During the rearrangement process, joint position coordinates serve as constraints, allowing structural adjustments only within adjacent node intervals to ensure structural continuity. The adjusted node sequence is then re-embedded into the semantically faithful broken propagation chain to form a new propagation chain structure.

[0063] Specifically, the basis for partially breaking the suppression path is the suppression relationship where the absolute difference between the semantic energy values ​​of the nodes at both ends of the path exceeds a preset threshold.

[0064] It should be noted that existing image-text semantic conflict detection schemes mainly focus on the hierarchical and structural relationships of conflict content, emphasizing the identification of explicit fine-grained conflict features, while paying less attention to the decline in conflict visibility caused by modality dominance, feature suppression, and changes in signal strength during multimodal feature fusion. When the overall image-text matching degree is high and one modality's semantics is dominant, conflict features in another modality may be suppressed, diluted, and transformed into weak explicit conflict signals during the fusion process. These signals are not judged based on structural complexity but are primarily characterized by the attenuation of explicitness during fusion, thus differing from fine-grained conflicts in the usual sense. If such weak explicit conflict signals are not identified, subsequent consistency judgments easily classify them into pseudo-consistency regions, further weakening conflict cues through semantic generalization during the correction process, forming a chain of anomalies coupled with modality suppression, pseudo-consistency, and information impoverishment. Therefore, it is necessary to identify strong and weak nodes in this situation.

[0065] By constructing a chain-level semantic contribution distribution matrix on the semantic fidelity fracture propagation chain and decoupling the statistical distribution and source ratio of node energy, the semantic alignment process, which originally only stayed at the level of overall similarity, is transformed into an observable energy flow process along the path. This enables the identification of the specific location and direction of the gradual imbalance of semantics during propagation.

[0066] Based on this, the modal suppression index is obtained by calculating the contribution deviation, and the suppression propagation subchain is extracted by combining the path continuity constraint, so that single-point noise no longer interferes with the judgment, but only the continuous imbalance interval is located. This solves the problem that traditional methods cannot distinguish between local fluctuations and continuous suppression, and also avoids the accumulation of implicit shifts where strong expressions dominate for a long time while weak expressions are continuously weakened.

[0067] Furthermore, by performing nonlinear semantic stretching and local path rearrangement on the suppressed propagation subchain, the semantic response of the suppressed node is redistributed, and unreasonable suppression paths are structurally adjusted. Thus, the semantic balance within the chain is restored without destroying the overall semantic structure. The system no longer simply propagates strong semantics along the original path, but instead identifies the suppression interval and strengthens weak nodes, so that the stage differences are preserved and re-embedded into the propagation chain. This not only suppresses the semantic imbalance caused by modal suppression, but also repairs the structural offset problem caused by suppression diffusion.

[0068] In a preferred embodiment of the present invention, by analyzing the semantic differences between adjacent semantic nodes within the reconstructed semantic fidelity break propagation chain, a sequence of differences between nodes is obtained, including several sets of cosine difference values. The cosine difference value is the result of 1 minus the cosine similarity. This sequence is used to show the distribution of semantic differences on the entire chain, helping to identify intervals with significant differences. The joint position coordinates are used as a continuity constraint to obtain a set of difference segments. This set is used as the input for subsequent pseudo-consistency discrimination and temporal semantic analysis to locate semantic problem areas that need further processing.

[0069] Specifically, for each semantically faithful break propagation chain, the node sequence is traversed sequentially, and the semantic vectors of adjacent node pairs are extracted and the cosine difference value is calculated to form a difference sequence between nodes.

[0070] Subsequently, a sliding window scan is performed on the difference sequence to identify the node intervals where the continuous differences exceed the preset threshold. This interval indicates that the differences between semantic nodes on this path are very significant, and there may be semantic breaks or inconsistencies, so as to initially form candidate difference intervals.

[0071] Furthermore, path index and joint location coordinates are introduced as constraints to filter candidate intervals of difference, retaining only node intervals that are unbroken in the joint coordinate space (continuous in structural hierarchy, path index, and original location coordinates); then, the intervals that meet the conditions are aggregated to form a set of difference segments, and the starting node index, ending node index, path span, and corresponding joint location coordinate interval are recorded for each segment.

[0072] Based on the differential segment set, a pseudo-semantic consistent segment set is obtained through comparative analysis of local consistency and global consistency, including multiple sets of pseudo-semantic consistent segments. This set is the core input for subsequent nonlinear enhancement and pseudo-consistency discrimination, and is used to further identify and process semantic segments that seem to conflict but are actually consistent.

[0073] Specifically, the differential segment set is processed one by one. First, the local semantic consistency is calculated within each segment, which is the average cosine similarity between the semantic vectors of all nodes within the segment. This is used to measure whether the semantics within the segment are coherent and consistent.

[0074] Subsequently, global consistency is calculated within the entire semantic fidelity break propagation chain to which the segment belongs. Specifically, cosine similarity is calculated between each node within the segment and all other nodes in the semantic fidelity break propagation chain (nodes outside the segment), and then the average is taken. This average is used to measure the overall semantic relevance of the segment to the entire semantic propagation chain.

[0075] Based on this, a comparative analysis of local consistency and global consistency is performed to identify segments with significant local differences but overall consistency, i.e., segments with local consistency below a preset local threshold and global consistency above a preset global threshold. Subsequently, segments that meet the conditions are marked as pseudo-semantic consistent segments, and their segment number, node index range, and joint position coordinate interval in the semantic fidelity break propagation chain are recorded.

[0076] Extract the node sub-chains corresponding to each pseudo-semantic consistent segment in the pseudo-semantic consistent segment set, and use them as the initial propagation path to determine the starting point and core region of the path expansion.

[0077] Using energy continuity and path reachability as search conditions, a neighborhood search is performed on the initial propagation path, and after path expansion, a propagation sub-chain is obtained. This sub-chain is an extension and enhancement of the single initial propagation path, used to capture more semantic information related to pseudo-semantic consistent segments, and to ensure the integrity of the propagation sub-chain.

[0078] Energy continuity means that the semantic energy value of a node is continuous with the energy value of nodes in the initial propagation path, that is, the absolute difference of the semantic energy value does not exceed a preset threshold, thus avoiding the introduction of irrelevant or conflicting information; path reachability means that a node is reachable in the path structure of the semantic fidelity broken propagation chain, that is, it can be connected to the initial propagation path through the path, thus ensuring the structural integrity of the propagation sub-chain.

[0079] In a semantic propagation chain, the position of a node is often related to its importance. Nodes located at the center of a segment typically carry more core semantic information, while peripheral nodes are relatively less important. Therefore, positional weights can reflect the importance of a node to some extent.

[0080] Based on the propagation subchain, the relative positional relationship of each semantic node in the corresponding pseudo-semantic consistent segment is identified, and the semantic node weight and offset propagation subchain are obtained; this subchain can reflect information of superficially inconsistent but essentially consistent features.

[0081] In subsequent pseudo-semantic consistency discrimination and temporal semantic analysis, nodes with higher weights will be given more attention, thereby improving the accuracy and relevance of the analysis.

[0082] The node subchain is formed relative to the semantically faithful break propagation chain; neighborhood search refers to searching around the adjacent nodes at both ends of the pseudo-semantic consistent segment;

[0083] Specifically, for pseudo-semantic consistency segments, a path expansion operation is performed. First, the node sub-chain corresponding to each segment is extracted from the semantic fidelity break propagation chain and used as the initial propagation path. Then, based on the energy distribution, a neighborhood search is performed on the adjacent nodes at both ends of the segment, and nodes that satisfy energy continuity and path reachability are included in the propagation path, thereby forming the expanded propagation sub-chain.

[0084] During the construction process, each node is assigned a weight, which is determined by the position of each node in the propagation subchain within its corresponding pseudo-semantic consistency segment. The weighted node sequence is then organized into an offset propagation subchain, i.e., a weighted propagation subchain. This allows subsequent pseudo-semantic consistency discrimination and temporal semantic analysis to more accurately focus on core nodes and key information with higher weights. (The formula is used to...) Obtain the semantic node weights, where For the first Semantic node weights of each semantic node, and the index of the center node of the pseudo-semantic consistent segment. , To round down, This represents the total number of nodes within the pseudo-semantically consistent segment.

[0085] By performing cosine difference-based serialization analysis on the semantically fidelity-preserving break propagation chain and combining joint position coordinates to impose structural continuity constraints on the difference intervals, the semantic differences that were originally scattered between nodes are organized into difference segments with path significance, thereby avoiding misjudging isolated fluctuations as structural problems. Furthermore, through the comparison mechanism of local consistency and global consistency, the chain segments that show local differences but whose overall semantics remain related are separated from real conflicts, thus solving the defect of misjudging such segments as conflicts in traditional methods, and also avoiding the problem of semantic information reduction caused by over-correction.

[0086] Building upon this, the initial path is extended using energy continuity and path reachability constraints, and offset propagation sub-chains are constructed by combining position weights. This allows semantic judgment to no longer rely on uniform processing, but instead unfold around the core nodes of the chain segment, thereby improving the ability to preserve key semantics and the accuracy of discrimination.

[0087] While local actions may differ, the overall events remain consistent. The system can identify the essential relationships by concentrating weights on the central node of the chain segment, thus avoiding misjudgment as conflict and preventing the semantics from being simplified into a generalized description. At the same time, it overcomes two types of problems: misidentification of local differences and loss of structural information.

[0088] In a preferred embodiment of the present invention, S4 includes: extracting the temporal semantic units of each semantic node in the offset propagation subchain, and obtaining a temporal semantic sequence after node mapping, which is used to identify the distribution of strong and weak signals;

[0089] First, the semantic segments corresponding to the chain nodes are located in the text semantic structure, and time indicators, stage descriptions, and sequential relationship expressions are extracted from them. Then, these time information are mapped into time semantic units, and the time units are sorted according to the dependencies in the original semantic structure to form a local time series. In this process, the joint position coordinates are used as constraints to ensure that the time semantic units are mapped in a consistent manner with the node positions in the semantic fidelity break propagation chain. Finally, the time semantic units are mapped one-to-one with the chain nodes to generate a time semantic sequence.

[0090] Stage descriptors are extracted, and the temporal semantic sequence is segmented and labeled with nodes based on the stage descriptors to construct a stage matrix. This matrix is ​​used to provide a basis for temporal cross-sectional offset analysis. By comparing the node semantics of different stages, semantic offsets and conflicts in the time dimension can be identified.

[0091] Stage descriptors are words or phrases that characterize the development stage or temporal features of an event. They are the core basis for dividing the stage intervals of a temporal semantic sequence and the basic identifiers for constructing a stage matrix.

[0092] The temporal semantic sequence is divided into stages. First, the sequence is segmented according to the stage descriptors in the temporal semantic units, and different stages are divided into discrete intervals. Then, the stage information is mapped to the nodes in the semantic fidelity break propagation chain, each node is assigned a stage label, and its sequential position in the stage is recorded.

[0093] For nodes that do not explicitly contain time information, stage inference is performed by combining their position in the chain and the stage information of adjacent nodes. Here, interpolation is used to ensure that all nodes have stage identifiers. Based on this, the node stage information is organized into a matrix structure, where the matrix rows are pseudo-semantic consistent chain segments, the columns are node numbers, and the elements represent stage labels.

[0094] Based on the stage matrix, time difference calculation and path mapping analysis are performed on each semantic node to obtain a time offset chain segment; only node intervals with both semantic offset and time offset are retained in this offset chain segment, which are used to locate node intervals with abnormal offset in the time dimension.

[0095] Specifically, the stage matrix is ​​differentially calculated. First, the stage label of each node is converted into a numerical value through encoding. Then, the difference between its stage label and the reference stage is calculated. By quantizing the time offset, abnormal situations of nodes in the time dimension are identified, forming a node-level time offset sequence.

[0096] Based on this, the node-level time offset sequence is mapped back to the semantically faithful break propagation chain structure, and node intervals with consecutive time offset values ​​exceeding the threshold in the time dimension are identified and aggregated to form time offset chain segments; the node-level time offset sequence includes multiple sets of time offset values; here, anomalies are found from the time dimension.

[0097] In this invention, all parameters are dimensionless by using dimensionless processing technology to remove their dimensions; and all thresholds can be obtained by the mean-standard deviation method.

[0098] Based on time-offset chain segments, event consistency judgment and chain structure update are performed on semantic nodes to obtain mapping classification results. These mapping classification results guide subsequent semantic repair strategies, avoiding blind processing and improving the accuracy and efficiency of semantic repair; specifically including:

[0099] Set classification constraints and use these constraints to make a comprehensive judgment on the time offset chain segments;

[0100] The classification constraints include path continuity determination and stage sequence consistency determination. Path continuity ensures that nodes within a chain segment are coherent in the semantic propagation path and belong to the same semantic structure. Stage sequence consistency ensures that the development of nodes within a chain segment in the time dimension conforms to the logic of a single event and does not exhibit logical inconsistencies. It should be noted that the path continuity determination here verifies anomalies from a structural dimension.

[0101] If a semantic node is discontinuous in the reconstructed semantic fidelity break propagation chain path structure or the stage order does not conform to the single event progression logic, the corresponding time offset chain segment is marked as a real conflict chain segment, and a fidelity constraint instruction is triggered. If a semantic node is continuous in the reconstructed semantic fidelity break propagation chain path structure and the stage order conforms to the single event progression logic, the corresponding time offset chain segment is marked as an asynchronous chain segment and corrected. If the corrected asynchronous chain segment no longer has time offsets or semantic conflicts, the processing ends. If conflicts still exist, it is marked as a real conflict chain segment.

[0102] The asynchronous chain segment and the real conflict chain segment are combined to obtain the classification result. The classification result is then mapped back into the reconstructed semantic fidelity break propagation chain to obtain the mapping classification result, so that its structure contains temporal semantic information.

[0103] An asynchronous chain segment refers to a time-shifted chain segment in the reconstructed semantic fidelity break propagation chain where the semantic node paths are continuous, the overall stage sequence conforms to the same event progression logic, and the time shift is caused only by ambiguous, delayed, or advanced time descriptions, thus triggering secondary and recoverable semantic incoherence. This time-shifted chain segment can be corrected by adjusting the time description using linear interpolation.

[0104] A real conflict chain segment refers to a chain segment where nodes do not belong to the same event chain and there is a real temporal semantic conflict. Such conflicts are usually caused by inconsistencies or errors in the event descriptions in the image and text data. Semantic correction needs to be performed first, and then stage correction can be considered.

[0105] The semantic fidelity break propagation chain path structure refers to the connection relationship and order between nodes in the semantic fidelity break propagation chain. It reflects the path and structure of semantics in the propagation process, including the order of nodes, connection methods, and hierarchical relationships.

[0106] By introducing temporal semantic units and constructing a stage matrix on the offset propagation subchain, the node relationships that were originally expressed only in the semantic space are extended to the time dimension for joint modeling, so that semantic offsets are no longer judged in isolation, but can be verified in the event development chain.

[0107] Based on this, time offset segments are extracted by stage difference and path mapping, and dual constraint discrimination is performed by combining path continuity and stage sequence consistency. This effectively distinguishes segments that belong to the same event but are in different stages from segments that do not belong to the same event in essence. This solves the defect of misjudging time section differences as semantic conflicts in traditional methods, and avoids the problem of incorrect correction caused by ignoring time logic based solely on semantic similarity.

[0108] Furthermore, by writing the classification results back into the propagation chain structure, the chain structure is given a temporal semantic identifier, which can guide the selection of subsequent processing strategies under the dual constraints of structure and time.

[0109] In a preferred embodiment of the present invention, S5 includes: receiving a fidelity constraint instruction, extracting the node probability of each semantic node in the real conflict chain segment, and using the product of the modality suppression index and the semantic node weight as a correction term to correct the node probability; and combining the concept of self-information in information theory to calculate the information content of each semantic node, the larger the value, the rarer and more important the node is in the chain segment.

[0110] For nodes in the actual conflict chain, adjusting their probability increases, leading to a decrease in information content. Reduced information content means lower node rarity, but this doesn't necessarily mean lower importance. In fact, by adjusting the probability, we artificially increase the node's occurrence probability, giving it greater weight in information density calculations and thus strengthening its importance to avoid information impoverishment. That is, by adjusting the probability, we can enhance the information density of nodes, thereby preventing information impoverishment and ensuring that key information can be accurately identified. Information impoverishment typically refers to excessively low information density, leading to the loss of key information.

[0111] The purpose of adjusting node probabilities is not simply to reduce the amount of information by increasing the probability, but to strengthen the importance of core nodes and improve the accuracy of information density analysis by adjusting node probabilities in a targeted manner, thereby avoiding information depletion.

[0112] Semantic node weights reflect the structural importance of a node in the chain, while modal suppression indexes reflect the intensity of conflict that a node is suppressed. By using semantic node weights as weighting coefficients for the suppression indexes, we can control which nodes' suppression is more worth amplifying. The overall logic is self-consistent and without loopholes.

[0113] Through formula Obtain the corrected node probabilities, where For the first The corrected node probabilities of semantic nodes in the actual conflict chain segment. The original node probabilities, This represents the semantic node weight, with a value greater than 0. The modal suppression index;

[0114] The formula for calculating the information content after node probability correction is: ,in This is the amount of information after node probability correction. It is a logarithmic function with base 2. For the index of semantic nodes; For the first One semantic node;

[0115] The information content of adjacent semantic nodes is differentially divided to obtain a differential sequence, which is then compared with a preset change threshold to identify editable domains, that is, the set of conflict areas that can be edited and repaired.

[0116] Specifically, by scanning the sequence of information density changes through a sliding window, the node intervals where continuous changes exceed the threshold are identified and treated as editable fields. These fields reflect drastic information abrupt changes, indicating that semantic transmission is not smooth and are likely to be conflicting, broken, or abnormal regions.

[0117] The differential sequence contains multiple sets of information density variations, which are used to identify information break segments in order to determine the editable domain;

[0118] S6 includes: obtaining the node position index and coordinate mapping relationship based on the editable field and the joint position coordinates;

[0119] The coordinate mapping relationship is a correspondence table. The left side is the number index of the semantic node in the semantic fidelity break propagation chain, and the right side is the actual position coordinates of the node in the original image and text.

[0120] Based on node position index, the front and back boundary nodes of the editable field are located, and the semantic vectors and stages of the boundary nodes are extracted respectively as front and back anchor points. Combined with arithmetic mean interpolation, new semantic vectors are generated to eliminate semantic bias and modal conflict.

[0121] Here, the boundary node refers to the normal node preceding the editable domain, that is, the node that has not been affected by semantic or temporal offsets;

[0122] Based on the new semantic vector, linear interpolation is used. The method corrects the stages of semantic nodes within the editable domain to obtain new stages, making them sequentially continuous and smooth, and repairing time offsets; among which... This is the corrected stage mask, i.e., the new stage; This is the stage mask for the previous normal node. This is the index of the current node's position in the semantically faithful break propagation chain. This is the position index of the previous normal node. This serves as the position index for the next normal node. This is the stage mask for the next normal node;

[0123] The new semantic vector is combined with the new stage to form a reconstruction result. Then, according to the node position index, the reconstruction result is written back to the original semantic fidelity break propagation chain structure node by node to complete the semantic write-back operation.

[0124] By introducing a probability correction mechanism based on the coupling of modal suppression index and node weight in the real conflict chain, the node probability is no longer determined solely by the original distribution, but can reflect its degree of suppression and structural importance in the conflict structure. This allows for targeted reinforcement of key nodes during the information calculation process, preventing them from being diluted or obscured during continuous propagation.

[0125] Based on this, by performing differential analysis on the amount of information and combining it with sliding window identification of editable fields, the scope of correction is transformed from overall rewriting to local precise positioning. This effectively solves the problem of excessive modification caused by the inability to distinguish key conflict areas in traditional methods, while avoiding the defect that semantic expression is simplified or even lost during the repair process.

[0126] Furthermore, node index mapping is completed by combining position coordinates, and semantic vectors and stage interpolation are performed based on the front and rear boundary nodes, thereby suppressing the semantic weakening problem caused by information impoverishment and repairing the structural conflict diffusion problem caused by modality suppression.

[0127] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A multimodal graph-text semantic conflict detection and correction method based on contrastive learning, characterized in that, The method includes: Multimodal image and text data is acquired, and data parsing is performed based on contrastive learning-driven natural language processing technology to obtain a semantically faithful break propagation chain. Reconstruction of semantically fidelity-preserving broken propagation chains is performed through modal suppression coupling recognition; Based on the pseudo-semantic consistency segmentation and chain strengthening of the reconstructed semantic fidelity broken propagation chain, the offset propagation sub-chain is obtained. Feature extraction is performed on the offset propagation sub-chain to obtain the time offset chain segment, and event consistency judgment and chain structure update are performed on the semantic nodes to obtain the mapping classification results, including asynchronous chain segments and real conflict chain segments. If it is an asynchronous link segment, the information content of the semantic node is extracted and corrected, and after differential processing, the editable field is identified; Perform semantic write-back operations based on editable fields.

2. The multimodal graph-text semantic conflict detection and correction method based on contrastive learning according to claim 1, characterized in that, Multimodal image and text data is acquired, and data parsing is performed based on contrastive learning-driven natural language processing techniques to construct a semantically faithful break propagation chain, including: Based on natural language processing technology, the text description in multimodal image and text data is semantically decomposed to obtain text nodes. The non-text description is structurally parsed through structured semantic modeling and indexing encoding to obtain structure nodes. The text nodes and structure nodes are mapped to the same semantic vector space to obtain a joint set of semantic primitives. Based on a joint set of semantic primitives, semantic node pairs are identified through comparative learning of semantic matching. Path expansion analysis is then performed using a connected component algorithm to construct a semantic association domain and generate joint location coordinates. Based on the semantic association domain, the energy anomaly domain is identified and a set of fracture candidate regions is generated through semantic energy distribution modeling and gradient change analysis. Based on the set of fracture candidate regions, a semantically faithful fracture propagation chain is constructed through path tracing and structural connection to characterize the contribution distribution, path transmission relationship and conflict signal attenuation state of different modal semantic nodes during the comparison and fusion process.

3. The multimodal graph-text semantic conflict detection and correction method based on contrastive learning according to claim 2, characterized in that, Based on semantic association domains, semantic energy distribution modeling and gradient change analysis are used to identify energy anomaly domains and generate a set of fracture candidate regions, including: Combining text nodes and structure nodes yields semantic nodes; Extract the vector response intensity of semantic nodes and use it as the semantic energy value; Based on the semantic association domain, the energy change between adjacent semantic nodes is calculated to obtain the gradient sequence. Each feature in the gradient sequence is compared with a preset gradient threshold to obtain the break nodes and generate the energy anomaly domain. The continuity of the energy anomaly domain is verified by combining the location coordinates to generate a set of fracture candidate regions.

4. The multimodal graph-text semantic conflict detection and correction method based on contrastive learning according to claim 3, characterized in that, Reconstruction of the semantically faithful broken propagation chain is performed through modal suppression coupling recognition, including: Based on the semantic fidelity break propagation chain, a chain-level semantic contribution distribution matrix is ​​constructed through intra-chain semantic energy statistics and distribution analysis. Based on the chain-level semantic contribution distribution matrix, the modal suppression index is obtained by calculating the degree of contribution deviation; combined with path constraint analysis, the suppression propagation interval with the same direction of contribution deviation is identified, and the suppression propagation sub-chain is obtained. Based on the suppression propagation subchain and modal suppression index, a reconstructed semantically faithful break propagation chain is obtained through nonlinear semantic stretching and path structure rearrangement.

5. The multimodal graph-text semantic conflict detection and correction method based on contrastive learning according to claim 4, characterized in that, Based on the pseudo-semantic consistency segmentation and chain strengthening of the reconstructed semantically faithful break propagation chain, the offset propagation sub-chain is obtained, including: By analyzing the semantic differences between adjacent semantic nodes in the reconstructed semantic fidelity fracture propagation chain, a sequence of differences between nodes is obtained, and the joint position coordinates are used as a continuity constraint to obtain a set of difference segments. Based on the differential segment set, a pseudo-semantic consistent segment set is obtained through comparative analysis of local consistency and global consistency, including multiple sets of pseudo-semantic consistent segments; Extract the node sub-chain corresponding to each pseudo-semantic consistent segment in the pseudo-semantic consistent segment set, and use it as the initial propagation path; Using energy continuity and path reachability as search conditions, a neighborhood search is performed on the initial propagation path, and after path expansion, the propagation sub-chains are obtained. Based on the propagation subchain, the relative positional relationship of each semantic node in the corresponding pseudo-semantic consistency segment is identified, and the semantic node weight and offset propagation subchain are obtained.

6. The multimodal graph-text semantic conflict detection and correction method based on contrastive learning according to claim 5, characterized in that, Feature extraction is performed on the offset propagation subchain to obtain the time offset chain segment. Event consistency judgment and chain structure update are then performed on the semantic nodes to obtain the mapping classification results, including: Extract the temporal semantic units of each semantic node in the offset propagation subchain, and obtain the temporal semantic sequence after node mapping; Extract stage descriptors, and perform stage segmentation and node labeling on the temporal semantic sequence based on the stage descriptors to construct a stage matrix; Based on the stage matrix, time difference calculation and path mapping analysis are performed on each semantic node to obtain the time offset chain segment; Based on the time offset chain segment, the semantic nodes are subjected to event consistency judgment and chain structure update to obtain the mapping classification result.

7. The multimodal graph-text semantic conflict detection and correction method based on contrastive learning according to claim 6, characterized in that, Based on the time-offset chain segment, event consistency judgment and chain structure update are performed on the semantic nodes to obtain the mapping classification results, including: Set classification constraints and use these constraints to make a comprehensive judgment on the time offset chain segments; The classification constraints include path continuity determination and stage sequence consistency determination; If a semantic node is discontinuous in the reconstructed semantic fidelity break propagation chain path structure or the stage order does not conform to the single event advancement logic, it is marked as a real conflict chain segment and a fidelity constraint instruction is triggered; if a semantic node is continuous in the reconstructed semantic fidelity break propagation chain path structure and the stage order conforms to the single event advancement logic, it is marked as an asynchronous chain segment and correction processing is performed. The asynchronous chain segment and the real conflict chain segment are combined to obtain the classification result, and the classification result is mapped back into the reconstructed semantic fidelity break propagation chain to obtain the mapped classification result.

8. The multimodal graph-text semantic conflict detection and correction method based on contrastive learning according to claim 7, characterized in that, If it is an asynchronous link segment, the information content of the semantic node is extracted and corrected, and after differential processing, the editable field is identified, including: Receive the fidelity constraint instruction, extract the node probability of each semantic node in the real conflict chain segment, and use the product of the modality suppression index and the semantic node weight as a correction term to correct the node probability. Combine information theory to calculate the information content of each semantic node. The information content of adjacent semantic nodes is differentially divided to obtain a differential sequence containing multiple sets of information density changes. The information density changes are compared with a preset change threshold. If the information density changes exceed the preset change threshold, it is recorded as an editable field.

9. The multimodal graph-text semantic conflict detection and correction method based on contrastive learning according to claim 8, characterized in that, Perform semantic write-back operations, including: Based on the editable field and the joint position coordinates, the node position index and coordinate mapping relationship are obtained; Based on the node position index, the front and back boundary nodes of the editable field are located, and the semantic vectors and stages of the boundary nodes are extracted respectively, which are used as front and back anchor points to generate new semantic vectors. Based on the new semantic vector, the stages of semantic nodes in the editable domain are corrected to obtain new stages, making them continuous and smooth in time sequence and correcting time offsets; The new semantic vector is combined with the new stage to form a reconstruction result. Then, according to the node position index, the reconstruction result is written back to the original semantic fidelity break propagation chain structure node by node to complete the semantic write-back operation.