Network false information detection method and system based on large model fine tuning

By associating semantic fragments and tracing chains of misplaced and false information texts in Chinese online corpora, the logical order is reconstructed, solving the problem of inconsistent semantic judgment in existing technologies and achieving stable identification of false information in multi-source corpus scenarios.

CN121503499BActive Publication Date: 2026-05-29PEOPLES POLICE UNIV OF CHINA (INT LAW ENFORCEMENT COOP INST OF THE MINISTRY OF PUBLIC SECURITY CHINA PEACEKEEPING POLICE TRAINING CENT)

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PEOPLES POLICE UNIV OF CHINA (INT LAW ENFORCEMENT COOP INST OF THE MINISTRY OF PUBLIC SECURITY CHINA PEACEKEEPING POLICE TRAINING CENT)
Filing Date
2025-11-24
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies lack global tracking of semantic connection order and contextual logic when processing cross-sentence or multi-level quoted text, resulting in decreased accuracy of semantic judgment. In particular, under the influence of sentence breaks or emotional bias, the model has difficulty identifying online misinformation.

Method used

By acquiring misplaced and false information texts from Chinese online corpora, we extract source tag fields, skipped sentences, and contextual information from text paragraphs, determine the semantic span order, separate factual statements, emotional statements, and information source sentences from paragraphs, trace the semantic extension path along the sentence order, identify inconsistent parts of sentence combinations, adjust paragraph positions, reconstruct semantic chain groups, and identify truth and falsehood judgments and cited content.

Benefits of technology

Maintaining the continuity and stability of semantic judgment in multi-source corpus scenarios improves the accuracy and reliability of false information identification, eliminates semantic shifts caused by paragraph jumps, and ensures causal connection and contextual coherence of content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121503499B_ABST
    Figure CN121503499B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of cloud collaboration, in particular to a network false information detection method and system based on large model fine-tuning, comprising the following steps: extracting mispositioned text and context information, judging semantic span to form an association set, separating factual sentences, emotional sentences and provenance sentences to track semantic chains, comparing original texts to identify offset segments, adjusting the order of analysis to analyze cause and effect, and obtaining false information determination conclusions. In the present application, by correlating and extracting the mispositioned relationship between semantic segments and tracking the chain, the logical order can be reconstructed at the sentence level, the semantic deviation caused by paragraph jumps can be eliminated, the content cause and effect connection and the context coherence can be maintained, the corresponding relationship between factual sentences, emotional tendencies and time elements can be identified simultaneously in semantic path analysis, the semantic judgment remains continuous and stable, the dynamic comparison and offset correction of the quoted content and the original context are completed in the multi-source corpus scene, and the stability and conclusion reliability of false information identification are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of cloud-based collaborative technology, and in particular to a method and system for detecting fake information on the network based on large model fine-tuning. Background Technology

[0002] The field of cloud-based collaborative technology encompasses intelligent collaborative methods for unified data processing and model scheduling across multiple terminals and scenarios, relying on cloud computing resources. Core aspects include model parameter distribution, multi-source synchronous data training task delivery, and inference process result integration. It focuses on solving the problem of sharing model computing resources and promoting the dynamic adaptation and collaborative execution of AI models across different computing nodes. This technology is widely applied in practical scenarios involving complex semantic judgments, such as public opinion analysis, content review, and information security. Traditional methods for detecting online misinformation based on large-scale model fine-tuning utilize pre-trained large-scale language models. By constructing a Chinese rumor dataset corresponding to the task scenario, efficient parameter fine-tuning strategies, such as low-rank adaptation techniques, are employed in a local environment to retrain the model and improve its ability to identify specific information. This approach typically enhances the structured expression of the model output by constructing a sample labeling system and introducing time-stamped information, thereby enabling the detection and discrimination of complex semantic phenomena such as forged authority and misinterpretation in online texts.

[0003] Existing technologies rely on fixed corpora and labeling systems for semantic analysis. When processing cross-sentence or multi-level quoted texts, they lack global tracking of semantic connection order and contextual logic. They are easily affected by sentence breaks or emotional biases, which can affect the accuracy of judgment. When the quoted sentences are misplaced or the time sequence is disordered, the model tends to rely on surface features and ignore deep semantic relationships, resulting in a larger deviation between factual expression and quoted information. In scenarios with multi-source corpus fusion or large semantic span, the output results often have problems of semantic incoherence and ambiguous judgment. Summary of the Invention

[0004] To address the technical problems existing in the prior art, embodiments of the present invention provide a method for detecting network misinformation based on large model fine-tuning, comprising the following steps:

[0005] To achieve the above objectives, the present invention adopts the following technical solution: a network misinformation detection method based on large model fine-tuning, comprising the following steps:

[0006] S1: Obtain misquoted and false information from Chinese online corpus, extract source tag fields, jump sentences and context information of text paragraphs, determine the semantic span order, and obtain a semantic fragment association set;

[0007] S2: Based on the semantic fragment association set, separate the sentences that state facts, the sentences with emotional tendencies, and the sentences that explain the source information in the paragraph. Determine the connection between the preceding and following sentences according to the position of each sentence in the paragraph, and trace the semantic extension path along the sentence order to obtain the semantic judgment chain structure.

[0008] S3: Based on the semantic determination chain structure, extract the corresponding statement of the quoted content and compare its position with the quoted segment in the original text, identify the inconsistent statement combination, and obtain the semantic offset fragment set;

[0009] S4: Based on the set of semantic offset segments, identify the location of sequential jumps, analyze the starting words and connecting words of the sentences on both sides of the breakpoint, and adjust the position of the segments according to the causal relationship and reverse expression between the sentences to obtain a semantic reconstruction chain group.

[0010] S5: Based on the semantic reconstruction chain group, identify the true / false judgment, reference information and time description, analyze whether each statement includes the corresponding related information, retain the complete information statement, and obtain the false information judgment conclusion group.

[0011] As a further aspect of the present invention, the semantic fragment association set includes source tag fields, jump sentences, context information, sequential logic, semantic span order, and misdirection relationships; the semantic judgment chain structure includes factual sentences, sentences with emotional bias, source information sentences, connection between preceding and following sentences, semantic extension paths, breakpoints, word order jumps, and consistency of citation direction; the semantic offset fragment set includes cited content sentences, inconsistent sentence combinations, citation start and end positions, context connection methods, and misaligned fragments; the semantic reconstruction chain group includes sequential jump positions, sentence start words, connecting words, causal relationships, reverse expressions, time descriptions, background clues, and missing sentence positions; and the false information judgment conclusion group includes truth / false judgment, cited content, time information, sentences containing complete information, and sentences lacking information.

[0012] As a further aspect of the present invention, the skipping statement refers to a statement in the text that breaks away from the original logical order of the context and directly points to the content across intermediate segments;

[0013] The semantic span refers to the semantic interval between sentences, the logical and temporal distance that information traverses from one segment to another.

[0014] As a further aspect of the present invention, the sentence order direction statement refers to the logical direction of the arrangement order in the paragraph and the semantic progression, and determines whether the information flow is forward or backward;

[0015] The semantic extension path refers to the logical chain in which semantics are sequentially connected and propagated between sentences and segments, reflecting the semantic connection between facts, emotions, and cited information.

[0016] As a further aspect of the present invention, the specific steps of S1 are as follows:

[0017] S101: Obtain misquoted false information text from Chinese online corpus, extract source tag fields, skipped sentences and context identifiers from text paragraphs, determine the positional order of source fields and skipped sentences, and obtain the citation order judgment result;

[0018] S102: Based on the citation order judgment result, compare the semantic span order between the jump segment statement and the context identifier, combine the time sequence information and the citation direction field, identify the part where the position is reversed and the semantic connection direction is inconsistent, and obtain the segment citation offset information;

[0019] S103: Based on the segment reference offset information, filter the positions of fields with misaligned order and direction, compare them with the context connection interval, extract the logical offset part between segments, and obtain the semantic segment association set.

[0020] As a further aspect of the present invention, the specific steps of S2 are as follows:

[0021] S201: Based on the semantic fragment association set, separate factual statements, emotional statements and information source segments in the segment, identify the sentence punctuation structure and sentence ending punctuation mark type of the sentence, and build a segment index list according to type to obtain a sentence classification structure set;

[0022] S202: Based on the sentence classification structure set, extract the start and end positions of each type of sentence segment in the original paragraph, perform a linear comparison of the order position difference between adjacent sentence segments, mark the order jump points, remove disordered segments, and obtain the sentence sequence connection diagram.

[0023] S203: Based on the segment sequence connection diagram, locate the jump breakpoint, analyze the connecting words and transition expressions of the adjacent segments, determine the sequential relationship and consistency of the citation direction between adjacent segments, filter semantic conflict intervals and record the connection path to obtain the semantic judgment chain structure.

[0024] As a further aspect of the present invention, the specific steps of S3 are as follows:

[0025] S301: Based on the semantic determination chain structure, extract the content fields identified as references in each sentence, call the time description field and topic guidance statement in the original text content, retrieve the sentence position of the matching field in the original text, record the sentence index interval, and obtain the set of reference positions;

[0026] S302: Based on the reference position corresponding set, extract each group of referenced segments and the connecting words between the preceding and following sentences, construct the guide word sequence and the response sentence sequence, calculate the distribution frequency offset ratio of the guide word sequence in the sentence group, determine the arrangement pattern of the referenced fields in the text, and obtain the sentence segment order offset set;

[0027] S303: Based on the sentence segment order offset set, locate the start and end positions of the offset sentence group, and divide the reference range and corresponding reference method of each sentence segment according to the first appearance order of the referenced field in the text and the context paraphrase method to obtain the semantic offset fragment set.

[0028] As a further aspect of the present invention, the specific steps of S4 are as follows:

[0029] S401: Based on the semantic offset fragment set, retrieve the arrangement order of the segments in the context, track the segment group to which the jump points, identify the index position in the preceding and following segment structures, compare and map the sequence distribution state to obtain the jump position index set;

[0030] S402: Based on the jump position index set, locate the first word and connecting word of the two segments on both sides of the jump area, and determine the logical connection direction between the preceding and following segments according to the connection order and part-of-speech matching relationship of the words on both sides in the semantic chain, and obtain the reverse connection position group;

[0031] S403: Based on the reverse connection position group, extract the time identifier field and background description statement in the corresponding context, filter the relevant paragraph segments, match the missing part position according to the context coverage and semantic coherence path, and obtain the semantic reconstruction chain group.

[0032] As a further aspect of the present invention, the specific steps of S5 are as follows:

[0033] S501: Based on the semantic reconstruction chain group, retrieve the true / false judgment statements, quoted content statements and time description statements in the segment, locate the starting position and grammatical relationship of the statements in the paragraph, identify the connection path and semantic inheritance structure between statements, and obtain the structure set corresponding to the statements.

[0034] S502: Based on the structure set corresponding to the statement, identify the semantic extension relationship between the statements, identify the order of the true / false judgment content and the quoted statement according to the logical sequence of the paragraph, and aggregate the statement combination with the characteristics of related order and semantic coherence to obtain a logically consistent combination set.

[0035] S503: Based on the logically consistent combination set, determine whether the segment simultaneously contains a true / false judgment statement, a quoted content statement, and a time description statement; remove content that does not have a complete statement; extract semantically coherent and structurally complete sentence segments to obtain a false information judgment conclusion group.

[0036] A network misinformation detection system based on large model fine-tuning includes:

[0037] The data processing module acquires false information text from Chinese online corpora, extracts source tags, jump sentences and context information, determines the order logic of source fields and jump sentences, determines the semantic span order, and obtains a semantic fragment association set.

[0038] Based on the semantic fragment association set, the semantic judgment module separates sentences stating facts, statements with emotional tendencies, and sentences that provide source information from the segment, judges the connection between the preceding and following sentences, and traces the semantic extension path along the sentence order to obtain the semantic judgment chain structure.

[0039] Based on the semantic determination chain structure, the offset recognition module extracts the corresponding sentences of the quoted content and compares their positions with the quoted segments in the original text, identifies inconsistent sentence combinations, and obtains a set of semantic offset fragments.

[0040] Based on the semantic offset fragment set, the reconstruction processing module identifies the sequential jump position, analyzes the starting words and connecting words of the sentences on both sides of the breakpoint, adjusts the position of the segment before and after, and adjusts the position of the segment before and after according to the causal relationship and reverse expression method between the sentences before and after to obtain the semantic reconstruction chain group.

[0041] The judgment output module extracts the true / false judgment, reference content and time information based on the semantic reconstruction chain group, analyzes whether each statement includes complete information, retains the statements that include all information, and obtains the false information judgment conclusion group.

[0042] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0043] In this invention, by extracting the associations and chain-tracking the misaligned relationships between semantic fragments, the logical order can be reconstructed at the sentence level, eliminating semantic shifts caused by paragraph jumps, maintaining causal connection and contextual coherence, and simultaneously identifying the correspondence between factual statements, emotional tendencies and time elements in semantic path analysis, so that semantic judgment remains continuous and stable. In multi-source corpus scenarios, dynamic comparison and offset correction of quoted content and original context are completed, improving the stability and reliability of false information identification. Attached Figure Description

[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, the accompanying drawings can be obtained from these drawings without creative effort.

[0045] Figure 1 This is a schematic diagram of the steps of the present invention;

[0046] Figure 2 This is a detailed schematic diagram of S1 of the present invention;

[0047] Figure 3 This is a detailed schematic diagram of S2 of the present invention;

[0048] Figure 4 This is a detailed schematic diagram of S3 of the present invention;

[0049] Figure 5 This is a detailed schematic diagram of S4 of the present invention;

[0050] Figure 6 This is a detailed schematic diagram of S5 of the present invention;

[0051] Figure 7 This is a system module diagram of the present invention. Detailed Implementation

[0052] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0053] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0054] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.

[0055] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0056] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0057] Please see Figure 1 This invention provides a method for detecting network misinformation based on large model fine-tuning, comprising the following steps:

[0058] S1: Obtain misquoted false information text from Chinese online corpus, extract source tag fields, skipped sentences and context information from text paragraphs, determine the logical order of source fields and skipped sentences, compare the semantic span order according to the context position, identify the misdirection relationship between paragraphs through the difference in time and citation direction, and obtain a semantic fragment association set;

[0059] S2: Based on the semantic fragment association set, separate the sentences that state facts from the sentences with emotional tendencies and the information segments used to explain the source in the paragraph. Determine the connection between the preceding and following sentences according to the position of each sentence in the paragraph, trace the semantic extension path along the sentence order, rearrange the break points, word order jumps and parts with inconsistent citation directions, and obtain the semantic judgment chain structure.

[0060] S3: Based on the semantic judgment chain structure, extract the sentences corresponding to the quoted content and compare their positions with the quoted segments in the original text. Identify inconsistent sentence combinations by sentence order and semantic connection order. Segment the misaligned segments according to the start and end positions of the citation and the context connection method to obtain a set of semantic offset segments.

[0061] S4: Based on the semantic offset fragment set, identify the location of the sequential jump, analyze the starting words and connecting words of the sentences on both sides of the breakpoint, adjust the position of the sentence segments according to the causal relationship and reverse expression between the sentences, and filter the position of the missing sentence segments in the original context by combining time description and background clues to obtain the semantic reconstruction chain group.

[0062] S5: Based on the semantic reconstruction chain group, extract the true / false judgment, reference content and time information, combine the extracted content in order, analyze whether each statement includes the extracted content, retain the complete corresponding related information, output in order, exclude statements with missing information, and obtain the false information judgment conclusion group;

[0063] The semantic fragment association set includes source tag fields, jump sentences, contextual information, sequential logic, semantic span order, and misdirection relationships. The semantic judgment chain structure includes factual sentences, sentences with emotional bias, source information sentences, connection between preceding and following sentences, semantic extension paths, breakpoints, word order jumps, and consistency of citation direction. The semantic offset fragment set includes quoted content sentences, inconsistent sentence combinations, citation start and end positions, context connection methods, and misplaced fragments. The semantic reconstruction chain group includes sequential jump positions, sentence start words, connecting words, causal relationships, reverse expressions, time descriptions, background clues, and missing sentence positions. The false information judgment conclusion group includes truth / false judgments, quoted content, time information, sentences containing complete information, and sentences lacking information.

[0064] As a further aspect of the present invention, the skipping statement refers to a statement in the text that breaks away from the original logical order of the context and directly points to the content across intermediate segments;

[0065] The semantic span refers to the semantic interval between sentences, the logical and temporal distance that information traverses from one segment to another.

[0066] As a further aspect of the present invention, the sentence order direction statement refers to the logical direction of the arrangement order in the paragraph and the semantic progression, and determines whether the information flow is forward or backward;

[0067] The semantic extension path refers to the logical chain in which semantics are sequentially connected and propagated between sentences and segments, reflecting the semantic connection between facts, emotions, and cited information.

[0068] As a further aspect of the present invention, the specific steps of S1 are as follows:

[0069] S101: Obtain misquoted false information text from Chinese online corpus, extract source tag fields, skipped sentences and context identifiers from text paragraphs, determine the positional order of source fields and skipped sentences, and obtain the citation order judgment result;

[0070] First, monitor the text corpus containing source tag fields and identify whether there are skipped sentences in each text segment. These sentences are usually represented by conjunctions or special identifiers, such as "responding to this" or "the reporter previously mentioned". The context information before and after these sentences can be marked, and the search scope can be set, such as from the beginning of the text segment to the location of the skipped sentence, or from the skipped sentence to the end of the text segment. During the extraction process, citation identifiers in the text segment, such as "according to xx" or "pointed out", need to be located, their position index values ​​extracted, and marked as source field position indexes. Simultaneously, a skipped sentence identifier value is assigned to the location of the skipped sentence. By comparing the size relationship between the source field index and the skipped sentence identifier value, their order is determined. If the source field position is after the skipped sentence, it is recorded as a misalignment. Further, the contextual identifier content of the text segment is obtained, such as the ending of the previous paragraph and the first sentence of the current paragraph, to determine whether they belong to the same subject or whether there is a referential relationship. For example, if the references to "Zhang Mou" and "", or "event" and "this case" are logically inconsistent, they are marked as having a disjointed context. Then, the positions of the source fields, the jump paragraphs, and the contextual referential states need to be integrated to categorize the relationships between these three. For example, the source precedes the jump paragraph and the references are consistent; the source precedes the jump paragraph and the references conflict. Each state constitutes a combination pattern. Based on this, pattern matching is performed on all text to extract combinations with abnormal citation positions and logical inconsistencies, and these are recorded as groups of texts with abnormal citation relationships. In online comment texts, analysis can be combined with user paraphrasing behavior on social media platforms. For example, if a comment paragraph contains "according to xx," followed immediately by "actually, that's not what they said," then there is a clear logical reversal in the citation position within that paragraph. Basic judgment logic can be established through such cases. Finally, the output results of the above comparison and judgment process are uniformly recorded to obtain the citation order judgment result.

[0071] S102: Based on the citation order judgment result, compare the semantic span order between the jump segment statement and the context identifier, combine the time sequence information and the citation direction field, identify the part where the position is reversed and the semantic connection direction is inconsistent, and obtain the segment citation offset information;

[0072] First, obtain the starting position of the jump statement and the position of the source tag field in the segment. By identifying the semantic span between the two, that is, the text path from the target reference content pointed to by the jump statement to the actual source position, extract the content of the context identifier fields involved in the path, and align the semantic guidance of the jump statement with the information flow carried by the target reference content in the context to determine whether there is a semantic jump interruption. In the process, it is necessary to obtain the subject-predicate structure and logical conjunction type of all statements on the path, such as "but", "and", "therefore", etc., to identify the semantic progression direction and determine whether the semantic connection of the jump statement is valid. For example, in "According to Xinhua News Agency, Li Moumou was investigated. This is met with anger", if the jump statement "This is met with anger" points to the event of another news subject "Wang Moumou", then there is a semantic connection break in the context. At this time, it is necessary to record the position index of the jump statement as 'a' and the corresponding source field position index as 'b'. If position a comes before position b, This is a case of reversed citation order. Further, a citation direction field is introduced. Based on directional verbs or phrases such as "according to," "points out," etc., the consistency of information flow can be determined. The current statement is compared sequentially with time-marked statements in the previous paragraph. Descriptions like "the day before the event," "today," and "previously" can be used to determine the correctness of the citation's temporal order. If the jump statement is later than the statement containing the cited object and has no logical connection, it is recorded as a discrepancy between time and citation direction. For example, in a forum comment, if "Wang was mentioned as a fraudster in the previous text, but it was actually Li who did it," a sudden change in the semantic subject between "Wang was mentioned as a fraudster" and "it was actually Li who did it" can be identified, resulting in a distorted connection direction. In such cases, both the time sequence and the citation subject direction field must be combined to determine if a citation offset has occurred. If both are inconsistent, the statement is classified as an offset segment, and a segment citation relationship chain is constructed, recording the location and corresponding context connection status, ultimately obtaining the segment citation offset information.

[0073] S103: Based on the segment reference offset information, filter the positions of fields with misaligned order and direction, compare them with the context connection interval, extract the logical offset part between segments, and obtain the semantic segment association set;

[0074] First, identify the positions of paragraphs or fields with semantic jump breaks or abnormal information flow. These misaligned fields need to be located and processed in conjunction with their connection to preceding and following paragraphs in the context. First, filter out statements where the field's position points differently before and after the jump. A misaligned field might appear in paragraph A pointing to paragraph C, but the intermediate paragraph B fails to connect semantically or is skipped. In this case, it's necessary to check if there are abrupt changes in subject or timeline in the adjacent statements of the paragraph containing the field. For example, in a text, there might be "Zhang was arrested in the case," followed by "According to Wang, he was never present." If the content describing the connection between Zhang and Wang is missing, then Wang's statement is the misaligned field. Another example is common in online forums, where "Li was accused of misappropriation of property last year and responded in a post on WeChat Moments today" lacks a connection between the preceding statement "The accusation originated from the narrator Zhao," causing confusion. When there is a logical break between the use and the response, such fields need to be recorded separately and their original position information needs to be labeled. Then, the field is combined with the preceding and following paragraphs to form a connection interval. The content in the connection interval is compared semantically to identify whether the paragraphs jointly describe the same subject or time action. If the subject or verb action of the paragraph containing the field cannot be found in the preceding semantic content of the connection interval, it is judged as a segment with logical deviation. It is necessary to further extract the keyword content with segment jumping characteristics, such as keywords with directional or time-connected characteristics, such as "forward", "target", "respond", "explain", etc., and aggregate and compare them in multiple paragraphs. The judgment condition is established based on whether the subject quoted by the segment jumping field matches the subject of the surrounding paragraph. Then, combined with the relative position order of the field and the connection interval and the continuity of its connection state, the content of the segment with semantic alignment interruption phenomenon is identified, and finally the semantic segment association set is obtained.

[0075] Please see Figure 3 The specific steps of S2 are as follows:

[0076] S201: Based on the semantic fragment association set, separate factual statements, emotional statements and information source segments in the text segment, identify the sentence punctuation structure and sentence ending punctuation mark type of the sentence, and build a sentence segment index list according to type to obtain the sentence classification structure set;

[0077] First, factual statements, emotional statements, and information source segments are extracted from the text. To achieve semantic classification, each statement needs to undergo punctuation recognition and sentence segmentation. The number of commas, semicolons, periods, etc., and their positions in the semantic distribution are determined. Sentence type is determined by identifying the text intervals between punctuation marks. For example, sentences with a verb predicate and subject relationship are classified as factual statements; sentences containing interjections, emotional adverbs, or modal particles are classified as emotional statements; and sentences containing citations or phrases such as "according to a certain news" or "it has been reported" are classified as information source segments. In practice, the system scans each statement character by character when reading the text, marking punctuation marks as segmentation nodes, and segmenting based on the semantic density between nodes. For example, "The survey shows that some users question this message" can be split into two semantic units: "The survey shows" is the source identifier, and "some users question this message" is the factual statement. Semantic key comparisons are then used to further define the semantic units. The functional attributes of words determine the sentence type, and the sentence classification is determined by the sentence ending punctuation mark. If the sentence ends with a period and there is no rhetorical change, it is classified as a normal declarative sentence; if it ends with a question mark, it is classified as an interrogative sentence; and if it ends with an exclamation mark, it is marked as a sentence with emotional tendency. To improve the accuracy of semantic recognition, a threshold for the number of punctuation marks in a sentence can be set. When a single sentence has more than three punctuation marks and contains multiple conjunctions, the sentence is classified as a compound sentence. Its clauses are classified according to the above rules. For example, in "Some people questioned, but the official response said the information was true", "Some people questioned" is classified as emotional tendency, and "The official response said the information was true" is classified as factual statement. After all sentences have completed sentence structure recognition, an index list is built based on sentence structure type and sentence source label. The index list records the sequence number, sentence type, and semantic feature code of the sentence in the text, ensuring that each sentence corresponds to a unique index position. The index list can record the entry number, type identifier, and source mark field. Through multi-text cross-comparison, a unified semantic level labeling is achieved, and finally, a sentence classification structure set is obtained.

[0078] S202: Based on the sentence classification structure set, extract the start and end positions of each type of sentence segment in the original paragraph, perform linear comparison of the order position difference between adjacent sentence segments, mark the order jump points, remove disordered segments, and obtain the sentence sequence connection diagram.

[0079] First, the index records in the classification structure set are read to identify the order of declarative sentences, emotional sentences, and source sentences in the text. The continuity of the arrangement is determined by comparing the index numbers of adjacent sentence segments. If there is a jump between the numbers, the interval is recorded as an abnormal sequence interval. During implementation, the starting and ending character positions of each sentence segment need to be mapped. For example, in a text like "According to a survey, users are generally worried about the veracity of the message, while some comments believe the matter has been exaggerated," "According to a survey" is marked with a starting position of 1 and an ending position of 8, while "users are generally worried about the veracity of the message" has a starting position of 9 and an ending position of 28. If the interval exceeds the set benchmark, it indicates a semantic gap between sentences. By comparing the size of such intervals, the continuity between sentence segments can be determined. In linear comparison, the difference between the starting positions of adjacent sentence segments is compared with the average sentence length interval. If the difference between the positions of sentences exceeds twice the average sentence length, it is considered a segment jump phenomenon and recorded as a sequence jump point in the annotation table. For example, in forum comment data, if a paragraph contains "Li was investigated" followed directly by "as seen in the on-site video," but lacks the connecting statement "the relevant parties made an explanation," then this constitutes a paragraph jump. After comparing all sentences and paragraphs, it is necessary to remove duplicate paragraphs or incorrectly indexed disordered segments in the text. These segments are usually caused by punctuation errors or semantic leaps, such as overlapping indexes when multiple exclamatory sentences are written consecutively. In data processing, sentences and paragraphs with the same starting position or overlapping intervals are identified and removed to ensure that the arrangement of sentences and paragraphs in the context conforms to the natural word order. Finally, the retained sentences and paragraphs and their adjacent relationships are mapped by connecting the end position of one sentence to the beginning position of the next sentence, drawing the path relationship of sentence and paragraph connection, so that the text forms a continuous mapping sequence in semantic structure, and finally obtains the paragraph sequence connection diagram.

[0080] S203: Based on the segment sequence connection diagram, locate the jump breakpoint, analyze the connecting words and transition expressions of adjacent segments, determine the sequential relationship and consistency of the citation direction between adjacent segments, filter semantic conflict intervals and record the connection path to obtain the semantic judgment chain structure.

[0081] First, interrupted or reverse-pointing connectors in the paragraph connection path are marked, recording the end character position of the previous paragraph and the beginning character position of the next paragraph. Discontinuous connection points are identified. For example, in a news commentary, if "Netizens questioned the source of the information, but the official subsequently issued a statement" is immediately followed by "The statement was refuted by multiple parties," the "but" in the middle constitutes a contrastive connection, but the following sentence does not continue the original subject, forming a break. Next, the connecting words of adjacent paragraphs need to be detected, extracting words such as "but," "however," "although," and "therefore" that indicate contrast, progression, and cause and effect, and determining their logical relationship between paragraphs. If the connecting words in adjacent paragraphs are of opposite types, for example, the first sentence is a contrastive connection while the second sentence is a progressive structure, then the logical direction is determined to be inconsistent. During the processing, the changes in subject-verb relationship in adjacent sentences also need to be identified, and the coherence between verb actions and subject entities needs to be checked. For example, if "Zhang admitted the accusation" and "The relevant party denied the rumors" have the subject changing from an individual to an institution in the same narrative chain, then this position is marked as a semantic conflict point. In large-scale text applications, a table mapping conjunctions to sentence types can be established. By comparing the connection patterns of different sentence segments, the type of breakpoint can be determined. For example, in the sentence "According to reports, the case has been closed, but the comments section is still discussing it," the word "but" indicates a semantic reversal, and the following sentence introduces a new subject, "comments section." This shift constitutes a breakpoint. Subsequently, based on the breakpoint location, the sentence type and citation direction fields of the segments on both sides are extracted. If the preceding and following segments cite the same source tag but have opposite narrative directions, they are marked as directional conflict intervals in the judgment table. To filter semantic conflict intervals, a consistency search needs to be performed on the set of segments associated with all breakpoints, deleting non-conflicting statements that are repeatedly pointed to in the connection path, retaining only those with directional reversals, subject misplacement, or narrative deviation, and recording the connection path information for each conflict interval. Finally, these paths are arranged in semantic extension order, ensuring that each conflict interval maintains a continuous and traceable relationship with its context, ultimately resulting in a semantic judgment chain structure.

[0082] Please see Figure 4 The specific steps of S3 are as follows:

[0083] S301: Based on the semantic decision chain structure, extract the content fields marked as references in each sentence, call the time description field and topic guidance statement in the original text content, retrieve the sentence position of the matching field in the original text, record the sentence index interval, and obtain the set corresponding to the reference position.

[0084] First, the segments recorded in the judgment chain are traversed sentence by sentence, and field extraction rules are set to identify sentences containing verbal quotations or source noun phrases. Next, the time description field contained in the original text is called, which can be a descriptive time phrase or time adverb, such as "at that time," "then," or "a little earlier." This field is located by recording its start and end positions in the original paragraph using sentence position markers. Simultaneously, topic-leading statements are extracted and called. These statements generally appear at the beginning of a paragraph and explicitly indicate the current topic through demonstrative pronouns or subjects, such as "This report pointed out" or "The official statement." These statements are used as topic anchors for semantic connection matching. When processing quotation field matching, the extracted quotations are compared with the keyword groups in each sentence of the original text for keyword overlap. If a segment contains multiple combinations of quotations, it is marked as a high-match statement, and its position in the entire text is recorded. Sentence / segment positions are represented by paragraph numbers and relative order within the sentence, forming directly identifiable index intervals. For example, if "according to Xinhua News Agency" appears in the second sentence of paragraph 4, the index interval is recorded as "4-2". By combining the positions of multiple cited fields in the original text, multiple sets of sentence / segment index sequences can be formed. Then, by aggregating the contextual order between the cited fields and their corresponding time fields and topic statements, it is determined whether they are in the same semantic block. If the three are adjacent or arranged with short intervals, the group is considered a valid correspondence, ultimately yielding the set of cited position correspondences.

[0085] S302: Based on the reference position correspondence set, extract each group of referenced segments and the connecting words between the preceding and following sentences, construct the guide word sequence and the response sentence sequence, calculate the distribution frequency offset ratio of the guide word sequence in the sentence group, determine the arrangement pattern of the referenced fields in the text, and obtain the sentence segment order offset set;

[0086] The formula for calculating the frequency shift ratio of the introductory word sequence in the sentence group is as follows:

[0087] ;

[0088] in, Representing the The frequency shift ratio of the introductory word sequence in the sentence group. Representative introductory words Normalized frequency in the current sentence group Representative introductory words Normalized frequencies in the response sentence sequence Representing the The total number of conjunctions in a sentence group. Representing the The number of introductory words in the response sentence sequence within a sentence group. Representing the The guiding word frequency difference adjustment factor for sentence groups represents the positive number used to avoid a zero denominator in the group;

[0089] The operation logic of this formula is as follows: First, perform normalized frequency statistics on each guiding word in the sentence group to obtain the frequencies in the current sentence group and the response sentence sequence respectively. Subsequently, calculate the mean of the absolute values of the frequency differences of all guiding words in this group between the two types of text segments, and multiply it by the adjustment coefficient defined by the maximum and minimum response frequency differences in the group . Then, add the result to the current word frequency to form an adjusted value as the numerator of the ratio. Next, divide it by the denominator formed by the frequency of the corresponding word in the response sentence plus a very small positive number . Take the absolute value of the ratios obtained for all guiding words and calculate the average to obtain the overall deviation ratio of the distribution frequencies of guiding words in the upper and lower sentences of the sentence group. The entire formula uses frequency difference adjustment weighting, normalization averaging, and anomaly avoidance means to form a multi-layer composite ratio structure to depict the deviation degree of the guiding word frequency distribution in semantic links;

[0090] Select sentence groups from news texts , including 4 guiding words (including "therefore", "however", "moreover", "finally"). The frequencies of guiding words in the current sentence group are as follows:

[0091] "therefore" appears 3 times;

[0092] "however" appears 2 times;

[0093]

[0094] "moreover" appears 1 time;

[0095] "finally" appears 2 times;

[0096] The total number of words in the sentence group is 100, so we have:

[0097] ;

[0098] ;

[0099] ;

[0100] ;

[0101] The frequencies of these words in the response sentence sequence are: "therefore": 4 times / 80 total words →

[0102] ; However: 1 time / 80 → ;

[0103] "In addition": 3 times / 80 → ;

[0104] "Final": 1 time / 80 → ;

[0105] calculate ;

[0106] Calculate the mean absolute value of the word frequency difference:

[0107] ;

[0108] ;

[0109] ;

[0110] ;

[0111] set up ;

[0112] Substitute into the calculation:

[0113] For each arrive Calculate separately:

[0114] Item 1:

[0115] ;

[0116] Item 2:

[0117] ;

[0118] Item 3:

[0119] ;

[0120] Item 4:

[0121] ;

[0122] Summary average:

[0123] ;

[0124] Numerical results analysis: The range of the baseline offset ratio is set to [0.8~1.2]. Values ​​exceeding this range are marked as having unstable sequential distribution.

[0125] The current result is 1.0469 ∈ [0.8, 1.2], which falls within the range of moderate offset.

[0126] The results indicate that there is a slight shift in the distribution of introductory words in the current sentence group across the quoted and response segments, but this does not constitute an imbalance in order and does not trigger reordering.

[0127] The advantage of the formula lies in the introduction of an adjustment coefficient. Weighted adjustment of the frequency differences of the guiding words amplifies the semantic mutation points; at the same time, the normalized average offset operation is introduced to enhance the consistency of the comparison between the lengths of the differentiated sentence groups, thereby enabling a sensitive response to the detection of abnormal order of quoted paragraphs in the whole.

[0128] S303: Based on the sentence segment order offset set, locate the start and end positions of the offset sentence group, and divide the reference range and corresponding reference method of each sentence segment according to the first appearance order of the referenced field in the text and the context paraphrase method to obtain the semantic offset fragment set;

[0129] First, locate the exact order of the first and last sentences in each group within the original text. Then, combine this with paragraph numbers or intra-sentence markers in the context to number and label the beginning and end positions of the sentence groups, for example, "sentence 1 of paragraph 2 to sentence 4 of paragraph 2." This location operation relies on the index table established during the sentence segment annotation stage of the original corpus and can determine the structural integrity of the segment by examining the initial verb or subject. After determining the position of the sentence group, extract the fields marked as references from each segment sequentially. These fields commonly appear as phrases with a reference intent, such as "according to xx" or "xx pointed out." Count the first occurrence of each field in the entire text. If the first occurrence of a reference field is outside the current offset sentence group, record its initial position index and determine whether the referenced content appears earlier or later based on the positional relationship within the current sentence group. Next, identify the contextual paraphrasing methods within the sentence segments, including referential or indirect quotations such as "it is said" or "it was mentioned later," analyze their distribution within the sentence groups, and compare them with the original paragraphs to see if there is a change in the semantic subject. By employing a dual constraint of citation fields and contextual paraphrasing, the citation content within each group of sentence segments is grouped. Segments belonging to the same citation source or semantically consecutively describing the same event are grouped into the same citation scope. Further, based on the tone, tense, or referential characteristics of each type of citation, they are categorized as either "direct citation" or "indirect citation." For example, sentences directly using quotation marks are identified as direct citations, while those without quotation marks but containing indicator words are marked as paraphrased citations. During this process, if any citation content is found to be misaligned or discontinuous with the logical order of the context, this abnormal grouping is recorded and used as a marker of semantic connection anomalies, resulting in a semantically offset fragment set.

[0130] Please see Figure 5 The specific steps of S4 are as follows:

[0131] S401: Based on the semantic offset fragment set, retrieve the order of the segments in the context, track the segment group to which the jump points, identify the index position in the structure of the preceding and following segments, compare and map the sequence distribution state to obtain the jump position index set;

[0132] First, the paragraph number and sentence order number of each offset segment in the original text are obtained, and an initial index list is constructed in sequence. Then, the context paragraphs corresponding to each offset segment are indexed and retrieved, and the contextual paragraphs related to that segment are extracted. Secondary annotations are performed according to their paragraph position number and sentence marker order in the original text. Then, the boundary lines are determined by combining the punctuation separation positions within the text segments. Semantic linking is performed on multiple jump segments associated with the same reference object in different paragraphs. Jump sentence groups with cross-segment expression behavior are marked one by one. At the same time, part-of-speech identification is performed on the connecting verbs and referential nouns involved in the context. If there is a cross-segment misalignment in position and semantic reference, the relevant part of speech is identified. In such cases, the skipped segment pair is recorded as an abnormal segment pair. Then, the content of the two sentences before and after the skipped segment is retrieved from the original text as a comparison group for positional reference matching. By analyzing the connection relationship between the skipped segments in the paragraph structure, the position where each group of skipped segments is interrupted in the context is identified, and the paragraph number and intra-sentence position index of its starting and ending segments are extracted to construct the arrangement trajectory of the skipped segments. Subsequently, based on the distribution of the overall paragraph structure in the entire corpus, a two-level mapping is performed on the position indexes of the semantic skipped segments and their preceding and following related segments, so that each skipped chain can find a precise positional anchor point in the original corpus. Finally, the index results of all skipped segment groups are summarized to obtain the skipped position index set.

[0133] S402: Based on the jump position index set, locate the first word and connecting word of the two segments on both sides of the jump area. According to the connection order and part-of-speech matching relationship of the words on both sides in the semantic chain, determine the logical connection direction between the preceding and following segments and obtain the reverse connection position group.

[0134] First, segment location is performed at each jump position in the index set to obtain the sentence order and specific content of the two segments before and after the jump in the original text. Then, the first word of each segment is extracted, along with the connecting words at the connecting positions within the segment. The first three to five words of each segment are used to construct a subset for first word identification. Connecting words within the segment are extracted sequentially to form a set of connecting words, and their position index and part-of-speech category are labeled. Functional conjunctions indicating contrast, cause and effect, progression, and conditional relationships are prioritized in the connecting word set, and their semantic relationship with the last word of the preceding segment is labeled. Simultaneously, the semantic expansion table constructed in the semantic chain is called to analyze the semantic connection path between the first word and the connecting words. If semantic jumps or decreased semantic relevance occur in the path, further analysis is performed. If the connection is misaligned, for example, if the first sentence begins with "therefore" and the second sentence begins with "but", the connection direction is reversed. By comparing the main part of speech of the first word with the logical function part of speech of the connecting word, it is determined whether the pairing pattern of the two in the semantic chain is arranged in a forward direction. If there is a situation where the connection between the preceding and following paragraphs has a reverse logical flow, its index position is marked as a reverse connection point. At the same time, its sentence order difference, part of speech mismatch degree and semantic connection level are recorded. For example, in a certain online text, if the sentence before the jump starts with "according to the report" and the sentence after the jump starts with "however, the facts show", it means that the semantic orientation of the preceding and following paragraphs is completely reversed. Such jump connection can be classified as a misaligned type. Finally, all jump connection positions with reverse relationships are summarized to obtain the reverse connection position group.

[0135] S403: Based on the reverse connection position group, extract the time identifier field and background description statement in the corresponding context, filter the relevant paragraph fragments, match the missing part position according to the context coverage and semantic coherence path, and obtain the semantic reconstruction chain group;

[0136] First, a contextual search is performed on the inverse connection position of each record. Time markers and background descriptions appearing in the two sentences preceding and following this position are extracted and semantically categorized. Phrases containing time sequence words or event prompts are assigned to the time description set, while short sentences describing environment, event background, or subject-object relationships are assigned to the background description set. Based on this, the frequency and position range of the two types of fields are compared. If the two types of fields overlap or intersect within the same segment, they are marked as semantic coverage intervals. Next, for each inverse connection point, adjacent paragraphs in its contextual corpus are retrieved, and the range of related sentence groups is extracted. Their connection order in the semantic chain is compared. If both the preceding and following segments contain time fields but lack transitional sentences, the interval is marked as a potentially missing segment. Then, noun entities and subject components appearing in the background description statements are identified. Sentences with the same subject but semantically discontinuous are aggregated, and their coverage is recorded by segment number. By statistically analyzing the frequency of verb associations and the subject consistency ratio in the semantic chain, the approximate location of the missing segment in the context is determined. For example, if a text contains the sentence "The meeting issued a statement afterward" and the next sentence "Media reports sparked discussion," without any time-connecting phrases in between, the system will treat this as a jump segment and extract relevant contextual paragraphs to determine the insertion position of the missing sentence. Finally, based on the correspondence between all back join points and their covered segments, the time identifier field, background statements, and adjacent semantic segments are arranged sequentially according to matching priority to form a complete contextual extension chain, resulting in a semantic reconstruction chain group.

[0137] Please see Figure 6 The specific steps of S5 are as follows:

[0138] S501: Based on semantic reconstruction chain groups, retrieve true / false judgment statements, quoted content statements and time description statements in the segment, locate the starting position and grammatical relationship of the statements in the paragraph, identify the connection path and semantic inheritance structure between statements, and obtain the structure set corresponding to the statements.

[0139] First, the labeled true / false statements, quoted statements, and time-description statements are extracted from the text segment set. For each type of statement, keywords or verb predicates are identified, and their starting character indices in the original paragraphs are located. The corresponding sentence range in the original text is extracted using the index value, and grammatical structure analysis is performed to further decompose the subject-predicate and modification / dependency relationships between sentence components. For example, in the sentence "The statement is false information," "statement" is identified as the subject, "false information" as the predicate, and "is" as the judgment word, positioning it as a true / false statement. Second, when processing quoted statements, sentences containing quotation marks, reporting verbs, or dissemination phrases are located. Phrases such as "allegedly," "reported," and "sources indicate" are identified as marker phrases to confirm that they are quoted statements, and paragraphs are segmented based on the position of these phrases. Third, for time-description statements, words with temporal expressions, such as "morning," "then," and "last year," are extracted. The semantic extension before and after their positions is considered to determine whether the time expression plays a dominant role in the event, thus deciding whether to include it in the main clause's time attribute. After the structured extraction of various sentences, the relative positions of all sentences are merged. Based on the sentence index values, semantic connection paths are constructed. The existence of semantic continuity between sentences is determined by the identification of conjunctions and the recurrence of subject and object. For example, if "government spokesperson" appears as the subject in the previous sentence and "speaker's response" appears in the next sentence, it is considered a semantic connection. To verify the effectiveness of the path, it is also necessary to identify whether there are structural conflicts caused by misplaced conjunctions or improper word order. For example, if "although...but..." appears in a disjointed position, it is recorded as a broken link interval. In addition, to improve the accuracy of structure recognition, the grammatical structure type in each sentence needs to be classified and labeled. Judgment sentences are classified as type A, quotation sentences as type B, and time sentences as type C, and they are numbered and stored in the structure classification list respectively. Finally, based on the correlation between the sentence's starting position, the internal structure type, and the connection path between the preceding and following sentences, the corresponding structure set of the sentence is obtained.

[0140] S502: Based on the statement correspondence structure set, identify the semantic extension relationship between statements, identify the order of truth and false judgment content and the order of quoted statements according to the logical sequence of the paragraphs, and aggregate statement combinations with related order and semantic coherence features to obtain a logically consistent combination set.

[0141] First, semantic extension and identification are performed on the true / false judgment statements and quoted statements. The core subjects, event verbs, and modifiers shared between the statements are extracted. Semantic comparison determines their connection direction in the logical chain. Each statement is arranged vertically according to its order of appearance, and the dependencies between them are marked. For example, in the text "The report stated that the event was true, but the relevant parties subsequently denied it," "The report stated" is considered a quoted statement, and "true" is considered a true / false judgment statement. The semantic inversion structure is identified through their sequential relationship. Next, the semantic extension relationship between each group of statements is analyzed. The distribution of logical judgment words (such as "is," "is," "was identified as") in the true / false judgment statements is statistically analyzed and compared with the reporting verbs in the quoted statements. If the reporting statement appears before the true / false judgment, the combination is determined to be a causal sequential relationship; if it appears afterward, it is determined to be a result inversion relationship. Then, based on the logical sequence of the text segments, true / false judgment statements and quoted statements are paired and aggregated. Logically adjacent and semantically continuous statements are connected, for example, binding "some netizens say" with "the news is not true" to form a complete logical chain. When time-description statements are involved in the connection, their position in the text is checked to ensure a match between their time order and the corresponding judgment statement. If the time description precedes the quoted statement, the combination is marked as a forward extension; if it follows, it is classified as a reverse extension. Furthermore, semantic coherence features are extracted from all combinations. Statements with similar keywords or repeated subject-verb relationships are grouped together; for example, "the statement is released" and "the statement is false" are semantically coherent combinations. Finally, all aggregated statement combinations are numbered according to logical order, and the index intervals between sentences are rearranged to form a set of statements with unified semantic orientation and sequential consistency, resulting in a logically consistent set of combinations.

[0142] S503: Based on the logically consistent combination set, determine whether the segment contains true / false judgment statements, quoted content statements and time description statements at the same time, remove content that does not have complete statements, extract semantically coherent and structurally complete sentence segments, and obtain the false information judgment conclusion group.

[0143] First, each segment combination is searched item by item to identify whether it contains three core components simultaneously: true / false statements, quoted content statements, and time description statements. During the process, the semantic tag sequence of the segment is extracted, and the categories and quantities of semantic tags are counted. If a segment lacks any type of tag, it is marked as incomplete and removed from subsequent analysis. Then, for the retained segments, logical connectors and verb predicate components within the sentences are extracted sequentially to determine the subject consistency and temporal relationship between the true / false statements and the quoted statements. For example, when the text contains "the report states that an event is true" and the following sentence is "the relevant party verified and denied the news," the system recognizes that the subject "event" is consistent in this segment combination, and the statements "true" and "denied" form a logical opposition, thus confirming the integrity of its semantic relationship. Next, the time description statements within each segment are matched and detected to identify the adjacent position relationship between the time field and the quoted statement. If the time phrase appears before the quoted statement, it is determined to be a preceding time modifier; if it appears after the judgment statement, it is considered a following explanation. The existence of both ensures the temporal logical closure of the segment. Furthermore, the connection between sentences in the retained segments is structurally examined, analyzing the consistency of conjunctions, pronouns, and verb forms. If the semantic chain is continuous and structurally aligned, the segment is divided into a semantically coherent segment; otherwise, it is marked as a broken segment and removed. To improve the accuracy of the screening, the syntactic level of the segments is also compared, and subordinate sentences in multi-sentence complex segments are extracted. If their semantic content is independent and contains three types of sentence components, they are retained as independent segments. Taking actual text as an example, if a news segment contains "Netizens say the video is true, and relevant parties say the investigation is still ongoing," then the segment contains the quoted statement "netizens say," the truth judgment "true," and the time statement "still ongoing," and is therefore identified as a structurally complete segment. Finally, all the selected segments are logically merged and arranged in order. A unified index table is established based on their semantic coherence and component completeness. The output is a set of segments with judgment, reference and time, and the false information judgment conclusion group is obtained.

[0144] Please see Figure 7 A network misinformation detection system based on large model fine-tuning includes:

[0145] The data processing module acquires false information text from Chinese online corpora, extracts source tags, jump sentences and context information, determines the order logic of source fields and jump sentences, determines the semantic span order, and obtains a semantic fragment association set.

[0146] The semantic judgment module, based on the semantic fragment association set, separates sentences that state facts, statements with emotional bias, and sentences that provide source information from a segment, judges the connection between sentences, and traces the semantic extension path along the sentence order to obtain the semantic judgment chain structure.

[0147] The offset recognition module is based on the semantic decision chain structure. It extracts the corresponding sentences of the quoted content and compares their positions with the quoted segments in the original text. It identifies inconsistent sentence combinations and obtains a set of semantic offset fragments.

[0148] The reconstruction processing module identifies the sequential jump position based on the semantic offset fragment set, analyzes the starting words and connecting words of the sentences on both sides of the breakpoint, adjusts the position of the segment before and after, and adjusts the position of the segment before and after according to the causal relationship and reverse expression method between the sentences before and after to obtain the semantic reconstruction chain group.

[0149] The judgment output module is based on the semantic reconstruction chain group, extracts the true or false judgment, the quoted content and time information, analyzes whether each statement includes complete information, retains the statements that include all information, and obtains the false information judgment conclusion group.

[0150] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for detecting fake information on the network based on large model fine-tuning, characterized in that, Includes the following steps: S1: Obtain misquoted and false information from Chinese online corpus, extract source tag fields, jump sentences and context information of text paragraphs, determine the semantic span order, and obtain a semantic fragment association set; S2: Based on the semantic fragment association set, separate factual statements, emotionally charged statements and sentences that explain the source information in the paragraph. Determine the connection between the preceding and following sentences according to the position of each sentence in the paragraph. Track the semantic extension path along the sentence order to obtain the semantic judgment chain structure. S3: Based on the semantic determination chain structure, extract the corresponding statement of the quoted content and compare its position with the quoted segment in the original text, identify the inconsistent statement combination, and obtain the semantic offset fragment set; S4: Based on the set of semantic offset segments, identify the location of sequential jumps, analyze the starting words and connecting words of the sentences on both sides of the breakpoint, and adjust the position of the segments according to the causal relationship and reverse expression between the sentences to obtain a semantic reconstruction chain group. S5: Based on the semantic reconstruction chain group, identify the true / false judgment, reference information and time description, analyze whether each statement includes the corresponding related information, retain the complete information statement, and obtain the false information judgment conclusion group; The term "skipping sentence" refers to a sentence in the text that breaks away from the original logical order of the context, skips the middle paragraph, and directly points to other content. The semantic span refers to the semantic interval between sentences, the logical and temporal distance that information traverses from one segment to another; The sentence order direction refers to the logical direction of the arrangement order in the paragraph and the semantic progression, to determine whether the information flow is forward or backward; The semantic extension path refers to the logical chain in which semantics are sequentially connected and propagated between sentences and segments, reflecting the semantic connection between facts, emotions, and cited information; The specific steps of S1 are as follows: S101: Obtain misquoted false information text from Chinese online corpus, extract source tag fields, skipped sentences and context identifiers from text paragraphs, determine the positional order of source fields and skipped sentences, and obtain the citation order judgment result; S102: Based on the citation order judgment result, compare the semantic span order between the jump segment statement and the context identifier, combine the time sequence information and the citation direction field, identify the part where the position is reversed and the semantic connection direction is inconsistent, and obtain the segment citation offset information; S103: Based on the segment reference offset information, filter the positions of fields with misaligned order and direction, compare them with the context connection interval, extract the logical offset part between segments, and obtain the semantic segment association set; The specific steps of S2 are as follows: S201: Based on the semantic fragment association set, separate factual statements, emotional statements and information source segments in the segment, identify the sentence punctuation structure and sentence ending punctuation mark type of the sentence, and build a segment index list according to type to obtain a sentence classification structure set; S202: Based on the sentence classification structure set, extract the start and end positions of each type of sentence segment in the original paragraph, perform a linear comparison of the order position difference between adjacent sentence segments, mark the order jump points, remove disordered segments, and obtain the sentence sequence connection diagram. S203: Based on the segment sequence connection diagram, locate the jump breakpoint, analyze the connecting words and transition expressions of the adjacent segments, determine the sequential relationship and consistency of the citation direction between adjacent segments, filter semantic conflict intervals and record the connection path to obtain the semantic judgment chain structure. The specific steps for S3 are as follows: S301: Based on the semantic determination chain structure, extract the content fields identified as references in each sentence, call the time description field and topic guidance statement in the original text content, retrieve the sentence position of the referenced content field in the original text, record the sentence index interval, the sentence index interval is the position of the sentence in the paragraph, and obtain the reference position corresponding set; S302: Based on the reference position corresponding set, extract each group of referenced segments and the connecting words between the preceding and following sentences, construct a sequence of guiding words and a sequence of response sentences, calculate the distribution frequency offset ratio of the guiding word sequence in the response sentence sequence, determine the arrangement pattern of the referenced fields in the text, and obtain the sentence segment order offset set; S303: Based on the sentence segment order offset set, locate the start and end positions of the offset sentence segment set, and divide the reference range and corresponding reference method of each group of sentence segments according to the first appearance order of the referenced fields in the text and the context paraphrasing method to obtain the semantic offset fragment set; The specific steps of S4 are as follows: S401: Based on the semantic offset fragment set, retrieve the arrangement order of the segments in the context, track the segment group to which the jump points, identify the index position in the preceding and following segment structures, compare and map the distribution state of the segment index sequence to obtain the jump position index set; S402: Based on the jump position index set, locate the first word and connecting word of the two segments on both sides of the jump area, and determine the logical connection direction between the preceding and following segments according to the connection order and part-of-speech matching relationship of the words on both sides in the semantic chain, and obtain the reverse connection position group; S403: Based on the reverse connection position group, extract the time identifier field and background description statement in the corresponding context, filter the relevant paragraph segments, match the missing part position according to the context coverage and semantic coherence path, and obtain the semantic reconstruction chain group; The specific steps of S5 are as follows: S501: Based on the semantic reconstruction chain group, retrieve the true / false judgment statements, quoted content statements and time description statements in the segment, locate the starting position and grammatical relationship of the statements in the paragraph, identify the connection path and semantic inheritance structure between statements, and obtain the structure set corresponding to the statements. S502: Based on the structure set corresponding to the statement, identify the semantic extension relationship between the statements, identify the order of the true / false judgment content and the quoted statement according to the logical sequence of the paragraph, and aggregate the statement combination with the characteristics of related order and semantic coherence to obtain a logically consistent combination set. S503: Based on the logically consistent combination set, determine whether the segment simultaneously contains a true / false judgment statement, a quoted content statement, and a time description statement; remove content that does not have a complete statement; extract semantically coherent and structurally complete sentence segments to obtain a false information judgment conclusion group.

2. The network misinformation detection method based on large model fine-tuning according to claim 1, characterized in that, The semantic fragment association set includes source tag fields, jump sentences, context information, sequential logic, semantic span order, and misdirection relationships. The semantic judgment chain structure includes factual sentences, sentences with emotional bias, source information sentences, connection between preceding and following sentences, semantic extension paths, breakpoints, word order jumps, and consistency of citation direction. The semantic offset fragment set includes quoted content sentences, inconsistent sentence combinations, citation start and end positions, context connection methods, and misplaced fragments. The semantic reconstruction chain group includes sequential jump positions, sentence start words, connecting words, causal relationships, reverse expressions, time descriptions, background clues, and missing sentence positions. The false information judgment conclusion group includes truth / false judgments, quoted content, time information, sentences containing complete information, and sentences lacking information.

3. A network misinformation detection system based on large model fine-tuning, characterized in that, The system is used to implement the network misinformation detection method based on large model fine-tuning as described in any one of claims 1 and 2, and the system comprises: The data processing module acquires false information text from Chinese online corpora, extracts source tags, jump sentences and context information, determines the order logic of source fields and jump sentences, determines the semantic span order, and obtains a semantic fragment association set. Based on the semantic fragment association set, the semantic judgment module separates sentences stating facts, statements with emotional tendencies, and sentences that provide source information from the segment, judges the connection between the preceding and following sentences, and traces the semantic extension path along the sentence order to obtain the semantic judgment chain structure. Based on the semantic determination chain structure, the offset recognition module extracts the corresponding sentences of the quoted content and compares their positions with the quoted segments in the original text, identifies inconsistent sentence combinations, and obtains a set of semantic offset fragments. Based on the semantic offset fragment set, the reconstruction processing module identifies the sequential jump position, analyzes the starting words and connecting words of the sentences on both sides of the breakpoint, adjusts the position of the segment before and after, and adjusts the position of the segment before and after according to the causal relationship and reverse expression method between the sentences before and after to obtain the semantic reconstruction chain group. The judgment output module extracts the true / false judgment, reference content and time information based on the semantic reconstruction chain group, analyzes whether each statement includes complete information, retains the statements that include all information, and obtains the false information judgment conclusion group.