Intelligent chatting method and system based on large model

By identifying and supplementing semantic coherence and focus shifts in multi-turn dialogues, the problem of insufficient context parsing in traditional intelligent chat methods is solved, achieving higher semantic coherence and dialogue consistency, and improving the fluency and accuracy of human-computer dialogue.

CN121636671AInactive Publication Date: 2026-03-10上海笑聘网络科技有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-04
Publication Date
2026-03-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional intelligent chat methods, in multi-turn dialogue scenarios, have limited ability to analyze context and are unable to effectively identify semantic chain breaks. This results in responses that deviate from the current dialogue focus, lack contextual coherence and linguistic structural coherence, and affect the fluency of human-computer dialogue and the accuracy of semantic understanding.

Method used

By identifying semantic transition regions in adjacent structures where the target tends to be consistent, a set of semantic coherence chain labels is generated. The direction of semantic focus shift and the logical progression relationship between sentences are measured. Semantic stable fragment aggregation groups are constructed and combined with a large language model to generate response content. Candidate sentence groups with a deviation range greater than the semantic center distribution threshold are identified and supplemented to optimize the contextual consistency of the response content.

Benefits of technology

It enhances the semantic coherence and topic consistency of response content in multi-turn dialogues. By locating the semantic chain break point and detecting focus shift, it generates topic derailment indicators, optimizes the contextual semantic path of response content, and improves the fluency and semantic understanding accuracy of human-computer dialogue.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121636671A_ABST
    Figure CN121636671A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent chatting, in particular to an intelligent chatting method and system based on a large model, and the method comprises the following steps: extracting verb combinations and target phrases, labeling semantic transition and fracture, generating coherent chain tags, measuring offset to judge derailment, aggregating stable chunks, aligning centroid lexical items, recognizing focus offset, and reconstructing low-efficiency sentence segments. And generating a control response text. According to the method, the topic content, separated from the current semantic path, in the response statement can be recognized by extracting the action verb combination, limiting the target word group and the logic connection mark item, constructing the structure sequence according to the word order and combining semantic unit lexical position matching and offset distance measurement in the focus transfer direction; statistics of sentence pattern templates with consistent semantic directions is carried out, role change or target replacement phrases are eliminated, sentence phrases with low scores are eliminated, and reconstruction and rewriting at a paragraph level are carried out, so that semantic coherence and topic consistency of response contents in multiple rounds of dialogues are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent chat, and in particular to an intelligent chat method and system based on a large model. BACKGROUND

[0002] The technical field of intelligent chat involves language communication between a computing device and a user through natural language processing, machine learning, and human-computer interaction, including semantic understanding, context modeling, dialogue generation, intent recognition, and response generation. The goal of this technical field is to improve the system's ability to analyze natural language input and generate language output that fits the context, enabling smooth human-computer dialogue. The technology is applied in online customer service, virtual assistants, educational question and answer systems, and enterprise knowledge management, and relies on language model training, corpus construction, and label mechanism design in the implementation process.

[0003] Among them, the traditional intelligent chat method refers to matching and responding to user input statements through rule engines, retrieval dialogue systems, or shallow neural network models. It mainly generates responses based on pre-defined templates, keyword recognition, or retrieval results from question and answer corpora. In this process, a pre-defined intent classifier or pattern matcher is usually used to analyze user input, and a static corpus and a limited dialogue state machine are used to call and organize response content. Although this type of solution is controllable, it has limited semantic coverage and insufficient response flexibility when dealing with open, multi-turn, or context-dependent dialogue tasks.

[0004] Traditional methods rely on rule engines or static corpora to match user input and generate response content. In multi-turn dialogue scenarios, due to limited context analysis capabilities, there are breaks in the semantic chain that cannot be effectively identified. When the dialogue topic shifts in multiple input rounds, traditional methods cannot effectively track the semantic trend in continuous turns, causing the response content to easily deviate from the current dialogue focus. For example, when a user asks multiple sub-questions, the system tends to generate answers independently based on the latest input, ignoring the intent continuation relationship in previous turns. In addition, due to the poor adaptability of template statements and finite state machine structures to changes in input language, it is difficult to capture fine-grained target changes and sentence variants, resulting in a lack of context coherence and language structure continuity in response content, affecting the smoothness of human-computer dialogue and the accuracy of semantic understanding. SUMMARY

[0005] The purpose of the present application is to solve the problems in the prior art and to provide an intelligent chat method and system based on a large model.

[0006] To achieve the above purpose, the present application adopts the following technical solution: An intelligent chat method based on a large model, comprising the following steps: S1: According to the user input sentence in the current session, identify the semantic transition area of the target trend consistency in the adjacent structure, judge whether the frequency and the length of the syntax sequence offset exceed the set turn continuity critical value, call the large language model structured semantic labeling interface to generate the semantic segmentation boundary, and generate the user input semantic coherent chain label set; S2: Call the user input semantic coherent chain label set, measure the offset distance according to the semantic focus shift direction and the inter-sentence logical progression relationship, compare the measured value with the preset topic offset boundary range, and generate the response sentence topic derailment indication item; S3: According to the action sentence group marked as logical target clear in the first five turns of user input, count the number of sentence templates with consistent semantic direction, and exclude the language segments that trigger role change or target replacement, construct the language block with consistent direction in the dialogue context, and generate the semantic stable segment aggregation group; S4: Based on the semantic stable segment aggregation group, combine the attention concentration position information in the large language model generated response, identify the candidate sentence group whose deviation range is greater than the semantic gravity distribution threshold, and determine it as the semantic focus shift node, and generate the response content focus shift node set.

[0007] As a further scheme of the application, the user input semantic coherent chain label set includes semantic structure stability markers, syntax logical break points and semantic transition unit boundaries, the response sentence topic derailment indication item specifically is a focus word offset value, a logical progression direction deviation rate and a candidate sentence segment position index, the semantic stable segment aggregation group includes a target word frequency set, a semantic direction consistent template set and a context window segment index table, and the response content focus shift node set specifically refers to a semantic gravity alignment matrix, a sentence segment number deviating from the focus and an attention value distribution aggregation statistics.

[0008] As a further scheme of the application, the acquisition step of the user input semantic coherent chain label set is specifically: S101: According to the user input sentence in the current session, extract the action verb combination, the limited target word group and the logical connection marker in the continuous two turns of input sentences, arrange and construct the original semantic structure sequence according to the syntax sequence, label according to the part of speech and grammatical role of the semantic segment in the sentence structure, and generate the semantic instruction construction sequence; S102: Call the semantic instruction construction sequence, identify whether the target trend is consistent between adjacent combinations according to each group of action target relationship, calculate the position difference value and frequency of the logical jump, compare the difference value result with the set semantic continuity critical interval, obtain the structure paragraph exceeding the interval and mark the break position, and generate the target consistency offset label group; S103: Based on the target consistency offset label group, according to the break position index, call the large language model structured semantic labeling interface to identify the subject, predicate, object, time adverbial and conditional restriction word group in the input sentence, combine the break boundary information in the label group to perform semantic level segmentation, and generate a user input semantic coherent chain label set.

[0009] As a further scheme of the present application, the obtaining step of the response sentence question derailment indication item is specifically: S201: Call the user input semantic coherent chain label set, according to the semantic unit marked as tending to break, match the candidate response sentence generated by the current large language model piece by piece, extract the word position index sequence of the key word in the semantic unit in the response sentence, establish the word position mapping table of the input and the response semantic unit, and generate a semantic position mapping value; S202: According to the corresponding word position index in the semantic position mapping value, extract the context block of the focus word before and after in the response sentence, obtain the sequence change direction of the focus word in the inter-sentence sequence, and extract the progressive, turning or causal structure combined with the main sentence dependent structure label in the context block, utilize the inter-sentence semantic migration direction angle and the structure turning label frequency, and calculate to obtain a semantic focus offset value; S203: Call the semantic focus offset value, compare the interval with the set topic offset boundary range, filter the sentence position index whose offset value exceeds the boundary, and aggregate the offset index, obtain the coverage proportion of the topic focus derailment segment in each response, and generate a response sentence question derailment indication item.

[0010] As a further scheme of the present application, the formula for calculating to obtain a semantic focus offset value is specifically ; Wherein, represents the semantic focus offset value, represents the normalized value of the semantic transfer angle of the two rounds of focus words, represents the normalized value of the turning connection label frequency in the context structure, represents the number of candidate sentence segments in the current response text, represents the word vector distance between the first main sentence focus word and the previous round of focus words, represents the syntactic connection strength of the first focus word in the context, represents the length of the sentence segment corresponding to the first focus word, is a constant offset value, represents the number of focus words in the current sentence.

[0011] As a further scheme of the present application, the obtaining step of the semantic stable fragment aggregation group is specifically: S301: According to the action sentence group marked as a logical target explicit in the first five rounds of user input, extract the target phrase item in the word order arrangement, count the frequency of occurrence in each round of input, calculate the proportion of the frequency in all sentences, select the phrase content with a frequency exceeding a set semantic representative threshold, and generate a target phrase distribution rate; S302: Based on the target phrase distribution rate, extract the sentence structure fragment matching the phrase set according to the frequency ranking, judge whether the predicate type and the logical direction of the word order are consistent, count and generate a structure statistics table for the sentence pattern that meets the consistent condition of semantic progression direction, exclude the sentence fragment including role subject switching or purpose verb replacement, and generate the number of semantic direction consistent templates; S303: Call the number of semantic direction consistent templates, select the continuous template coverage fragment according to the start and end position index of each template corresponding to the sentence segment, combine the semantic cache information retained by the context window of the large language model, judge the residual number of the sentence segment in the window and extract the sentence pattern fragment with an occurrence frequency greater than the retention reference frequency, construct the sentence group block with convergent semantic direction in the turn, and generate the semantic stable fragment aggregation group.

[0012] As a further scheme of the present application, the obtaining step of the response content focus shift node set is specifically: S401: Based on the semantic stable fragment aggregation group, extract the action verb, result target word and scene modifier included in the sentence segment, select the repeated frequency of the word item in each sentence segment, record the position information of the high-frequency word group in the original sentence segment, and mark the expression intensity maximum word item at the beginning and end of the sentence, generate a semantic gravity word group comparison table; S402: According to the semantic gravity word group comparison table, identify the key expression word items at the beginning and end of the current response candidate sentence, detect the vector distance, semantic direction matching degree and content structure overlap ratio between the extracted target semantic gravity word group, calculate the deviation intensity between each group of sentences and the semantic gravity, judge whether it exceeds the semantic boundary according to the set semantic gravity distribution interval, and generate semantic separation shift information; The formula for calculating the deviation intensity between each group of sentences and the semantic gravity is specifically: ; Wherein, is the semantic separation shift value, is the word vector normalization value of the i-th target semantic gravity word, is the word vector normalization value of the i-th response focus word, is the word vector normalization value of the i-th target semantic gravity word, is the word vector normalization value of the i-th response focus word, is the word vector normalization value of the i-th target semantic gravity word, Normalized value of semantic direction matching factor for each semantic term For the first The normalized value of the term density of the segment containing each term. For the first The normalized value of the semantic diffusion range of each focal word in the response sentence. For the current candidate sentence segment The mean, To align the number of semantic pairs, In response to the number of focus words; S403: Based on the semantic decoupling offset information, according to the identified sentence group numbers whose offsets exceed the distribution threshold, summarize the candidate sentence segment index and offset frequency of semantic decoupling, count the proportion of offset nodes appearing in each round of candidate responses, establish the position information and matching value table of offset nodes, obtain the continuous decoupling blocks in the sentence group, and generate the focus offset node set of the response content.

[0013] As a further aspect of the present invention, the method further includes the following steps: S5: Based on the topic off track indicator of the response statement and the focus offset node set of the response content, identify the sentence segments in the sorting results whose scores are lower than the lower limit of the context alignment score standard, call the large language model to generate the previous round of content reconstruction fragments in the cache, delete the marked fragments at the paragraph level and back-write the semantic anchor points, and generate the response output text after semantic path control. The response output text after semantic path control includes control paragraph text fragments, semantic anchor replacement position table, and focus backtracking supplementary terms.

[0014] As a further aspect of the present invention, the step of obtaining the response output text after semantic path control specifically includes: S501: Based on the topic derailment indicator of the response statement and the focus offset node set of the response content, extract the semantic span value and focus word position index of the corresponding sentence segment respectively, calculate the matching degree score between the keyword span interval and the overlapping position of the context focus in each sentence, construct the content validity ranking benchmark of the statement in the context, establish the sentence segment ranking matrix, and generate content validity ranking information. S502: Call the content validity sorting information, extract the original index number and semantic anchor position according to the sentence and segment content whose score is lower than the lower limit of the set context alignment score standard, retrieve the corresponding content block in the previous generation cache, compare the missing focus words in the current content with the anchor words in the previous round, extract the words before and after the missing keywords and establish a supplementary index table, and generate a semantic anchor supplementary parameter group. S503: Based on the semantic anchor point completion parameter group, locate the boundary range of the sentence segment to be deleted in the current response segment according to the position index value and the term structure content, and perform paragraph-level deletion and structural replacement in combination with the semantic anchor point cache content of the previous round, semantically fill the focus word group towards the original anchor word group, and generate the response output text after semantic path control.

[0015] A large-model-based intelligent chat system, wherein the large-model-based intelligent chat system is used to implement the above-mentioned large-model-based intelligent chat method, the system comprising: The input coherence analysis module identifies semantic transition regions with consistent targets in adjacent structures based on user input statements in the current session, determines whether the frequency and word order offset length exceed the set turn continuity threshold, calls the large language model structured semantic annotation interface to generate semantic segmentation boundaries, and generates a set of user input semantic coherence chain tags. The topic derailment identification module calls the user input semantic coherence chain tag set, measures the offset distance according to the semantic focus shift direction and the logical progression relationship between sentences, compares the measured value with the preset topic offset boundary range, and generates a response statement topic derailment indicator. The semantic stability analysis module counts the number of sentence templates with consistent semantic direction based on the action statement groups marked with clear logical goals in the first five rounds of user input, and filters out the segments that trigger role changes or goal replacements, constructs language blocks with consistent direction in the dialogue context, and generates semantically stable fragment aggregation groups. The focus shift identification module, based on the semantically stable segment aggregation group, combines the attention concentration position information generated in the response by the large language model, identifies candidate sentence groups whose deviation range is greater than the semantic center of gravity distribution threshold, and determines them as semantic focus shift nodes, generating a set of focus shift nodes in the response content. The backtracking and supplementing output module identifies sentence segments in the sorting results whose scores are lower than the lower limit of the context alignment score standard based on the topic off track indicator and the focus offset node set of the response statement. It calls the large language model to generate the previous round of content reconstruction fragments in the cache, performs paragraph-level deletion and semantic anchor backtracking and supplementing on the marked fragments, and generates the response output text after semantic path control.

[0016] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In this invention, by extracting action verb combinations, defining target phrases and logical connection markers, and constructing a structural sequence according to word order, it is possible to locate semantic transition regions that tend towards consistency in the semantic chain and detect turn-ending interruptions, effectively marking the semantic coherence between input statements. By combining semantic unit word position matching and focus shift direction offset distance measurement, it is possible to identify topic content in response statements that deviates from the current semantic path and generate topic derailment indicators. By statistically analyzing sentence templates with consistent semantic direction and eliminating sentence segments with role changes or target replacements, it is possible to construct semantically stable segment aggregation groups while retaining context sentence fragments, achieving the filtering and optimization of response content under the dimension of contextual consistency. Relying on the weighted sorting mechanism of focus alignment degree and semantic span, it is possible to exclude low-scoring sentence segments and perform paragraph-level reconstruction and supplementation, enhancing the alignment and focus aggregation degree of response content with the contextual semantic path, and improving the semantic coherence and topic consistency of response content in multi-turn dialogues. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of the workflow of the present invention; Figure 2 This is a detailed flowchart of S1 of the present invention; Figure 3 This is a detailed flowchart of the S2 process of the present invention; Figure 4 This is a detailed flowchart of the S3 process of the present invention; Figure 5 This is a detailed flowchart of the S4 process of the present invention; Figure 6 This is a detailed flowchart of S5 of the present invention; Figure 7 This is a system flowchart of the present invention. Detailed Implementation

[0019] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0020] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0021] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.

[0022] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0023] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0024] Please see Figure 1 This invention provides a technical solution: an intelligent chat method based on a large model, comprising the following steps: S1: Based on the user input statements in the current session, extract the action verb combinations, target phrases, and logical connection markers from the two consecutive input statements, arrange them into a structural sequence according to the original word order, identify semantic transition regions where the target tends to be consistent in adjacent structures, determine whether the frequency of the break position in the semantic chain segment and the word order offset length exceed the set turn continuity threshold, call the large language model structured semantic annotation interface to generate semantic segment boundaries, and generate a set of user input semantic coherence chain tags; S2: Call the user input semantic coherence chain tag set, match the corresponding semantic unit word position in the candidate response statement currently generated by the large language model according to the semantic unit marked as tending to break, and measure the offset distance according to the semantic focus shift direction and the logical progression relationship between sentences. Compare the measured value with the preset topic offset boundary range to generate the response statement topic derailment indicator. The semantic focus shift direction is the change in the semantic trajectory represented by the target word that is significantly focused on in the output of BERT or T5 type models. The logical progression relationship between sentences depends on the topic shift level identified in the syntactic tree from the main clause to the subordinate clause and the subject-verb agreement path. The topic offset boundary range is generally set to the upper and lower limits of three times the standard deviation of the cosine distance of the continuous input focus words. S3: Based on the action statement groups marked as having clear logical goals in the user input of the first five rounds, extract the set of target phrases that appear frequently in the word order, count the number of sentence templates with consistent semantic direction, and filter out the segments that trigger role changes or target replacements. Combine the sentence fragments retained by the context window mechanism of the large language model to construct language blocks with consistent direction in the dialogue context and generate semantically stable fragment aggregation groups. S4: Based on the semantically stable fragment aggregation group, a control group is established according to the extracted target semantic core words and the key expression words at the beginning and end of the current response candidate content. Combined with the large language model, attention concentration position information in the response is generated, candidate sentence groups with a deviation range greater than the semantic core distribution threshold are identified and determined as semantic focus offset nodes, and a set of focus offset nodes in the response content is generated. The semantic centroid distribution threshold can be defined as the boundary of the interval twice the standard deviation of the average vector distance between the TF-IDF vector of the target term and the vector distance between each term in the response sentence. It is used to determine whether the response text deviates from the user's target term focus area as a whole. S5: Based on the topic off track indicator of the response statement and the focus offset node set of the response content, a content validity ranking table is established by assigning weights according to the semantic span and focus alignment degree within the sentence. The content of the sentence segment with a score lower than the lower limit of the context alignment score standard in the ranking result is identified. The large language model is called to generate the previous round of content reconstruction fragments in the cache. The marked fragments are deleted at the paragraph level and semantic anchor points are back-written to generate the response output text after semantic path control. The context alignment score can be defined as the average similarity between the response fragment and the content of the previous 5 rounds of conversation in the embedding space. The BERTScore or BLEURT index is often used to filter out non-matching content outside the logical dialogue context. The user input semantic coherence chain tag set includes semantic structure stability markers, word order logic breakpoints, and semantic transition unit boundaries. The response statement topic derailment indicators specifically include focus word offset values, logical progression direction deviation rates, and candidate sentence segment position indices. The semantic stable segment aggregation group includes target word frequency sets, semantic direction consistency template sets, and context window segment index tables. The response content focus offset node set specifically refers to the semantic center alignment matrix, derailed sentence segment numbers, and attention value distribution aggregation statistics. The response output text after semantic path control includes control paragraph text segments, semantic anchor replacement position tables, and focus backtracking supplementary terms.

[0025] Please see Figure 2 The specific steps for obtaining the semantic coherence chain tag set input by the user are as follows: S101: Based on the user input statements in the current session, extract the action verb combinations, target phrases and logical connection markers from the two consecutive input statements, arrange them in word order to construct the original semantic structure sequence, and mark the semantic fragments according to their parts of speech and grammatical roles in the sentence structure to generate a semantic instruction construction sequence. To obtain two consecutive rounds of user input statements in the current session, the first step is to extract the action verb combinations contained in each round of statements. This step can be achieved through rule-based recognition based on part-of-speech tagging. All words with verbs tagged as "VB" are included in the candidate set, and their subjects and objects are extracted to form the basic action structure, combined with syntactic dependency structures. For example, in the first round of input "I want to query orders," "query" is extracted as the action verb, and in "I need to cancel orders," "cancel" is extracted as the action verb. Then, target phrases are identified. Here, based on the object position or the direct object of the verb, phrases such as "orders" are selected as target phrases, and their word positions and intra-sentence dependencies are recorded. Logical connection markers are then identified from the statements, extracting markers that express logical relationships between sentences, such as "then," "at the same time," and "but," and establishing a connection index between these markers and the action target pairs. Next, the extracted action phrases, target phrases, and logical markers are sorted and constructed according to the original word order to form a sequence. The relative order number of each combination in the original sentence is recorded. When tagging each semantic segment with part of speech, a standard word class classification table is used, with action words tagged as VB, target words as NN, and logical terms as CC. At the same time, hierarchical structural labels of grammatical roles are added, such as SVO, SVC, etc., and semantic roles are assigned to its sentence structure. For example, if the input is "I submit an application and the system processes it automatically", then "submit" is extracted as the action verb, "apply" as the target phrase, and "after" as the logical connector. In its sentence structure, "I" is the subject, forming an SVO structure. Finally, the above extraction results are integrated into a triple combination structure, and the original word order index and grammatical annotation information are attached to it to generate a semantic instruction construction sequence.

[0026] S102: Call semantic instructions to construct a sequence, identify whether the target trends between adjacent combinations are consistent based on the target relationship of each group of actions, calculate the position difference and frequency of logical jumps, compare the difference results with the set semantic continuity critical interval, obtain the structural segments that exceed the interval and mark the break position, and generate a target consistency offset label group. After constructing the action-target relationship for each group in the sequence by invoking semantic instructions, it is necessary to determine the consistency of target tendencies between adjacent combinations based on semantic positional continuity. A comparison operation is performed on all adjacent action-target groups to determine whether their target word meanings belong to the same category. For example, the thesaurus set defined by WordNet is used to calculate their affiliation. If both groups of targets belong to the category of "order processing," they are considered to have a consistent tendency; otherwise, they are considered to have a broken tendency. Simultaneously, the positional difference of the semantic structure's jump degree is calculated by subtracting the index number of the current combination in the original sequence from the index number of the previous combination to obtain the logical jump value. If the jump value exceeds a set critical interval, it is determined as a semantic interruption. The interval can be set as a combination of semantic sequence interval values ​​greater than 3. When the difference between a pair of combinations is 4, the combination is marked as a breakpoint. Then, the frequency of all breakpoints in the instruction sequence is counted. If the frequency exceeds 30% of the total number of combinations, a continuity break judgment is formed. For example, if there are 3 combinations with jump values ​​greater than 3 in 10 action target groups, it is determined that the sentence structure has a strong break trend. A break label is generated, and the index value of the combination at the breakpoint is located and marked as a break structure segment. It is then bound to the semantic block corresponding to the sentence structure to obtain the mapping relationship between the semantic continuity structure and the break area. Finally, a target consistency offset label group is generated.

[0027] S103: Based on the target consistency offset label group, according to the break position index, call the large language model structured semantic annotation interface to identify the subject, predicate, object, time adverbial and conditional restriction phrases in the input sentence, and combine the break boundary information in the label group to perform semantic hierarchical segmentation to generate a set of semantic coherence chain labels for user input.

[0028] Based on the breakpoint index in the target consistency offset label group, the large language model structured semantic annotation interface needs to be called at the current sentence level to obtain the subject, predicate, object, and modifying components involved. During this process, the original input sentence is fed into the model to obtain its output structured annotation results, and the subject-predicate-object combination structure in each sentence segment is filtered out. For parts containing multiple main clauses or parallel complex sentence structures, each clause is extracted using clause segmentation rules, and the sentence beginning, sentence ending, and parallel components are extracted separately. Additional components such as time adverbs, place adverbs, and conditional phrases are identified in the structured results. Next, the structured semantic components are compared with the target... The break index in the consistency offset label group is matched, and structural components are extracted before and after the break position. The semantic segmentation boundary at the break point is constructed. If the structure before the break is a complete subject-verb-object structure and the start after the break is a conditional modifier, it is determined to be a semantic level segmentation point. Based on this, the statement is divided into blocks. Each structural block is further labeled with its hierarchical relationship. For example, the main clause is level one, the conditional statement is level two, and the modifying clause is level three. A semantic level labeling table structure is formed. Finally, the divided semantic structural blocks are integrated and the break position information and semantic backbone label are added to generate the user input semantic coherence chain label set.

[0029] Please see Figure 3 The specific steps for obtaining the off-topic indicator in the response statement are as follows: S201: Call the user input semantic coherence chain tag set, match the candidate response sentences generated by the current large language model one by one according to the semantic units that have been marked as tending to break, extract the word position index sequence of keywords in the semantic units in the response sentences, establish a word position mapping table between the input and response semantic units, and generate semantic position mapping values. The system retrieves semantic units marked as tending towards breakage from the user-input semantic coherence chain tag set. First, it extracts all semantic structure fragments labeled with breakage tags from the tag set. For each breakage unit, it extracts its internal keywords, including the central action verb, the instruction target phrase, and logical connectors, and establishes a correspondence between these keywords and their original word index values. Simultaneously, it binds the labeled grammatical function role information to the semantic relation structure. Then, it obtains the candidate response sentences generated by the current round of large language model, and splits these candidate response sentences into chunks according to semantic boundaries such as periods, semicolons, and connecting adverbs. Each segment is then read sequentially, and the words appearing within are labeled with part-of-speech tags and semantic vector embeddings for matching. A keyword lookup table is established, and the keywords are then used to... The similarity matching strategy identifies terms in the candidate response that have a similarity of more than 0.75 with the keywords in the input fragment, records their position index in the sentence segment, and further constructs a word position mapping relationship matrix between the input keywords and the response keywords to ensure the degree of position matching of the same keyword in the input and response sentences. At the same time, it retains the list of contextual word terms in the sentence segment for each term. For example, if the fragment in the input is "cancel application", and the word position corresponding to the keyword "cancel" in the response sentence "the request has not been canceled" is the 4th, then its response word position index is recorded as 4, forming a mapping pair with the input word position index. The above operation is repeated until all fragments have completed the construction of the mapping table, and finally the semantic position mapping value is output.

[0030] S202: Based on the corresponding word index in the semantic position mapping value, extract the context chunks of the focus words before and after the response sentence, obtain the direction of the order change of the focus words in the sentence sequence, and extract the progressive, turning or causal structure by combining the main clause attachment structure marker in the context chunk. Calculate and obtain the generated semantic focus offset value by using the angle of semantic migration direction between sentences and the frequency of structural turning marker. The specific formula for obtaining the semantic focus offset value is as follows: ; in, Indicates the semantic focus offset value. The normalized value representing the angle of semantic shift between the two focus words. This represents the normalized value indicating the frequency of redirection markers appearing in the context structure. This indicates the number of candidate segments in the current response text. Indicates the first The word vector distance between the focus word of the main clause and the focus word of the previous round. Indicates the first Each focal word corresponds to the syntactic connectivity strength in the context. Indicates the first The length of the sentence segment corresponding to each focal word This is a constant offset value to prevent the denominator from being zero. Indicates the number of focus words in the current sentence; Regarding the semantic focus offset value (ΔP) Meaning: ΔP reflects the overall deviation between the current response text and the semantic focus of the previous round in terms of semantic shift direction, logical connection density, and semantic alignment of focus words. The larger the value, the more the response content deviates from the intended direction of the dialogue in terms of semantic logic.

[0031] Purpose: Used to determine whether the current response has deviated from the topic focus, in order to help determine whether the conversation path needs to be restructured.

[0032] All participating items have been normalized or are dimensionless variables to avoid the problem of inconsistent dimensions; The overall expression reflects the comparison between "transition direction × connection density × response complexity" and "the degree of matching between focus words"; The result ΔP is essentially a measure of the semantic matching deviation strength, and does not require specific physical units.

[0033] Based on the response word index information of all keywords in the semantic position mapping value, the contextual blocks of 5 words before and after each group of keywords in the response sentence segment are extracted to capture the semantic context of the segment where the focus word is located. The syntactic structure information of the segment where the focus word is located is deconstructed and marked, including whether it is attached to the main clause structure, whether it participates in the construction of subordinate clauses, and whether it is the starting item of the adversative logical chain. The attachment position and connection structure identifier of the main clause are further recorded. Then, the occurrence order of the focus word in the current sentence segment in the previous and next round sequences is identified to determine its semantic migration path in the response language output. If the difference between the word positions before and after exceeds 3 and the function of its sentence structure changes, it is judged as a semantic migration mutation. The angle of its migration path is calculated, that is, the cosine value of the angle between the current focus word vector and the corresponding focus word vector in the previous round is defined as... For example, when the angle between two vectors is 45 degrees, we get Simultaneously, the frequency of all connection markers used for semantic shifting in the response sentence is counted, such as the number of occurrences of phrases like "but" and "however." If a marker appears 3 times in a response with a total of 5 sentence segments, its frequency is [missing value]. This frequency value is set to .

[0034] Next, the number of candidate segments in the response text is defined as... Then calculate For each focal word in the main clause, calculate the Euclidean distance between its word vector and the vectors of the previous focal words. An example is the first... The vector distance between the focus words is Its syntactic connectivity strength (i.e., dependency hierarchy weight) is The length of the sentence segment is To prevent the denominator from being 0, set Then calculate in sequence: ; Continue calculating the first Key words: , , ,but: ; Summing yields: ; Substitute it into the formula again: ; Therefore, the semantic focus offset value calculated in this round of response is 0.479.

[0035] The parameters are explained as follows: Indicates the semantic focus offset value. The normalized value representing the angle of semantic shift between the two focal words reflects the trend of semantic orientation change. This represents the normalized frequency value of semantic conjunctions appearing in the context. This indicates the number of candidate segments in the response text. Indicates the first The semantic distance between the current focus word and the previous focus word. Indicates the first The syntactic dependency strength of a word Indicates the first The number of words in the sentence segment containing each focal word. It is a constant, let's set it to 1. It represents the number of terms involved in semantic focus comparison.

[0036] The logic of formula operators is as follows: The first half of the parentheses It reflects a comprehensive expression of "directional shift × semantic connection strength × sentence density"; the second half is the total semantic shift weight of multiple focal terms in the structure; the absolute value of the difference between the two parts is used to measure the absolute strength of the overall shift risk value.

[0037] The advantage of the formula is that by combining terms such as angle, syntactic strength and structural length, it not only considers changes in semantic direction, but also introduces an assessment of grammatical structure stability, which can effectively improve the system's offset sensitivity and structural fault tolerance assessment capabilities in focus shift detection.

[0038] The system sets the topic offset boundary range to be an interval. The results were obtained through statistical analysis of 4000 manually annotated question-and-answer pairs in the training set, with a mean of 0.38 and a standard deviation of 0.09, as set according to the rules. , The offset interval is [0.245, 0.515].

[0039] The result indicates that the current semantic focus offset value is 0.479, which is within the set offset range but close to the upper limit, suggesting a potential risk of marginal shift in the current response focus. This value will be used to further filter and determine whether the response content needs backtracking correction. In the next step, it will be used in conjunction with the context alignment score as an initial screening criterion for content removal and candidate addition, directly affecting the determination and annotation structure of subsequent response statement topic deviation indicators.

[0040] S203: Call the semantic focus offset value, compare it with the set topic offset boundary range, filter the position index of the sentence segment whose offset value exceeds the boundary, and perform aggregate annotation on the offset index to obtain the coverage ratio of the topic focus off track segment in each response, and generate the topic off track indicator item of the response statement.

[0041] The semantic focus offset value is retrieved and compared item by item with the topic offset boundary range set by the system. This boundary range is defined by the offset value tolerance band preset by the system, and is usually established based on the upper and lower limit threshold intervals of the mean offset of samples in the training corpus ± 1.5 standard deviations, denoted as . If a certain sentence or paragraph If the value exceeds the range, the segment number and word index of the sentence segment are added to the offset record table, and a sentence segment index list is created for all sentences with offsets. By statistically analyzing the proportion of these offset segments in the current candidate response, the ratio of the number of offset segments to the total number of sentences segments is calculated to form a coverage ratio value. For example, if there are 3 offset segments in 10 response sentences, it is recorded as 30%. Then, all offset segment indexes are aggregated and classified. If the length of consecutively occurring offset segments exceeds two sentences, a continuous offset chain is constructed and marked with a focus off-track segment mark. Finally, a response statement topic off-track indicator is generated.

[0042] Please see Figure 4 The specific steps for obtaining semantically stable fragment aggregation groups are as follows: S301: Based on the action statement groups marked as having clear logical goals in the first five rounds of user input, extract the target phrase items in the word order, count the frequency of occurrence in each round of input, calculate the proportion of frequency in all statements, filter phrase content with frequency exceeding the set semantic representativeness threshold, and generate the target phrase distribution rate. To obtain action statement groups marked with clear logical objectives from the first five rounds of user input, the complete sentence content of each round of input needs to be read sequentially. Based on the configured grammar recognition rules, the parts of the sentence containing the "verb + object" structure are extracted. During this process, action target combinations such as "extract semantic structure" and "generate offset labels" need to be extracted sequentially. The text positions of target phrases are recorded through forward word order traversal. Then, the sequence positions of all target phrases in each round of input are recorded in the target index pool, forming a matrix structure corresponding to rounds and targets. In the example, if the phrase "generate offset labels" appears 2 times in round 1, 1 time in round 3, and 1 time in round 5, its frequency in the total 20 statements is 4 / 20 = 0.2. Next, the semantic representativeness threshold is set to 0.15. This value is calculated based on the median and standard deviation of the frequency distribution of label generation in the training set. The logic is that when the average frequency is 0.12, the corresponding standard deviation is 0.02, so the upper limit is taken. Using a representative threshold, occasional target phrases can be effectively eliminated. All phrases are sorted from highest to lowest frequency, and the proportion of each phrase appearing in all sentences is calculated. Only target phrases with a proportion greater than 0.15 are retained. The resulting list of selected phrases, such as "extract keywords," "analyze semantic chains," and "generate semantic tags," represents the target phrase distribution rate, primarily used for quantifying the semantic importance of phrases in subsequent structural clustering.

[0043] S302: Based on the distribution rate of the target phrases, extract sentence structure fragments that match the set of phrases with the highest frequency, determine whether the predicate type and word order logic direction in the structure fragments are consistent, accumulate the count of sentence templates that meet the condition of consistent semantic progression direction and generate a structure statistics table, exclude the segment fragments including the switching of role subject or the replacement of purpose verb, and generate the number of templates with consistent semantic direction. Based on the high-frequency phrase set obtained from the target phrase distribution rate, it is necessary to traverse the first five rounds of input text, extracting sentence structure fragments that match the set sentence by sentence. Specifically, each sentence needs to be divided into predicate verb and target phrase positions according to syntactic structure rules, and its complete overlap with the target set needs to be verified. Simultaneously, it is determined whether the predicate type in the sentence is consistent with the target word. For example, if the target phrase is "analyze semantic chain," then when the matching predicate is "analyze," it is retained. Furthermore, logical connectors before and after the sentence are used to determine whether it belongs to a progressive structure. The appearance of words such as "secondly," "furthermore," and "subsequently" can be considered as indicators of consistent word order logic direction. For each structural template that meets the condition of progressive predicate logic direction, a cumulative count is performed, ultimately resulting in a total of 42 sets of structural templates that meet the conditions. A template structure statistics table is then established. During this process, it is necessary to detect the switching of the subject role in the sentence structure. If the subject changes from "system" to "user," or the predicate verb changes from "judgment" to "generation," it is considered that the semantic role has changed or the purpose has shifted. Such sentence structures should be removed from the statistics. The final retained structural templates constitute a set of templates with consistent semantic direction. Their structural stability and logical consistency provide a sentence structure basis for subsequent semantic segmentation.

[0044] S303: Call the number of templates with consistent semantic direction, filter the continuous template coverage segments according to the start and end position index of the corresponding segment for each template, combine the semantic cache information retained by the context window of the large language model, determine the number of residual segments in the window and extract the sentence fragments whose occurrence frequency is greater than the retained benchmark frequency, construct the sentence blocks with similar semantic direction in the turn, and generate semantically stable segment aggregation groups.

[0045] All structural template fragments retained from the semantically consistent template set are retrieved, and their start and end points are extracted according to their corresponding position indices in the original input. A template coverage index list is constructed, and the regions where consecutive sentence templates appear are segmented. For example, if sentences 2 to 4 are all template fragments, their start and end indices are [2, 4]. Subsequently, the context window retention mechanism of the large language model is used to count the number of times each sentence appears in the context. The context window capacity is set to the first 50 words, and the baseline frequency is retained at 2 times. This baseline value comes from the minimum repeated feedback required for the model to generate a stable response during dialogue generation. When the frequency of a sentence fragment is greater than 2, it indicates that it has a stable structural tendency and content focus in the user's semantic expression, and therefore can be included in the semantically stable structure analysis as a semantically convergent block. Finally, these high-frequency, sentence-stable, and logically convergent sentence sets are combined into sentence blocks with semantically convergent direction in the turn, and the semantically stable fragment aggregation group is finally output as the structural input source for subsequent focus recognition, offset analysis, and other modules.

[0046] Please see Figure 5The specific steps for obtaining the focus offset node set of the response content are as follows: S401: Based on semantically stable fragment aggregation groups, extract action verbs, result target words and context modifiers included in the segment, filter the repetition frequency of words in each segment, record the position information of high-frequency words in the original segment, and mark the words with the greatest expressive strength at the beginning and end of the sentence to generate a semantic core word group comparison table. Based on semantically stable segment aggregation groups, structural expression units in each segment are analyzed sequentially to extract verbs, objects, and modifying phrases. Action verbs must have actual instructional meaning, such as "extract," "judge," and "generate." Target words must have semantic carrying functions, such as "semantic focus," "offset node," and "keyword combination." Modifiers mainly cover time, condition, and sequence phrases. The frequency of each extracted word in each segment is counted, and phrases with repetition exceeding a set frequency threshold are selected. The high-frequency phrase selection threshold is set to 3 times, derived from the standard median of sentence repetition rate within the aggregated segment. If an "offset node" appears 5 times in six segments, it is retained. When the original occurrence position of a high-frequency phrase is located at the beginning or end of a sentence, its expressive strength is marked accordingly. The expressive importance of the phrase is determined by calculating the TF-IDF value of the phrase containing the word and the subject-verb position factor in the sentence. An example of expressive strength calculation is as follows: If the word "generate" appears at the beginning of a sentence, the TF-IDF of the sentence is 0.47, and the subject position factor is 1.0, then the expressive strength is... These strong expressive terms are recorded at the beginning and end of the sentence respectively, and a sentence beginning-to-end lookup table is constructed. Finally, a semantic focus phrase lookup table is formed, which is used as the basis for subsequent focus alignment and offset value judgment.

[0047] S402: Based on the semantic core word group comparison table, identify the key words at the beginning and end of the current response candidate sentences, detect the vector distance, semantic direction matching degree and content structure overlap ratio between the extracted target semantic core word groups, calculate the deviation strength between each group of sentences and the semantic core, and combine the set semantic core distribution interval to determine whether it exceeds the semantic boundary, and generate semantic deviance offset information. The specific formula for calculating the deviation strength between each group of sentences and the semantic focus is as follows: ; in, The semantic decoupling offset value is a dimensionless measure of the degree of semantic deviation of the focus word in a response segment. For the first The normalized value of the word vector of each target semantic core word represents the dimensionless vector component of its directional position in the word embedding space. For the first The normalized word vector values ​​of each response focus word have the same meaning as the previous item and are derived from the word embedding output of the large language model. For the first The normalized value of the semantic direction matching factor for each semantic term represents the standardized output of the cosine of the angle between two vectors, used to measure the degree of semantic direction similarity. For the first The normalized word density value of a segment containing a word is derived from the ratio of the number of effective words per unit length to the average word density of the entire sentence. For the first The normalized value of the semantic diffusion range of each focal word in the response sentence is obtained by normalizing the maximum inter-word semantic distance. For the current candidate sentence segment The mean, also a dimensionless quantity, To align the number of semantic pairs, In response to the number of focus words; Regarding the semantic decoupling offset value (δ): Meaning: δ describes the difference in semantic deviation strength and diffusion characteristics between the semantic features of the focal word in the current sentence segment and the historical semantic focus phrase. This indicator quantifies whether the candidate response sentence deviates from the semantic main line.

[0048] Purpose: Used to identify candidate sentence segments that have lost focus semantically, and is the basis for determining focus offset nodes.

[0049] The first term is the weighted offset value of semantic direction, position and density between focus pairs; the second term is the standard deviation of the semantic diffusion degree of the response focus word in the sentence segment, which measures the degree of semantic center concentration; all terms are normalized quantization results to ensure the consistency of units in addition, subtraction, multiplication and division operations.

[0050] Based on the established semantic focus word pair lookup table, the initial and final words of the current response candidate sentence are identified. Word embedding vectors are extracted, and their differences from the semantic focus words are calculated. A direction matching degree judgment is then performed on each pair of words. The calculation method requires calling... and The directional matching factor is calculated by representing the directional positions of the target word and the response word in the vector space and using the cosine of the angle between the two vectors. In the example, if the included angle is 30°, the matching factor is... The data is then processed and normalized. Further, the term density per unit length in the current response sentence is obtained. If a sentence has 20 words and a length of 10, then the density is 2. Assuming the average word density is 1.5, the normalized value is... Subsequently, the semantic diffusion range of the current response focus word is extracted, defined as the maximum semantic vector distance from each focus word to all other words. For example, if the maximum semantic distance of a focus word is 0.84, after normalization... With all The mean is Construct the formula: ; The operational logic is explained as follows: The first term is the sum of the absolute values ​​of the directional differences between each target-response word pair multiplied by the matching degree and divided by the word density penalty factor, used to measure the structural offset; the second term is the standard deviation of the semantic dispersion of the focal words in the sentence, used to reflect the semantic clustering; overall... This represents the offset strength value; a larger value indicates a higher degree of semantic decoupling. In the example, it is set to... ,like: - - - -

[0051] Then the first item: ; Item 2: ; Therefore ; Based on the set semantic boundary threshold of 0.18, it can be determined that the sentence segment exhibits semantic shift. The advantage of this shift value lies in its ability to differentiate the strength of semantic focus by jointly evaluating it using word vector difference terms, direction matching factors, and syntactic density factors. It also clearly identifies the semantic decoupling location, avoiding ambiguity. This result indicates a focus shift between the response sentence and the user input, necessitating the removal or correction of the content structure. The shift strength value provides a quantitative basis for subsequent structural adjustments.

[0052] S403: Based on semantic decoupling offset information, according to the identified sentence group numbers whose offsets exceed the distribution threshold, summarize the candidate sentence segment index and offset frequency of semantic decoupling, count the proportion of offset nodes appearing in each round of candidate responses, establish the position information and matching value table of offset nodes, obtain continuous decoupling blocks in sentence groups, and generate a set of focus offset nodes for response content.

[0053] Based on the semantic decoupling offset values ​​generated from all candidate statements, an upper limit of 0.18 is set for the offset distribution. All candidate sentence group numbers and their corresponding offset values ​​are then screened sequentially. Sentence segment numbers with offset values ​​greater than this value are extracted and aggregated at the current response position. If the response contains 8 sentences, and the 3rd, 5th, and 6th sentences all have offset values ​​higher than the set threshold, they are considered decoupling segments with an offset frequency of 3, representing a percentage of [missing information]. Then, an index matching table for the offset segments is created, recording information such as the number, value, and location. It also checks for consecutive offset segments; for example, if sentences 5 and 6 are consecutive offsets, they are merged into one group. The final output offset structure is as follows. The output response content focus offset node set is combined with this information and used as input for subsequent content clipping and focus restoration strategies.

[0054] Please see Figure 6 The specific steps for obtaining the response output text after semantic path control are as follows: S501: Based on the topic derailment indicator of the response statement and the focus offset node set of the response content, extract the semantic span value and focus word position index of the corresponding sentence segment respectively, calculate the matching degree score between the keyword span range and the overlapping position of the context focus in each sentence, construct the content validity ranking benchmark of the statement in the context, establish the sentence segment ranking matrix, and generate content validity ranking information. Based on the topic derailment indicator and the focus offset node set of the response content, the marked sentence segment number and the corresponding focus word index position are extracted. All sentence segments are traversed sequentially, and keyword combinations for each sentence are extracted. The semantic span value is calculated by combining the word order distance between the verb and target word in the combination within the sentence. In the example sentence segment "identify semantic derailment and mark it", the verb "identify" and the target word "mark" are separated by 2 word terms, so the semantic span value is 2, which is represented as 0.2 after normalization. The index position of the corresponding focus word in the context of the keyword group is obtained and alignment is performed. Successful alignment is defined as the difference between the index number of the context focus word and the index of the current keyword position not exceeding 3 positions. If "mark" appeared at position 12 in the previous round and is currently at position 14, with a difference of 2, then a valid alignment is considered to exist. The keyword span value of each sentence is normalized. Matching ratio with context focus The content validity ranking baseline score is obtained by weighting the data, using the following weighting formula: ; in For the first The validity score of the sentence / paragraph, ranging from... The closer the score is to 1, the stronger its ability to align with the context. In the example, if the span value of a certain sentence segment is 0.3 and the matching ratio is 0.75, then the validity score is... The calculation results for all sentence segments Construct a sentence segment sorting matrix, sort the segments by value from largest to smallest, and output the content validity sorting value for subsequent sentence segment filtering and supplementation processing.

[0055] S502: Call the content validity ranking information, extract the original index number and semantic anchor position based on the sentence and segment content whose score is lower than the lower limit of the set context alignment score standard, retrieve the corresponding content block in the previous generation cache, compare the missing focus words in the current content with the anchor phrases in the previous round, extract the words before and after the missing keywords and build a supplementary index table, and generate semantic anchor supplementary parameter groups. Based on the generated content effectiveness ranking values, all segments with scores below the set lower limit of the context alignment scoring standard are filtered out. The lower limit threshold is set at 0.55, which is the weighted midpoint of the sum of the differences between the median and the first percentile of the context content information coverage in empirical statistics. All segments are extracted sequentially. The system retrieves the sentence segment number and the original position of its keywords. It then calls upon anchor phrases from the previous generation cache for the same numbered segment and compares them to the current segment to check for missing phrases. If keywords such as "derailed" or "connected" are missing, they are recorded as missing. The system obtains the two word blocks before and after the missing keyword as the context for completion, creating a completion index table. Each entry includes the missing keyword, the word blocks before and after the missing keyword, and its position index in the cache. For example, if "complete" is missing at position 9, and the preceding word is "semantic" and the following word is "term," then the completion index is 9, and the word block is "semantic...term." The final output is a semantic anchor completion parameter set, serving as the basis for paragraph structure replacement.

[0056] S503: Based on the semantic anchor point completion parameter group, according to the position index value and term structure content, locate the boundary range of the sentence segment to be deleted in the current response segment, and combine the semantic anchor point cache content of the previous round to perform paragraph-level deletion and structural replacement, semantically fill the focus word group towards the original anchor word group, and generate the response output text after semantic path control.

[0057] Based on the term structure in the semantic anchor completion parameter group and the original sentence segment number, the boundary range of each sentence segment to be processed is located. All sentence segments to be replaced are aggregated and merged in index order. The segment containing the semantic anchor term group in the previous round of cache is called, and paragraph-level structural removal and replacement processing is performed. Taking the segment "judgment offset" missing the "focus" word as an example, before replacement, the block sequence positions of "judgment" and "offset" in the current sentence are first confirmed. Then, with "focus judgment offset" in the original cache segment as a reference, intra-segment structure rewriting is performed, "judgment offset" is reconstructed into "focus judgment offset", and replaced to the corresponding sentence segment position in the current response. In this process, the structural replacement is limited to the beginning and end positions of the current sentence segment to avoid cross-segment interference. This process is repeated until all completion index items are executed, and the updated response statement is output. Finally, the response output text after semantic path control is generated to provide feedback on the context consistency of the generated content of the large model.

[0058] Please see Figure 7 A large-model-based intelligent chat system is used to execute the aforementioned large-model-based intelligent chat method. The system includes: The input coherence analysis module identifies semantic transition regions with consistent targets in adjacent structures based on user input statements in the current session, determines whether the frequency and word order offset length exceed the set turn continuity threshold, calls the large language model structured semantic annotation interface to generate semantic segmentation boundaries, and generates a set of user input semantic coherence chain tags. The topic derailment identification module calls the user input semantic coherence chain tag set, measures the offset distance according to the semantic focus shift direction and the logical progression relationship between sentences, compares the measured value with the preset topic offset boundary range, and generates a response statement topic derailment indicator. The semantic stability analysis module counts the number of sentence templates with consistent semantic direction based on the action statement groups marked with clear logical goals in the first five rounds of user input, and filters out the segments that trigger role changes or goal replacements, constructs language blocks with consistent direction in the dialogue context, and generates semantically stable fragment aggregation groups. The focus shift recognition module is based on semantically stable fragment aggregation groups, combined with the attention concentration position information generated in the response by the large language model, identifies candidate sentence groups whose deviation range is greater than the semantic center distribution threshold, and determines them as semantic focus shift nodes, generating a set of focus shift nodes in the response content. The backtracking and supplementing output module identifies sentence segments in the ranking results whose scores are lower than the lower limit of the context alignment score standard based on the response statement topic off track indicator and the response content focus offset node set. It calls the large language model to generate the previous round of content reconstruction fragments in the cache, performs paragraph-level deletion and semantic anchor backtracking and supplementing on the marked fragments, and generates the response output text after semantic path control.

[0059] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.

[0060] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0061] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0062] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0063] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0064] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0065] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0066] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0067] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0068] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A large model-based intelligent chat method, characterized in that, Comprise the following steps: S1: According to the user input statement in the current session, identify the semantic transition area of the target trend consistency in the adjacent structure, judge whether the frequency and the length of the syntax sequence offset exceed the set turn continuity critical value, call the large language model structured semantic labeling interface to generate the semantic segmentation boundary, and generate the user input semantic coherent chain label set; S2: Call the user input semantic coherent chain label set, measure the offset distance according to the semantic focus shift direction and the inter-sentence logical progression relationship, compare the measurement value with the preset topic offset boundary range, and generate the response sentence topic derailment indication item; S3: According to the first five rounds of user input marked as logical target clear action statement group, count the number of sentence templates with consistent semantic direction, and exclude the sentence segments that trigger role change or target replacement, construct the language block with consistent direction in the dialogue context, and generate the semantic stable segment aggregation group; S4: Based on the semantic stable segment aggregation group, combine the attention concentration position information in the response generated by the large language model, identify the candidate sentence group whose deviation range is greater than the semantic gravity distribution threshold, and determine it as the semantic focus shift node, and generate the response content focus shift node set.

2. The large model-based intelligent chat method according to claim 1, characterized in that, The user input semantic coherent chain label set includes semantic structure stability mark, syntax logical break point and semantic transition unit boundary, the response sentence topic derailment indication item is specifically focus word offset value, logical progression direction deviation rate and candidate sentence segment position index, the semantic stable segment aggregation group includes target word frequency set, semantic direction consistent template set and context window segment index table, and the response content focus shift node set is specifically semantic gravity alignment matrix, sentence segment number deviating from the focus and attention value distribution aggregation statistics.

3. The large model-based intelligent chat method according to claim 2, characterized in that, The acquisition step of the user input semantic coherent chain label set is specifically: S101: According to the user input statement in the current session, extract the action verb combination, limited target word group and logical connection mark item in the continuous two rounds of input sentence, arrange and construct the original semantic structure sequence according to the syntax order, mark according to the part of speech and grammatical role of the semantic segment in the sentence structure, and generate the semantic instruction construction sequence; S102: Call the semantic instruction construction sequence, identify whether the target trend is consistent between adjacent combinations according to each group of action target relationship, calculate the position difference value and frequency of appearing logical jump, compare the difference value result with the set semantic continuity critical interval, obtain the structure paragraph exceeding the interval and mark the break position, and generate the target consistency offset label group; S103: Based on the target consistency offset label group, according to the break position index, call the large language model structured semantic labeling interface to identify the subject, predicate, object, time adverbial and conditional restriction word group in the input sentence, combine the break boundary information in the label group to perform semantic level segmentation, and generate the user input semantic coherent chain label set.

4. The large model-based intelligent chat method according to claim 3, characterized in that, The acquisition step of the response sentence topic derailment indication item is specifically: S201: Call the user input semantic coherence chain label set, according to the semantic unit marked as trend fracture, match the candidate response sentence generated by the current large language model one by one, extract the word position index sequence of the key words in the response sentence, establish the word position mapping table of input and response semantic unit, and generate semantic position mapping value; S202: According to the corresponding word position index in the semantic position mapping value, extract the context block of the front and rear focus words in the response sentence, obtain the sequence change direction of the focus words in the sentence, and extract the progressive, turning or causal structure combined with the main sentence dependent structure mark in the context block, calculate the semantic focus shift value by using the inter-sentence semantic migration direction angle and structure turning mark frequency; S203: Call the semantic focus shift value, compare with the set topic shift boundary range, filter the sentence segment position index whose shift value exceeds the boundary, aggregate the shift index, obtain the coverage ratio of the topic focus derailment segment in each response, and generate the response sentence topic derailment indication item.

5. The large model-based intelligent chat method according to claim 4, characterized in that, The formula for calculating the semantic focus shift value is ; wherein, denotes a semantic focus shift value, denotes a normalized value of the semantic shift angle of the two focus words, denotes a normalized value of the frequency of the turn-taking marker in the context structure, denotes the number of candidate sentence segments in the current response text, denotes the word vector distance between the th main sentence focus word and the previous focus word, denotes the syntactic connection strength in the context corresponding to the th focus word, denotes the length of the sentence segment corresponding to the th focus word, is a constant offset value, denotes the number of focus words in the current sentence.

6. The large model-based intelligent chat method according to claim 5, characterized in that, The acquisition step of the semantic stable segment aggregation group is: S301: According to the action sentence group marked as logical target clear in the first five rounds of user input, extract the target phrase item in the sequence, count the frequency of each round of input, calculate the proportion of the frequency in all sentences, filter the phrase content with frequency exceeding the set semantic representative threshold, and generate the target phrase distribution rate; S302: Based on the target phrase distribution rate, according to the phrase set ranked by frequency, extract the sentence structure segment matched with the set, judge whether the predicate type and the sequence logical direction in the structure segment are consistent, count and generate the structure statistical table for the sentence pattern that meets the consistent semantic progression direction condition, exclude the sentence segment including role subject switching or purpose verb replacement, and generate the number of semantic direction consistent templates; S303: Call the number of semantic direction consistent templates, according to the start and end position index of each template corresponding segment, filter the continuous template coverage segment, judge the residual times of the segment in the context window reserved by the large language model according to the semantic cache information, extract the sentence pattern segment with appearance frequency greater than the reserved reference frequency, construct the sentence group block with consistent semantic direction in the turn, and generate the semantic stable segment aggregation group.

7. The large model-based intelligent chat method according to claim 6, characterized in that, The acquisition step of the response content focus shift node set is: S401: Based on the semantic stable segment aggregation group, extract the action verbs, result target words and scene modifier words included in the sentence segment, filter the repetition frequency of the word items in each sentence segment, record the position information of the high-frequency word group in the original sentence segment, mark the expression intensity maximum word item at the beginning and end of the sentence, and generate the semantic gravity word group comparison table; S402: According to the semantic focus word group table, the key expression word items at the beginning and end of the current response candidate sentence are identified, the vector distance, semantic direction matching degree and content structure overlap ratio between the extracted target semantic focus word group and the semantic focus are detected, the deviation intensity between each group of sentences and the semantic focus is calculated, whether the deviation exceeds the semantic boundary is judged by combining the set semantic focus distribution interval, and semantic deviation offset information is generated; The formula for calculating the deviation intensity between each group of sentences and the semantic focus is specifically: ; wherein, is a semantic disengagement offset value, is a word vector normalization value of the th target semantic focus word, is a word vector normalization value of the th response focus word, is a semantic direction matching factor normalization value of the th semantic term, is a term density normalization value of the th sentence segment where the term is located, is a semantic diffusion range normalization value of the th focus word in the response sentence, is the mean value of the th candidate sentence segment, is the number of aligned semantic pairs, is the number of response focus words; S403: Based on the semantic deviation offset information, the candidate sentence segment index and offset frequency of the semantic deviation are summarized according to the identified sentence group number whose deviation exceeds the distribution threshold, the proportion of the offset node appearing in each round of candidate response is counted, the position information and matching value table of the offset node are established, the continuous deviation block in the sentence group is obtained, and the response content focus offset node set is generated.

8. The large model-based intelligent chat method according to claim 7, characterized in that, The method further comprises the following steps: S5: Based on the response sentence question derailment indicator and the response content focus offset node set, the sentence segment content with a score lower than the lower limit of the context alignment score standard in the sorting result is identified, the previous round of content reconstruction segment in the large language model generation cache is called, the marked segment content is deleted at the paragraph level and the semantic anchor point is backfilled, and the response output text after semantic path control is generated; The response output text after semantic path control includes control paragraph text segment, semantic anchor point replacement position table and focus backfill word item.

9. The large model-based intelligent chat method according to claim 8, characterized in that, The acquisition step of the response output text after semantic path control is specifically: S501: Based on the response sentence question derailment indicator and the response content focus offset node set, the semantic span value and focus word pair position index of the corresponding sentence segment are extracted respectively, the matching degree score of the key word span interval in each sentence and the context focus overlap position is calculated, the content validity sorting benchmark of the sentence in the context is constructed, the sentence segment sorting matrix is established, and the content validity sorting information is generated; S502: The content validity sorting information is called, the original index number and semantic anchor position are extracted according to the sentence segment content with a score lower than the lower limit of the set context alignment score standard, the corresponding content block is retrieved in the previous round of generation cache, the focus word missing item in the current content is compared with the previous round of anchor word group, the front and rear text blocks of the missing key word are extracted and the backfill index table is established, and the semantic anchor point backfill parameter group is generated; S503: Based on the semantic anchor point backfill parameter group, the position index value and word item structure content are used to locate the sentence segment boundary range to be deleted in the current response segment, and the paragraph level deletion and structure replacement are performed in combination with the previous round of semantic anchor cache content, the focus word group is filled in the direction of the original anchor word group, and the response output text after semantic path control is generated.

10. A large model-based intelligent chat system, characterized by, The system is used to implement the intelligent chat method based on a large model according to any one of claims 1-9, and the system comprises: The input coherence analysis module identifies the semantic transition area of the target trend consistency in the adjacent structure according to the user input sentence in the current session, judges whether the frequency and the sequence offset length exceed the set turn continuity critical value, calls the large language model structured semantic labeling interface to generate the semantic segmentation boundary, and generates the user input semantic coherence chain label set; The topic derailment recognition module calls the user input semantic coherence chain label set, measures the offset distance according to the semantic focus shift direction and the inter-sentence logical progression relationship, compares the measured value with the preset topic offset boundary range, and generates the response sentence topic derailment indication item; The semantic stability analysis module counts the number of sentence templates with consistent semantic direction according to the action sentence groups marked as clear logical targets in the first five turns of user input, filters out the segments that trigger role changes or target replacements, constructs the language blocks with consistent direction in the dialogue context, and generates the semantic stability segment aggregation group; The focus offset recognition module generates the attention concentration position information in the response based on the semantic stability segment aggregation group and the large language model, identifies the candidate sentence group whose deviation range is greater than the semantic barycenter distribution threshold, and determines it as the semantic focus offset node to generate the response content focus offset node set; The backfill output module identifies the sentence content with a score lower than the lower limit of the context alignment score standard in the sorting result based on the response sentence topic derailment indication item and the response content focus offset node set, calls the large language model to generate the previous round of content reconstruction segment in the cache, deletes the marked segment content at the paragraph level and performs semantic anchor backtracking and supplement, and generates the response output text after semantic path control.

Citation Information

Patent Citations

  • Method and system for processing data based on large language model

    CN120471023A

  • Artificial intelligence interaction robot

    CN121071110A

  • Intelligent question and answer matching method based on machine learning

    CN121328728A