Method and system for processing data based on large language model

Through inter-sentence punctuation and part-of-speech annotation, a semantic chain triggers the index sequence and label path breakpoints are solved, and the problem of overlapping and misalignment of label paths in the prior art is realized, and efficient conversion and accurate mapping of text unstructured to structure is realized.

CN120471023AActive Publication Date: 2025-08-12BEIJING SHENZHOU BANGBANG TECH SERVICE CO LTD

Patent Information

Application Number
CN202510976080.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-08-12
Estimated Expiration
2045-07-16

AI Technical Summary

Technical Problem

When processing text-type unstructured data, the prior art lacks an explicit judgment mechanism based on semantic content, resulting in the label path being easily overlapped, misaligned or omitted, especially in long text or multi-chain semantic structures, which affects the accuracy of downstream semantic deconstruction and reasoning operations.

Method used

By identifying and dividing sentence blocks based on inter-sentence punctuation positioning, part-of-speech annotation and subject verb structure recognition, segment function partitioning is generated, semantic chain trigger index sequences of topics and relational words are extracted, logical jump points and breakpoints are identified, label path nested hierarchical tables are generated, and mapping corrections are performed in large language models.

Benefits of technology

It enhances the accuracy, coherence and hierarchical clarity of semantic reconstruction, ensures the integrity and correct mapping of label paths, and improves the structured expression quality of language model output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471023A_ABST
    Figure CN120471023A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a method and system for processing data based on a large language model.The method comprises the following steps that based on inter-sentence punctuation positioning and part-of-speech tagging, sentence blocks are divided to generate functional partitions, themes and relational words are extracted to judge semantic chain starting points, a trigger index sequence is constructed, and logic jump points and breakpoint positions are recognized; and mapping the label structure to the language model output analysis deviation, and generating a label mapping combination list. According to the method, semantic turning nodes can be captured by analyzing the semantic direction change trend, the relation between the semantic turning nodes and verb and noun combinations is judged, the position of a trigger point of an actual information transfer effect is extracted, and the break point area of a semantic path is recognized through the positioning of key word starting and stopping blocks and the logical judgment of noun group cross combination; the path integrity has a clear fracture identifier, a traceable semantic mapping structure path is established in a language model, and the accuracy, coherence and hierarchy clearness of semantic reconstruction are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a method and system for processing data based on a large language model. Background Art

[0002] The field of data processing technology encompasses systematic methods and tools for collecting, transforming, analyzing, modeling, and structuring various types of data. The core of this field is to improve the efficiency of data organization and processing to support information extraction, decision-making, and intelligent application implementation. Data processing broadly encompasses aspects such as cleaning, feature extraction, data aggregation, classification indexing, and semantic analysis of both structured and unstructured data. Common applications include database management, image recognition, and predictive modeling. As data continues to expand and diversify, data processing technology is increasingly incorporating self-learning and reasoning mechanisms to enhance its processing logic and adaptive capabilities, playing a greater role in intelligent systems.

[0003] Among them, the method of processing data with a large language model refers to a method of using a large-scale language model based on a neural network to perform semantic modeling, structural reconstruction, and contextual reasoning on text data. The subject of this patent mainly focuses on the structural transformation, semantic association recognition, and knowledge expression problems of text-type unstructured data. It uses a pre-trained language model to generate context embedding vectors, and then determines the dependency between semantic fragments through vector similarity measurement. It also combines a rule generation mechanism to parse natural language text into content labels and entity relationship mappings that meet specific structural definitions to complete data structure expression. The method includes key steps such as large-scale corpus pre-training, semantic embedding generation, entity relationship extraction, and structural mapping assembly.

[0004] In the existing technology, when faced with the task of structuring unstructured text data, although the dependency relationship identification of semantic segments is completed through vector similarity measurement, the overall recognition logic is highly dependent on the context vector output by the language model, and lacks an explicit judgment mechanism based on changes in the semantic content itself. Since structural judgment is not performed on the semantic transition or logical interruption position in the text, label paths are prone to overlap, misalignment or omission in the actual mapping. When dealing with nested labels or multi-label parallel paths, existing methods use vector density similarity for rough merging, lacking a clear definition of hierarchical boundaries, resulting in disordered label distribution in the output layer of the structural mapping, which is prone to label attribution errors, especially in long texts or multi-chain semantic structures. For example, in policy texts or contract documents, if there is a lack of chain starting position identification and path jump breakpoint annotation capabilities, it is easy to mix labels from multiple logical paragraphs and output them, affecting the accuracy of downstream semantic deconstruction and reasoning operations. Summary of the Invention

[0005] In order to solve the problem in the prior art that the overall recognition logic is highly dependent on the context vector output by the language model, and lacks an explicit judgment mechanism based on changes in the semantic content itself. Since no structural judgment is made for the semantic transition or logical interruption position in the text, the label path is prone to overlap, dislocation or omission in the actual mapping. When dealing with nested labels or multi-label parallel paths, the existing methods roughly merge based on vector density similarity, lack a clear definition of hierarchical boundaries, resulting in disordered distribution of labels in the output layer of the structural mapping, which is prone to label attribution errors, especially in long texts or multi-chain semantic structures. For example, in policy texts or contract documents, if there is a lack of chain starting position recognition and path jump breakpoint marking capabilities, it is easy to mix the labels in multiple logical paragraphs and output them, affecting the accuracy of downstream semantic deconstruction and reasoning operations. The present invention provides a method and system for processing data based on a large language model. The technical solution is as follows:

[0006] In one aspect, a method for processing data based on a large language model is provided, the method comprising: S1: Based on the sentences in the text paragraph, locate the punctuation marks between sentences, tag the connectives of adjacent sentences with parts of speech, identify verb and subject pairs, divide the sentence blocks according to the order of phrase matching, and generate sentence segment functional partitions; S2: Calling the segment functional partition, extracting the subject words and relation words of the text paragraph, marking the position of the phrase in the context of the block, recording the semantic trend of the phrase according to the direction of position change, identifying the direction reversal nodes in two consecutive blocks, and generating a semantic chain trigger index sequence; S3: Using the semantic chain to trigger the index sequence, locate the start and end blocks of the keywords in the label path segment, extract the key noun groups at both ends, cross-combine the two groups of words and determine whether there are three or more groups of non-co-word structures. If so, mark them as logical jump points and generate a label path breakpoint position group; S4: calling the label path breakpoint position group, extracting the label phrase from the corresponding segment, recording the position sequence of the label phrase in the sentence, identifying whether there is path interleaving or position nesting, analyzing the nesting hierarchical relationship, and generating a label structure nesting hierarchical table.

[0007] As a further solution of the present invention, the sentence segment functional partition includes semantic connection boundary points, language block structure boundary lines, and sentence trunk matching groups; the semantic chain trigger index sequence includes transition node sequence, semantic direction turning point, and language block combination starting position; the label path breakpoint position group includes label semantic jump area, keyword disjoint fragment, and path continuity breakpoint; the label structure nested hierarchical table includes label hierarchical index set, nested path number column, and label interleaved sequence diagram.

[0008] As a further solution of the present invention, the steps for obtaining the functional partition of the sentence segment are specifically as follows: S101: Based on the sentences in the text paragraph, punctuation marks between sentences are scanned, the connection relationship between the sentences is analyzed, the conjunctions after the sentence-end punctuation are extracted as analysis objects, the part-of-speech tags are identified, and whether they are conjunctions or adverbs indicating parallel, transitional, or progressive relationships is determined. The words are then marked as connection nodes to generate a connection part-of-speech tag structure set; S102: Using the connection part-of-speech tagging structure set, locate the verb and subject in the sentence corresponding to the connection node, extract the subject phrases before and after the verb position and determine whether they form a subject-predicate structure, identify subject-predicate pair combinations with related meanings in continuous sentences, record the start and end positions and semantic attribution of the subject-predicate structure in the original text, and generate a subject-predicate phrase combination mapping sequence; S103: According to the subject-predicate group combination mapping sequence, each group of semantic units is arranged according to the order of the subject-predicate combination in the original text, sentence boundaries are divided and assembled into semantic sentence blocks, and according to the number of semantic repetitions between the subject and predicate in the semantic sentence blocks, whether continuity between semantic groups is formed is determined, and sentence blocks with related subject-predicate combinations are classified into the same interval to obtain the functional partitioning of sentence segments.

[0009] As a further solution of the present invention, the steps of obtaining the semantic chain trigger index sequence are specifically as follows: S201: calling the sentence units in the sentence segment functional partition, extracting the noun groups and verb groups with the highest frequency of occurrence in the sentence block, classifying the noun groups as subject words, classifying the verb groups as relation words, and recording the starting position and ending position in the sentence block as a semantic index range to obtain a subject relation word position information set; S202: Based on the topic-related word position information set, the index ranges of the phrases are vertically compared according to the original arrangement order of the language blocks in the text, and the movement direction of the index position in the previous language block and the next language block are compared. The forward movement, backward movement, or unchanged state is marked respectively, and the original text position is recorded to obtain a semantic transition node index table; S203: Call the semantic transition node index table, analyze the sentence content corresponding to the transition node, extract the verb and noun combination segments near the node, and determine whether the node falls in the combination segment after the verb and before the noun. If the position relationship is met, mark it as the semantic chain starting point and generate a semantic chain trigger index sequence.

[0010] As a further solution of the present invention, the step of obtaining the label path breakpoint position group is specifically as follows: S301: Using the semantic chain trigger index sequence, locate the start chunk and the end chunk corresponding to the trigger index in the text, extract the noun groups with the highest frequency of occurrence in the two chunks, record the position index range and the chunk identifier in the original text, and generate a chunk key noun group set; The formula for extracting the noun groups with the highest frequency of occurrence in the two language blocks is: ; in, Score the priority of the noun group, represents the frequency of the i-th noun group in the chunk, represents the average frequency of noun groups. represents the relevance score between the i-th noun group and the context, and k represents the number of noun groups; S302: Call the noun groups of the starting language block and the ending language block in the set of key noun groups of the language block, perform cross-combinations in pairs, identify whether there is vocabulary overlap in each combination, count the number of combinations without common words, and if the number of combinations without common words reaches three or more, mark the starting and ending language block segments as logical jump points, and obtain the label path breakpoint position group.

[0011] As a further solution of the present invention, the steps for obtaining the tag structure nesting level table are specifically as follows: S401: Calling the text segment corresponding to the jump point in the tag path breakpoint position group, extracting noun phrases that appear more than twice in the text segment as candidate tag phrases, recording the first word position and the last word position of the tag phrase in the sentence, and obtaining a tag phrase position indexing table; S402: Arrange the tag phrases in order within the sentence based on the start and end position indexes of each group of phrases in the tag phrase position indexing table, compare the front and back boundary ranges to see if there is an overlapping area or nested relationship, and generate a tag path nested interleaved tag set; S403: calling the nested interleaved tag set of the tag path, progressively establishing tag level index numbers according to the nesting layers, classifying phrases with a common nesting starting point into the same level path group, and obtaining a tag structure nesting level table.

[0012] As a further solution of the present invention, the comparison of whether the front and back boundary ranges have an overlapping area or a nested relationship is performed using the formula: ; in, Represents the boundary intersection strength index between the p-th group and the q-th group of label phrases, Represents the starting position index value of the pth group of label phrases, Represents the end position index value of the pth group of label phrases, Represents the starting position index value of the qth group of label phrases, Represents the ending position index value of the qth group of label phrases, Represents the starting position index value of the rth group of label phrases, Represents the average value of the starting position index value of the first u groups of label phrases, Represents the ending position index value of the rth group of label phrases, Represents the average value of the end position index value of the first u groups of label phrases, Represents the span value of the rth group of label phrases, is the number of label pairs.

[0013] As a further solution of the present invention, the method further includes step S5: S5: Extracting the context position of the label segment from the large language model pre-training corpus based on the label structure nesting hierarchy table, mapping the label structure corresponding to the nesting hierarchy to the language model output layer, comparing the deviation of the label word distribution at the corresponding position in the mapped structure and the default output of the language model, identifying the label trunk combination order in the traceable path, and generating a language model label mapping combination list; The language model label mapping combination list includes a semantic nested mapping path, a label structure corresponding relationship set, and a word-direction distribution difference mark sequence.

[0014] As a further solution of the present invention, the steps of obtaining the language model label mapping combination list are specifically as follows: S501: Call the hierarchical index number of each label path in the nested hierarchical table of the label structure, search the context segment of the corresponding label in the original text in the pre-trained corpus of the large language model, extract the position identification range of the label context segment in the model input, and establish a context input index set of the label segment to obtain a label context mapping location set; S502: Based on the label context mapping positioning set, the corresponding label structure nesting level information is injected into the output mapping sequence of the large language model, the label word direction representations generated by the output layer at the corresponding positions before and after the injection are extracted, and the distribution trajectories between each group of label word directions are compared. The offset direction, angle and spacing are located and classified to obtain label word direction distribution deviation information; S503: Based on the label combination trajectory with continuous offset and consistent direction in the label word distribution deviation information, extract the word path that can be used for semantic backtracking, determine whether the head and tail logical vocabulary combination connecting the labels in the word path meets the original semantic chain order, filter the label paths that meet the sequential logic and arrange them into nested combinations, and obtain the language model label mapping combination list.

[0015] On the other hand, the system for processing data based on a large language model is used to perform the above-mentioned method for processing data based on a large language model, and the system includes: The chunk recognition module identifies punctuation marks within a text paragraph, detects the part of speech of the conjunctions between the punctuation marks, and determines whether the sentences adjacent to the conjunctions contain subject phrases and predicate verb structures. It then divides the entire paragraph into multiple chunks based on the order in which the phrases appear in the sentence, generating functional partitions for the segments. The semantic indexing module calls the sentence segment functional partition, extracts the subject words and relation words in the sentence block, records the position of the phrase in the context block, identifies whether the movement direction of the phrase is reversed, determines whether the reversal point is located between the verb and noun combination segment, and obtains the semantic chain trigger index sequence; The logic jump module triggers the index sequence according to the semantic chain, locates the corresponding starting and ending language blocks in the text paragraph, extracts all noun groups in the starting and ending language blocks, and arranges them according to the noun groups at both ends to generate a label path breakpoint position group; The label extraction module uses the label path breakpoint position group to extract the label phrase from the sentence block pointed to by the jump node, identifies the sentence position number of the label phrase in the sentence block, determines whether there is an interleaved position or nested arrangement, and obtains a label structure nesting level table; The label mapping module uses the nested hierarchical table of the label structure to extract the context of the paragraph that matches the label phrase in the language model, identifies the word vector position distribution of the label group in the output layer of the language model, compares the corresponding difference with the default word vector distribution, and generates a language model label mapping combination list.

[0016] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least: By identifying sentence boundaries within a text paragraph and using the structural positioning of connective part-of-speech tags and verb-subject pairs as the basis for sentence chunking, the resulting functional segmentation logically organizes semantics at the initial stage of the text, providing a structured input foundation for the precise location of the starting points of subsequent semantic chains. By leveraging the semantic direction trends reflected by the positional differences in the occurrence of phrases within the preceding and following contexts of a chunk, semantic transition nodes can be captured and their relationships with verb-noun combinations can be determined, thereby extracting trigger points that effectively transfer information. By locating keyword start-end chunks and logically analyzing the intersection of noun groups, breakpoints in the semantic path are identified, ensuring clear breakpoints in the path integrity. The nested hierarchical structure within the path is extracted by layering the labels based on their relative position within the sentence and combining the interlaced features of the path to define the logical hierarchical relationships between the labels. After mapping the nested structure to the output layer of a large language model, the path rationality is determined by identifying word-direction distribution deviations, enabling the transition from unstructured to structured text expression. A traceable semantic mapping structure path is established within the language model, enhancing the accuracy, coherence, and hierarchical clarity of semantic reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a schematic diagram of the workflow of the present invention; Figure 2 It is a system flow chart of the present invention. DETAILED DESCRIPTION

[0018] The technical solution of the present invention is described below in conjunction with the accompanying drawings.

[0019] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as an "exemplary" in the present invention should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of the word "exemplary" is intended to present concepts in a concrete manner. Furthermore, in the embodiments of the present invention, "and / or" can mean both or either of the two.

[0020] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0021] See also Figure 1 The embodiment of the present invention provides a method for processing data based on a large language model. The processing flow of the method may include the following steps: S1: Based on the sentences in the text paragraph, locate the punctuation marks between sentences, tag the connectives of adjacent sentences with parts of speech, identify verb and subject pairs, divide the sentence blocks according to the order of phrase matching, and generate sentence segment functional partitions; S2: Invoke the segment function partitioning to extract the subject words and relation words of the text paragraph, mark the position of the phrase in the context of the block, record the semantic trend of the phrase according to the direction of position change, identify the direction reversal node between two consecutive blocks, and determine whether the node is located between the semantic verb and noun combination segment. If the combination structure is met, define the node as the starting point of the semantic chain and generate the semantic chain trigger index sequence; S3: Use semantic chain trigger index sequence to locate the start and end blocks of keywords in the tag path segment, extract the key noun groups at both ends, cross-combine the two groups of words and determine whether there are three or more groups of non-co-word structures. If so, mark them as logical jump points and generate a tag path breakpoint position group. S4: Call the tag path breakpoint position group, extract the tag phrase from the corresponding segment, record the position sequence of the tag phrase in the sentence, identify whether there is path interleaving or position nesting, analyze the nesting hierarchical relationship, and generate a tag structure nesting hierarchy table; S5: Extract the contextual position of the label segment from the large language model pre-training corpus based on the label structure nesting hierarchy table, map the label structure corresponding to the nesting hierarchy to the language model output layer, compare the deviation of the label word distribution at the corresponding position in the language model default output after mapping, identify the label trunk combination order in the traceable path, and generate a language model label mapping combination list; The functional division of sentence segments includes semantic connection boundary points, language block structure boundary lines, and sentence trunk matching groups. The semantic chain trigger index sequence includes transition node sequence, semantic direction turning point, and language block combination starting position. The label path breakpoint position group includes label semantic jump area, keyword disjoint fragment, and path continuity breakpoint. The label structure nested hierarchy table includes label hierarchy index set, nested path number column, and label interleaved sequence diagram. The language model label mapping combination list includes semantic nested mapping path, label structure correspondence relationship set, and word direction distribution difference mark sequence.

[0022] The specific steps for obtaining the functional partition of a segment are as follows: S101: Based on the sentences in the text paragraph, punctuation marks between sentences are scanned, the connection relationship between the sentences is analyzed, the conjunctions after the sentence-end punctuation are extracted as analysis objects, the part-of-speech tags are identified, and whether they are conjunctions or adverbs indicating parallel, transitional, or progressive relationships is determined. The words are then marked as connection nodes to generate a connection part-of-speech tag structure set; For the task of extracting and tagging conjunctions after punctuation marks, read the input text paragraph, use natural language processing tools to segment sentences and mark the part of speech of each word, identify the end punctuation of sentences through regular expressions or syntactic dependency analysis models, locate the end mark of each short sentence, count and extract the words immediately after the punctuation marks, filter out the words belonging to the category of conjunctions (such as CC) or adverbs (such as RB) in the part-of-speech tagging results, and use custom word lists (such as "but", "and", "therefore", etc.) to further judge whether they form semantic relationships such as parallelism, transition or progression. In the example, the original sentence can be set as "The equipment operation fails; therefore, the parameters need to be reconfigured", then identify " Therefore, "therefore" is a connecting adverb, which connects the previous and next sentences to establish a semantic connection and is marked as a progressive node. The corresponding annotation structure is: {"connection word": "therefore", "part of speech": "RB", "connection type": "progressive"}. A structure set list is established for each connecting word, and combined with the context syntactic analysis, its position information and connection type in the text are recorded. It is stored in JSON format, including fields such as "starting position", "connection word", "part of speech tag", "semantic connection type", etc., which will be used for subject-predicate structure recognition and semantic sentence block division. After completing this step, the connecting node will be used as the entry point for subsequent semantic analysis to generate a connecting part-of-speech tagging structure set.

[0023] S102: By connecting the part-of-speech tagging structure set, locating the verb and subject in the sentence corresponding to the connection node, extracting the subject phrases before and after the verb position and determining whether they form a subject-predicate structure, identifying the subject-predicate pair combinations with related meanings in the continuous sentences, recording the start and end positions and semantic attribution of the subject-predicate structure in the original text, and generating a subject-predicate phrase combination mapping sequence; Parse the set of connected part-of-speech tagging structures, index each connected node to the sentence it connects, extract the verb root position and track the subject phrases before and after it through the NLP dependency syntactic tree model (such as semantic dependency analysis based on BERT), and for structures such as "an error occurred; therefore, the module needs to be restarted", locate "need" as the predicate verb, and its dependent parent node is the subject "module". Based on the part-of-speech combination (NN+VB), determine whether it forms a subject-predicate structure. By tracing the dependency relationship nodes forward sentence by sentence, record the phrase combination that forms the subject-predicate structure, and extract the root, modifiers, qualifiers, etc. from each group of structures to form a complete subject-predicate group, such as "subject": module, "predicate": need to be restarted , establish a mapping entry for each group of structures, record the starting and ending character positions, the original position index and the semantic interval to which they belong, and form a subject-predicate combination mapping table, setting: "{subject: 'module', predicate: 'need to restart', starting position: 14, ending position: 22, semantics: 'restart operation'}", after each structure is verified to be semantically coherent (such as whether it shares a subject and whether the verb category is consistent), it is written into the mapping sequence for subsequent sentence block construction and functional partitioning processing. In the example, if there is "cannot be recognized; therefore, the interface needs to be adjusted", the two sentences share a subject and the verb structure matches, then they are merged into a pair of associated subject-predicate structures to generate a subject-predicate group combination mapping sequence.

[0024] S103: Based on the subject-predicate group mapping sequence, each semantic unit is arranged according to the order of the subject-predicate combination in the original text, sentence boundaries are divided and assembled into semantic sentence blocks, and based on the number of semantic repetitions between the subject and predicate in the semantic sentence blocks, whether a continuity between semantic groups is formed is determined, and sentence blocks with related subject-predicate combinations are grouped into the same interval to obtain the functional division of the sentence segments; Rearrange the order of the subject-verb pairs in the original text, define each combination as a basic semantic unit, and cut the original text into semantic sentence blocks by indexing its start and end positions. For example, the original text "The server is delayed; therefore, load balancing is performed; and capacity is further expanded" forms three subject-verb structures: "server-appears", "performs", and "capacity is expanded". Based on dependency analysis, "load balancing" and "further expansion" are judged as the continuation of the subject, forming a progressive structure. They are classified into the same interval according to the semantic sentence blocks, and the corresponding sentence segment function is set to "response processing flow". The semantic root between the subject and the predicate is used. Compare models (such as synonym discrimination and word vector distance comparison) to determine whether "load balancing" and "capacity expansion" constitute similar motivations or action groups. If the word vector similarity (cosine similarity) is greater than 0.75, it is considered a semantic continuation relationship and classified into the same functional partition. The output functional partition block includes: starting subject-verb combination, semantic attribution, and structure mapping index, such as: "{starting combination: 'carry out load balancing', ending combination: 'further expansion', interval semantics: 'response scheduling optimization'}". This partition is used for semantic scene merging or information extraction in subsequent tasks to obtain sentence segment functional partitioning.

[0025] The steps for obtaining the semantic chain trigger index sequence are as follows: S201: Calling sentence units in the sentence segment functional partition, extracting noun groups and verb groups with the highest frequency of occurrence in the sentence block, classifying the noun groups as subject words, and the verb groups as relation words, and recording the starting and ending positions in the sentence block as semantic index ranges to obtain a subject relation word position information set; For each semantic sentence block in the functional partition of the sentence segment, the noun group and verb group extraction operation is completed, the sentence is segmented, and then the part-of-speech tagging is performed on each word. The noun group includes general entities, tools, components, names, etc., and the verb group includes operation behaviors, state changes, functional reactions, etc. In practical applications, such as processing the text "restart service module, load driver component, user enters login information", the high-frequency nouns and verbs in each sentence block are counted. "Module" and "driver" will be counted as high-frequency nouns, while "restart", "load" and "input" belong to the verb group. According to the frequency of occurrence or the priority rule of occurrence position, the noun group is selected as the subject word and the verb group as the relation word, and their first and last occurrence positions are recorded respectively. It is set to appear in the first position of the first sentence for the first time and appear again in the middle position of the second sentence. The index range recorded as the subject word extends from 0 to the end of the sentence block. "Load" appears at the end of the second sentence for the first time. Its index range is determined, and the subject words and relation words in all semantic sentence blocks are identified. Accurate phrase positioning support is established for subsequent semantic trend judgment, and the subject relation word position information set is obtained.

[0026] S202: Based on the topic-related word position information set, the index ranges of the phrases are compared vertically according to the original arrangement order of the chunks in the text. The movement direction of the index position in the previous chunk and the next chunk is compared, and the forward movement, backward movement, or unchanged state is marked respectively. The original text position is recorded to obtain a semantic transition node index table; The positional index order of phrases within each semantic block is compared vertically, and their change trends between blocks are scanned block by block. From the previous block to the next block, whether the noun or verb group moves forward, backward, or remains unchanged is observed. If "module" appears near the beginning of a sentence in one block and appears in the middle or end of the next block, it indicates a backward shift. Conversely, if it moves from the end of the sentence to the beginning, it is a forward shift. In practical scenarios, such as describing "starting the module," "module initialization," and "driver module loading completion," the position of the noun "module" moves from the end of the sentence to the beginning of the sentence, forming a forward and backward movement trajectory. This process continuously compares the order of the blocks in the original text to determine whether the semantic core undergoes dynamic positional changes as the scenario progresses. Based on this, a directional change mark is recorded for each phrase. The mark content can be marked as "forward," "no change," or "backward." It serves as a key sequence for tracking concept evolution and is used for subsequent reversal point identification and chain trigger point screening to obtain a semantic turning point index table.

[0027] S203: Calling the semantic transition node index table, analyzing the sentence content corresponding to the transition node, extracting the verb and noun combination segments near the node, and determining whether the node falls in the combination segment after the verb and before the noun. If the position relationship is satisfied, it is marked as a semantic chain starting point, and a semantic chain trigger index sequence is generated; Based on the boundary position of the record, several words are extracted forward and backward to form the context scope. Grammatical dependency analysis is performed on the verbs and nouns within the scope to check whether there is a combination segment with a turning point after the verb and before the noun. Set the text "Tried to reconnect the interface module, but the connection failed" and determine that the verb "try" appears before the turning point and the noun "module" appears after the node. This position meets the semantic chain starting point condition. In application scenarios such as device control logic, the jump point of the instruction execution process appears with this structure. This place is located as the starting node of a new semantic clue. The character position of this node in the text and the sentence block number to which it belongs are recorded. This index is used to establish a connection mapping between semantic blocks. Through this chain starting point, all subsequent operation process statement segments controlled or affected by this node are tracked to form a starting path structure diagram of the semantic chain, providing basic data for subsequent logical path analysis and generating a semantic chain trigger index sequence.

[0028] The specific steps for obtaining the label path breakpoint position group are as follows: S301: Using the semantic chain trigger index sequence, locate the start chunk and end chunk corresponding to the trigger index in the text, extract the noun groups with the highest frequency of occurrence in the two chunks, record the position index range and the chunk identifier in the original text, and generate a set of chunk key noun groups; Extract the noun groups with the highest frequency of occurrence in the two chunks using the formula: ; in, Score the priority of the noun group, represents the frequency of the i-th noun group in the chunk, represents the average frequency of noun groups. represents the relevance score between the i-th noun group and the context, and k represents the number of noun groups; Parameter meaning and formula calculation derivation process: Representative The frequency of occurrence of the noun group in the text. This frequency is obtained by scanning a given chunk using a text analysis tool. In chunk A, noun group 1 appears 5 times and noun group 2 appears 3 times. The average value of the frequency of noun groups is obtained by calculating the frequencies of candidate noun groups and finding their arithmetic mean; It is The relevance score of the noun group is quantified through semantic analysis or contextual analysis. It is based on the weight of the noun group in the context and its influence on the text content. If a noun group has a high semantic match with the current chunk, it will be assigned a high relevance score (e.g., 0.8). is the number of noun groups to be analyzed. In this example, three noun groups are selected for analysis. ; Set the following data to the results obtained through actual text analysis tools or algorithms: Noun group frequency ( ): 5 times, 3 times, 2 times; Compute the mean frequency of noun groups: ; Relevance score ( ): 0.7, 0.8, 0.6; Calculate the absolute frequency difference for each noun group and multiply it by the logarithm: ; ; ; Sum and evaluate the numerator: ; Calculate the square of the difference between each frequency and the mean and sum them: ; ; ; ; Compute the product of each frequency and its relevance score and sum them: ; ; ; ; Substitute the above results into the formula to calculate: ; The results show that the noun group priority score is 0.97, indicating that the noun group has a high frequency priority in the language block and a high relevance score in the context. This priority score can be used as the basis for subsequent text analysis to determine that some noun groups are key elements in text semantic analysis.

[0029] S302: Calling the noun groups of the start and end chunks in the chunk key noun group set, cross-combining them in pairs, identifying whether there is lexical overlap in each combination, and counting the number of combinations without common words. If the number of combinations without common words reaches three or more, marking the start and end chunk segments as logical jump points, and obtaining a label path breakpoint position group; Read the three groups of high-frequency nouns of the marked starting and ending blocks in the key noun group set of the block, and generate cross-matching groups in pairs, that is, the first noun in the starting block is combined with the first to third nouns of the ending block, and then repeat the same operation for the second and third nouns, generating a total of 9 groups of combinations. Each group of combinations is compared to see if there is vocabulary overlap. The starting block is set to include "module", "interface", and "service", and the ending block is set to include "network" and "driver". After the combination, check one by one whether there is vocabulary consistency or high-similarity word roots. If there is no identical or similar word root in any combination, For semantic relationship words, statistics show that there is no common word combination. Continue to detect the remaining combinations. If the total number reaches 3 groups or more, it is determined that a semantic fault occurs between the start and end blocks, and this paragraph is marked as a logical jump point. In the actual corpus, if there is no identical entity or operation object between "loading driver module" and "user authorization access rights", the condition of logical isolation is met, that is, the label path breakpoint mark is triggered, the logical jump point is written into the path breakpoint position group, and its original text character start and end positions are recorded for subsequent semantic chain verification and path structure reconstruction, and the label path breakpoint position group is obtained.

[0030] The specific steps for obtaining the nested hierarchy table of the tag structure are as follows: S401: Calling the text segment corresponding to the jump point in the tag path breakpoint position group, extracting noun phrases that appear more than twice in the text segment as candidate tag phrases, recording the first word position and the last word position of the tag phrase in the sentence, and obtaining a tag phrase position indexing table; Enter the paragraph content analysis process corresponding to the jump point, extract noun phrases from the text in the paragraph, including common nouns, proper nouns and noun combinations in phrase form, use a grammar analyzer to mark the parts of speech, perform word frequency statistics on all noun phrases, and filter out phrases that appear more than twice as candidate label phrases. Set it in a description of "module loading failed, module interface response abnormality, module error detection, module restart initialization", "module" appears four times, and "interface" appears twice. After meeting the screening conditions, "module" and "interface" are used as candidate label phrases. The word order position of the first word in the sentence and the word order position of the last word in the sentence are recorded. For example, "module" first appears in the second word position and last appears in the 20th word position. Each record includes the phrase content, starting word order position, ending word order position and the paragraph number to which it belongs, providing structural index support for subsequent identification of relationships between labels, ensuring the accuracy of the relative positions between phrases, and obtaining a label phrase position indexing table.

[0031] S402: Arrange the tag phrases in order within the sentence based on the start and end position indexes of each group of phrases in the tag phrase position indexing table, compare the front and back boundary ranges to see if there is an overlapping area or nested relationship, and generate a tag path nested interleaved tag set; Compare the front and back boundaries to see if there is an overlapping area or nested relationship, using the formula: ; in, Represents the boundary intersection strength index between the p-th group and the q-th group of label phrases, Represents the starting position index value of the p-th group of label phrases, Represents the end position index value of the pth group of label phrases, Represents the starting position index value of the qth group of label phrases, Represents the ending position index value of the qth group of label phrases, Represents the starting position index value of the rth group of label phrases, Represents the average value of the starting position index value of the first u groups of label phrases, Represents the ending position index value of the rth group of label phrases, Represents the average value of the end position index value of the first u groups of label phrases, Represents the span value of the rth group of label phrases, is the number of label pairs; Parameter meaning and formula calculation derivation process: The tag phrase position indexing table contains the following five groups of tag phrases, and their starting position index values and ending position index values are as follows: Group 1: , ; Group 2: , ; Group 3: , ; Group 4: , ; Group 5: , ; Calculate the average of the starting position index values of the first five groups of label phrases: ; Calculate the average of the end position index values of the first five groups of tag phrases: ; Calculate the boundary intersection strength index between the first and second groups of label phrases : First calculation: ; Second calculation: ; ; Take the square root: ; Substitute into the formula for calculation : ; The results show that the boundary interleaving strength index between the first and second groups of label phrases is 8.294. A larger value indicates a higher degree of interleaving. This index can be used to determine whether there is an intersection area or nested relationship between label phrases and to generate a nested interleaving tag set for label paths.

[0032] S403: calling the tag path nested interleaved mark set, progressively establishing tag level index numbers according to the nesting layer, classifying phrases with a common nesting starting point into the same level path group, and obtaining a tag structure nesting level table; Start by building hierarchical numbers according to the nesting depth. Assign a first-level path group number to the outermost phrase in the nested structure. Continue to identify the nested sub-phrases and mark them as second-level, third-level path group numbers in the nesting order. If "module loading" nests "module loading", and "module loading" nests "module", the three are respectively classified as first-level, second-level, and third-level path groups. Each phrase is numbered according to its hierarchical structure, including phrase content, hierarchy, nesting starting point, phrase position interval, etc. If two or more phrases have the same nesting starting point but different ending points, they are determined to be parallel path groups at the same level and marked as independent path branches at the same level. For example, if "module configuration", "module loading", and "module detection" have the same starting position but different ending positions, they are classified under the same path group number and parallel labels are established, forming a clear label semantic hierarchy map, providing a complete label system support for text structure classification, abstract extraction, or semantic indexing, and obtaining a label structure nesting hierarchy table.

[0033] The steps for obtaining the language model label mapping combination list are as follows: S501: Call the hierarchical index number of each label path in the label structure nested hierarchical table, search the context segment of the corresponding label in the original text in the pre-trained corpus of the large language model, extract the position identification range of the label context segment in the model input, and establish a context input index set of the label segment to obtain the label context mapping location set; Extract the path structure from the nested hierarchical table of the label structure, read the hierarchical index number corresponding to each label path, and then construct a search request for the label content at each level to match the context segment in the pre-trained corpus of the large language model. Set "module loading" as the search label and call the semantic matching interface of the pre-trained model to filter out context segments containing this label or its synonymous variants in the model corpus. Then obtain the character or token position identifier range of the segment in the original model input. For example, segments such as "module loading exception" and "module loading process" correspond to token numbers between 115 and 125. This range is recorded as the input interval identifier of the label in the pre-trained data and is bound to the original label path hierarchical number. Each set of index records includes fields such as the label name, the level it is in, the context segment content, and the token start and end positions. This set is summarized as the label context mapping location set, which is used for subsequent model output trajectory monitoring and semantic sequence tracking to obtain the label context mapping location set.

[0034] S502: Based on the label context mapping location set, the corresponding label structure nesting hierarchical information is injected into the output mapping sequence of the large language model. The label word direction representations generated by the output layer at the corresponding positions before and after the injection are extracted. The distribution trajectories between each group of label word directions are compared, and the offset direction, angle, and spacing are located and classified to obtain label word direction distribution deviation information. Call the label context mapping to locate the input position index information of each label path in the set, and perform semantic embedding tracking on the input position in turn when the language model performs the generation task, and record the label word-directed representation of the corresponding position in the model output sequence when no structural information is injected. Then, based on the same input, the nested hierarchical information of the label structure is injected, and the output sequences before and after the injection are compared to see the changes in the word-directed representation generated at the corresponding label position. Focus on monitoring the direction of the word-directed deviation (such as concentration towards the semantic center or divergence towards the remaining label paths), angle (approximately represented by the angle between high-dimensional vectors), and relative spacing (such as the Euclidean distance between vectors). Indicators are used to determine the impact and degree of disturbance on the semantic representation after injecting structural information. This information describes the trajectory of semantic association changes between labels, and based on this, a word-directed migration model is constructed to provide a quantitative reference basis for semantic path optimization and obtain label word-directed distribution deviation information.

[0035] S503: Extracting word-direction paths that can be used for semantic backtracking based on label combination trajectories that have continuous offsets and consistent directions in the label word-direction distribution deviation information, determining whether the head-to-tail logical vocabulary combinations connecting the labels in the word-direction paths satisfy the original semantic chain order, screening label paths that meet the sequential logic and arranging them into nested combinations, and obtaining a language model label mapping combination list; Analyze the label combination trajectories with continuous offsets and consistent directions, and screen out word-oriented paths with stable semantic evolution directions. The paths continuously progress in a certain direction in space, which is manifested as an arrangement pattern in which the angles between several label word-oriented vectors are small and the spacing is gradually shortened. Set the word-oriented representations of the three groups of label paths "module detection → module initialization → module loading" to present a distribution direction with continuous offset to the upper right. Based on such trajectories, perform semantic backtracking on the word-oriented paths to determine whether there is a logical connection between the labels at both ends of the path in the original text, such as "detection" as the pre-operation and "loading" as the final state. If the two also show a sequence relationship in the original semantic chain, it is considered to meet the semantic order requirements and the path is retained as a legal label sequence. The label paths that meet the conditions are reorganized and arranged in a nested structure for multi-level label tracking, cross-layer semantic integration, and semantic diffraction analysis of nested structures to form a complete language model label mapping combination list.

[0036] See also Figure 2 , a system for processing data based on a large language model, the system includes: The chunk recognition module identifies punctuation marks within a text paragraph, detects the part of speech of the conjunctions between the punctuation marks, and determines whether the sentences adjacent to the conjunctions contain subject phrases and predicate verb structures. It then divides the entire paragraph into multiple chunks based on the order in which the phrases appear in the sentence, generating functional partitions for the segments. The semantic indexing module uses the sentence segment functional partitioning to extract the subject words and relational words within the sentence block, record the position of the phrase in the context block, identify whether the phrase movement direction has reversed, determine whether the reversal point is between the verb and noun combination segment, and obtain the semantic chain trigger index sequence; The logical jump module triggers the index sequence according to the semantic chain, locates the corresponding start and end blocks in the text paragraph, extracts all noun groups in the start and end blocks, and arranges them according to the noun groups at both ends to generate a label path breakpoint position group; The label extraction module uses the label path breakpoint position group to extract the label phrase from the sentence block pointed to by the jump node, identifies the sentence position number of the label phrase in the sentence block, determines whether there is an interleaved position or nested arrangement, and obtains the label structure nesting level table; The label mapping module uses a nested hierarchical table of label structures to extract the context of the paragraphs that match the label phrases in the language model, identify the word vector position distribution of the label group in the output layer of the language model, compare the corresponding differences with the default word vector distribution, and generate a list of language model label mapping combinations.

[0037] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A method for processing data based on a large language model, characterized in that: The following steps are involved: S1: Based on the sentences in the text paragraph, locate the punctuation marks between sentences, tag the connectives of adjacent sentences with parts of speech, identify verb and subject pairs, divide the sentence blocks according to the order of phrase matching, and generate sentence segment functional partitions; S2: Calling the segment functional partition, extracting the subject words and relation words of the text paragraph, marking the position of the phrase in the context of the block, recording the semantic trend of the phrase according to the direction of position change, identifying the direction reversal nodes in two consecutive blocks, and generating a semantic chain trigger index sequence; S3: Using the semantic chain to trigger the index sequence, locate the start and end blocks of the keywords in the label path segment, extract the key noun groups at both ends, cross-combine the two groups of words and determine whether there are three or more groups of non-co-word structures. If so, mark them as logical jump points and generate a label path breakpoint position group; S4: calling the label path breakpoint position group, extracting the label phrase from the corresponding segment, recording the position sequence of the label phrase in the sentence, identifying whether there is path interleaving or position nesting, analyzing the nesting hierarchical relationship, and generating a label structure nesting hierarchical table.

2. The method for processing data based on a large language model according to claim 1, characterized in that: The sentence segment functional partition includes semantic connection boundary points, language block structure boundary lines, and sentence trunk matching groups; the semantic chain trigger index sequence includes transition node sequence, semantic direction turning point, and language block combination starting position; the label path breakpoint position group includes label semantic jump area, keyword disjoint fragment, and path coherence breakpoint; the label structure nested hierarchical table includes label hierarchical index set, nested path number column, and label interleaved sequence diagram.

3. The method for processing data based on a large language model according to claim 1, characterized in that: The steps for obtaining the functional partition of the sentence segment are specifically as follows: S101: Based on the sentences in the text paragraph, punctuation marks between sentences are scanned, the connection relationship between the sentences is analyzed, the conjunctions after the sentence-end punctuation are extracted as analysis objects, the part-of-speech tags are identified, and whether they are conjunctions or adverbs indicating parallel, transitional, or progressive relationships is determined. The words are then marked as connection nodes to generate a connection part-of-speech tag structure set; S102: Using the connection part-of-speech tagging structure set, locate the verb and subject in the sentence corresponding to the connection node, extract the subject phrases before and after the verb position and determine whether they form a subject-predicate structure, identify subject-predicate pair combinations with related meanings in continuous sentences, record the start and end positions and semantic attribution of the subject-predicate structure in the original text, and generate a subject-predicate phrase combination mapping sequence; S103: According to the subject-predicate group combination mapping sequence, each group of semantic units is arranged according to the order of the subject-predicate combination in the original text, sentence boundaries are divided and assembled into semantic sentence blocks, and according to the number of semantic repetitions between the subject and predicate in the semantic sentence blocks, whether continuity between semantic groups is formed is determined, and sentence blocks with related subject-predicate combinations are classified into the same interval to obtain the functional partitioning of sentence segments.

4. The method for processing data based on a large language model according to claim 3, characterized in that: The steps for obtaining the semantic chain trigger index sequence are specifically as follows: S201: calling the sentence units in the sentence segment functional partition, extracting the noun groups and verb groups with the highest frequency of occurrence in the sentence block, classifying the noun groups as subject words, classifying the verb groups as relation words, and recording the starting position and ending position in the sentence block as a semantic index range to obtain a subject relation word position information set; S202: Based on the topic-related word position information set, the index ranges of the phrases are vertically compared according to the original arrangement order of the language blocks in the text, and the movement direction of the index position in the previous language block and the next language block are compared. The forward movement, backward movement, or unchanged state is marked respectively, and the original text position is recorded to obtain a semantic transition node index table; S203: Call the semantic transition node index table, analyze the sentence content corresponding to the transition node, extract the verb and noun combination segments near the node, and determine whether the node falls in the combination segment after the verb and before the noun. If the position relationship is met, mark it as the semantic chain starting point and generate a semantic chain trigger index sequence.

5. The method for processing data based on a large language model according to claim 4, characterized in that: The steps for obtaining the label path breakpoint position group are specifically as follows: S301: Using the semantic chain trigger index sequence, locate the start chunk and the end chunk corresponding to the trigger index in the text, extract the noun groups with the highest frequency of occurrence in the two chunks, record the position index range and the chunk identifier in the original text, and generate a chunk key noun group set; The formula for extracting the noun groups with the highest frequency of occurrence in the two language blocks is: ; in, Score the priority of the noun group, represents the frequency of the i-th noun group in the chunk, represents the average frequency of noun groups. represents the relevance score between the i-th noun group and the context, and k represents the number of noun groups; S302: Call the noun groups of the starting language block and the ending language block in the set of key noun groups of the language block, perform cross-combinations in pairs, identify whether there is vocabulary overlap in each combination, count the number of combinations without common words, and if the number of combinations without common words reaches three or more, mark the starting and ending language block segments as logical jump points, and obtain the label path breakpoint position group.

6. The method for processing data based on a large language model according to claim 5, characterized in that: The steps for obtaining the tag structure nesting level table are as follows: S401: Calling the text segment corresponding to the jump point in the tag path breakpoint position group, extracting noun phrases that appear more than twice in the text segment as candidate tag phrases, recording the first word position and the last word position of the tag phrase in the sentence, and obtaining a tag phrase position indexing table; S402: Arrange the tag phrases in order within the sentence based on the start and end position indexes of each group of phrases in the tag phrase position indexing table, compare the front and back boundary ranges to see if there is an overlapping area or nested relationship, and generate a tag path nested interleaved tag set; S403: calling the nested interleaved tag set of the tag path, progressively establishing tag level index numbers according to the nesting layers, classifying phrases with a common nesting starting point into the same level path group, and obtaining a tag structure nesting level table.

7. The method for processing data based on a large language model according to claim 6, characterized in that: The formula is used to determine whether there is an overlapping area or a nested relationship between the front and back boundary ranges of the comparison: ; in, Represents the boundary intersection strength index between the p-th group and the q-th group of label phrases, Represents the starting position index value of the pth group of label phrases, Represents the end position index value of the pth group of label phrases, Represents the starting position index value of the qth group of label phrases, Represents the ending position index value of the qth group of label phrases, Represents the starting position index value of the rth group of label phrases, Represents the average value of the starting position index value of the first u groups of label phrases, Represents the ending position index value of the rth group of label phrases, Represents the average value of the end position index value of the first u groups of label phrases, Represents the span value of the rth group of label phrases, is the number of label pairs.

8. The method for processing data based on a large language model according to claim 1, characterized in that: The method further comprises step S5: S5: Extracting the context position of the label segment from the large language model pre-training corpus based on the label structure nesting hierarchy table, mapping the label structure corresponding to the nesting hierarchy to the language model output layer, comparing the deviation of the label word distribution at the corresponding position in the mapped structure and the default output of the language model, identifying the label trunk combination order in the traceable path, and generating a language model label mapping combination list; The language model label mapping combination list includes a semantic nested mapping path, a label structure corresponding relationship set, and a word-direction distribution difference mark sequence.

9. The method for processing data based on a large language model according to claim 8, characterized in that: The steps for obtaining the language model label mapping combination list are specifically as follows: S501: Call the hierarchical index number of each label path in the nested hierarchical table of the label structure, search the context segment of the corresponding label in the original text in the pre-trained corpus of the large language model, extract the position identification range of the label context segment in the model input, and establish a context input index set of the label segment to obtain a label context mapping location set; S502: Based on the label context mapping positioning set, the corresponding label structure nesting level information is injected into the output mapping sequence of the large language model, the label word direction representations generated by the output layer at the corresponding positions before and after the injection are extracted, and the distribution trajectories between each group of label word directions are compared. The offset direction, angle and spacing are located and classified to obtain label word direction distribution deviation information; S503: Based on the label combination trajectory with continuous offset and consistent direction in the label word distribution deviation information, extract the word path that can be used for semantic backtracking, determine whether the head and tail logical vocabulary combination connecting the labels in the word path meets the original semantic chain order, filter the label paths that meet the sequential logic and arrange them into nested combinations, and obtain the language model label mapping combination list.

10. A system for processing data based on a large language model, characterized in that: The system is used to implement the method for processing data based on a large language model according to any one of claims 1 to 9, and the system includes: The chunk recognition module identifies punctuation marks within a text paragraph, detects the part of speech of the conjunctions between the punctuation marks, and determines whether the sentences adjacent to the conjunctions contain subject phrases and predicate verb structures. It then divides the entire paragraph into multiple chunks based on the order in which the phrases appear in the sentence, generating functional partitions for the segments. The semantic indexing module calls the sentence segment functional partition, extracts the subject words and relation words in the sentence block, records the position of the phrase in the context block, identifies whether the movement direction of the phrase is reversed, determines whether the reversal point is located between the verb and noun combination segment, and obtains the semantic chain trigger index sequence; The logic jump module triggers the index sequence according to the semantic chain, locates the corresponding starting and ending language blocks in the text paragraph, extracts all noun groups in the starting and ending language blocks, and arranges them according to the noun groups at both ends to generate a label path breakpoint position group; The label extraction module uses the label path breakpoint position group to extract the label phrase from the sentence block pointed to by the jump node, identifies the sentence position number of the label phrase in the sentence block, determines whether there is an interleaved position or nested arrangement, and obtains a label structure nesting level table; The label mapping module uses the nested hierarchical table of the label structure to extract the context of the paragraph that matches the label phrase in the language model, identifies the word vector position distribution of the label group in the output layer of the language model, compares the corresponding difference with the default word vector distribution, and generates a language model label mapping combination list.

Citation Information

Patent Citations

  • Document processing method and device, storage medium and electronic equipment

    CN110765237A

  • Intelligent search engine construction method based on large language model

    CN118964589A

  • Text classification method and system based on natural language processing

    CN119691179A

  • Intelligent equipment fault information extraction method based on multi-semantic knowledge interaction and dynamic pruning

    CN120068877A

  • Data conversion device and method, database managing device and method, and database search system and method

    WO2008007683A1

Cited By

  • AI-based sentiment analysis platform

    CN120874851A

  • AI-based sentiment analysis platform

    CN120874851B

  • Multi-modal digital publishing intelligent checking system and method based on large model

    CN120930637A

  • Big model-based multi-modal digital publishing intelligent proofreading system and method

    CN120930637B

  • Artificial intelligence interaction robot

    CN121071110A