Text extraction method and device, equipment and medium

By segmenting and encoding text, and using a pre-trained large text extraction model, the problem of low text extraction efficiency in the prior art is solved, and efficient and accurate text information extraction is achieved.

CN120045637APending Publication Date: 2025-05-27ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510206885.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing text extraction methods are inefficient and cannot efficiently extract effective information from large amounts of text.

Method used

By obtaining the text to be processed, performing slicing processing according to the preset slicing granularity, inserting index encoding, forming the encoded text, and inputting it into the pre-trained large text extraction model to output the target index encoding and determine the corresponding target subtext.

Benefits of technology

It improves the efficiency of text extraction, simplifies the generation process of text extraction model, reduces the amount of data processing, and ensures the accuracy of the extracted subtext.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045637A_ABST
    Figure CN120045637A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a text extraction method and device, equipment and a medium. According to the scheme, the method comprises the steps of performing segmentation processing on an obtained to-be-processed text according to a preset segmentation granularity to obtain a plurality of segmentation positions; inserting index codes for distinguishing different sub-texts at the plurality of segmentation positions to obtain a coded text; and inputting the encoded text into a pre-trained text extraction model, outputting a target index code meeting task requirements by utilizing the text extraction model, and determining a target sub-text corresponding to the target index code in the to-be-processed text according to the target index code.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a text extraction method, device, equipment and medium. Background Art

[0002] With the massive generation of text information in the Internet era, extracting text information to obtain effective information has become a research hotspot. Based on this, how to extract text has become a technical problem that needs to be solved urgently. Summary of the invention

[0003] The embodiments of this specification provide a text extraction method, apparatus, device and medium to solve the problem of low efficiency in existing text extraction methods.

[0004] To solve the above technical problems, the embodiments of this specification are implemented as follows:

[0005] A text extraction method provided in an embodiment of this specification includes:

[0006] Get the text to be processed;

[0007] According to a preset segmentation granularity, the text to be processed is segmented to obtain a plurality of segmentation positions;

[0008] Inserting index codes for distinguishing different subtexts at the plurality of segmentation positions to obtain encoded texts;

[0009] Inputting the encoded text into a text extraction model to obtain a target index code output by the text extraction model; the text extraction model is a pre-trained large model used to extract the index code of the sub-text that meets the task requirements from the input text;

[0010] According to the target index code, a target subtext corresponding to the target index code in the text to be processed is determined.

[0011] The embodiment of this specification provides a text extraction device, including:

[0012] An acquisition module is used to obtain the text to be processed;

[0013] A segmentation processing module is used to segment the text to be processed according to a preset segmentation granularity to obtain a plurality of segmentation positions;

[0014] An encoding module, used for inserting index codes for distinguishing different subtexts at the plurality of segmentation positions to obtain encoded texts;

[0015] An input module, used to input the encoded text into a text extraction model to obtain a target index code output by the text extraction model; the text extraction model is a pre-trained large model, used to extract the index code of the sub-text that meets the task requirements from the input text;

[0016] The extraction module is used to determine the target subtext corresponding to the target index code in the text to be processed according to the target index code.

[0017] A text extraction device provided in an embodiment of this specification includes:

[0018] at least one processor; and,

[0019] a memory communicatively connected to the at least one processor; wherein,

[0020] The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to:

[0021] Get the text to be processed;

[0022] According to a preset segmentation granularity, the text to be processed is segmented to obtain a plurality of segmentation positions;

[0023] Inserting index codes for distinguishing different subtexts at the plurality of segmentation positions to obtain encoded texts;

[0024] Inputting the encoded text into a text extraction model to obtain a target index code output by the text extraction model; the text extraction model is a pre-trained large model used to extract the index code of the sub-text that meets the task requirements from the input text;

[0025] According to the target index code, a target subtext corresponding to the target index code in the text to be processed is determined.

[0026] An embodiment of the present specification provides a computer-readable medium on which computer-readable instructions are stored. The computer-readable instructions can be executed by a processor to implement a text extraction method.

[0027] An embodiment of this specification can achieve the following beneficial effects:

[0028] The obtained text to be processed is segmented according to a preset segmentation granularity to obtain a number of segmentation positions and a number of sub-texts; index codes for distinguishing different sub-texts are inserted at the number of segmentation positions to obtain an encoded text; the encoded text is input into a pre-trained text extraction model, and the text extraction model is used to output a target index code that meets the task requirements, and the target sub-text corresponding to the target index code in the text to be processed is determined according to the target index code. The text extraction model can output the target index code of the sub-text that meets the task requirements, and the text extraction model does not need to output the text itself, which simplifies the generation process of the text extraction model, and can also reduce the data processing amount of the text extraction model and improve data processing efficiency.

[0029] On the other hand, the meaning of a sentence or word may be affected by the context, and the same sentence may have different meanings in different contexts. If the text to be processed includes multiple identical sentences or words, some of the multiple identical sentences or words may meet the task requirements, while others do not. In the embodiment of this specification, the text extraction model can output the index code of the subtext that meets the task requirements, so that the subtext that meets the task requirements can be clearly indicated, avoiding the inability to distinguish the subtext that meets the task requirements from other subtexts in the text to be processed that are the same as the subtext that meets the task requirements.

[0030] On the other hand, the text extraction model outputs the index encoding of the sub-text that meets the task requirements, which can also avoid the problem of inaccurate sub-text generated by the text extraction model, and thus avoid the abnormal situation caused by the inconsistency between the generated sub-text and the content in the text to be processed. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0032] Figure 1 This is a schematic diagram of an application scenario of a text extraction method provided in an embodiment of this specification;

[0033] Figure 2 is a flowchart of a text extraction method provided by an embodiment of this specification;

[0034] Figure 3 It is a swim lane diagram of a text extraction method provided in an embodiment of this specification;

[0035] Figure 4 is a structural schematic diagram of a text extraction device provided in an embodiment of this specification;

[0036] Figure 5 It is a structural schematic diagram of a text extraction device provided in an embodiment of this specification. DETAILED DESCRIPTION

[0037] In order to make the purpose, technical solutions and advantages of one or more embodiments of this specification clearer, the technical solutions of one or more embodiments of this specification will be clearly and completely described below in combination with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of one or more embodiments of this specification.

[0038] To facilitate understanding of the embodiments of this specification, the following is an explanation of the terms involved in the embodiments of this specification.

[0039] Position index: the position of each word, sentence or paragraph in the text.

[0040] Explicit position encoding: A concept defined to distinguish the method introduced by position indexing; explicit position encoding is to explicitly specify a position encoding vector for each position of the input sequence, so that the model can capture the position information of the elements in the sequence.

[0041] Implicit position encoding: This is also a concept defined to distinguish the method introduced by position index; implicit position encoding refers to implicitly incorporating position information into the input or feature representation of the model. Compared with explicit position encoding, implicit position encoding does not require explicitly specifying a position encoding vector for each position of the input sequence.

[0042] Sequence labeling: A basic task in the field of natural language processing (NPL), the goal is to assign a label to each element in a text sequence.

[0043] Large Language Model: Large Language Model (LLM), also known as large language model or large model, is a natural language processing technology based on deep learning. Examples include models GPT-4, Claude, Qwen, PaLM, Galactica, and LLaMA.

[0044] The technical solutions provided by the embodiments of this specification are described in detail below in conjunction with the accompanying drawings.

[0045] In the Internet era, a large amount of text information is generated, and extracting text information to obtain effective information has become a research hotspot.

[0046] Currently, large models can be used to extract and output text content that meets business needs from the text to be processed. However, it may take a long time and be inefficient for large models to directly output specific text information.

[0047] In order to solve the defects in the prior art, this solution provides the following embodiments:

[0048] Figure 1 The present invention is a schematic diagram of the overall solution flow of a text extraction method in an embodiment of the present specification.

[0049] like Figure 1 As shown, the scheme may include a user terminal 101, a server 102, and a text extraction model 103; the user may provide the text to be processed through the user terminal 101. The user terminal 101 may send the text to be processed to the server 102, and the server 102 may segment the text to be processed according to the preset segmentation granularity according to the processing requirements to obtain a number of segmentation positions; index codes for distinguishing different sub-texts may also be inserted at the several segmentation positions to obtain the encoded text. The server 2 may also call the text extraction model 103 to input the encoded text into the text extraction model 103, so that the text extraction model 103 outputs the target index code corresponding to the sub-text that meets the task requirements. The server 102 may extract the target sub-text corresponding to the target index code from the encoded text according to the target index code. If the user terminal needs to display the result of the extracted text, the terminal may also display the target sub-text. In the example Figure 1 In the scenario shown, the text extraction model 103 can be installed on the server 102 or on other servers that are in communication with the server 2 .

[0050] although Figure 1 , after receiving the to-be-processed text to be processed, the user terminal 101 can send the to-be-processed text to be processed to the server 102. In actual application, if the computing resources of the user terminal 101 meet the conditions for running the text extraction method, the scheme of extracting the target subtext that meets the task requirements from the to-be-processed text can also be executed on the user terminal 101. In this case, the called text extraction model 103 can be deployed on the user terminal 101, or can be deployed on other terminals or servers that are connected to the user terminal 101 in communication.

[0051] In such Figure 1 In the application scenario shown, the server may be connected to one or more terminal devices via a local area network connection, a wide area network connection, an Internet connection, or other types of data networks. Figure 1 The server 102 may include but is not limited to any device, equipment, platform, device cluster, etc. with computing and processing capabilities. Figure 1 The user terminal 101 may include but is not limited to a smart phone, a tablet computer, a laptop computer, a PDA, a personal computer, a smart home device, a vehicle-mounted device, etc.

[0052] It is understandable that the text extraction method provided in the embodiments of this specification can also be used to implement the intermediate process executed in certain business processes. For example, in the scenario of generating an article about a certain industry using a large language model or other intelligent tools, such as generating an industry report, key information, such as key sentences, can be first extracted from the acquired news reports about the industry, and then these key sentences are combined to generate an article related to the industry. The above-mentioned text to be processed can be obtained from the server that generates the industry article from the network or other databases without the need for the user to provide it himself. The above-mentioned user terminal can also be omitted in actual applications.

[0053] Next, a text extraction method provided in the embodiment of the specification will be specifically described with reference to the accompanying drawings:

[0054] Figure 2 A flowchart of a text extraction method provided in an embodiment of this specification. From a program perspective, the execution subject of the process can be a program or application client installed on an application server. It can be understood that the method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities.

[0055] like Figure 2 As shown, the process may include the following steps:

[0056] Step 202: Obtain the text to be processed.

[0057] In the embodiments of the present specification, the text to be processed may be a text that needs to be processed, and may be a text that only includes a plurality of phrases, a text that includes a plurality of sentences, or a text that includes a plurality of paragraphs.

[0058] In practical applications, the text to be processed can be text obtained by various means. As an example, the text to be processed can be text extracted from an existing file, a data set, or a database; as another example, the text to be processed can also be text obtained by using OCR to recognize a page image. The page image can be a photo, a webpage screenshot, etc. The page image can contain only text information, or it can contain multimodal information such as text and images, and the embodiments of this specification do not limit this.

[0059] Step 204: segment the text to be processed according to a preset segmentation granularity to obtain a plurality of segmentation positions.

[0060] In the embodiments of this specification, the preset segmentation granularity may be a preset encoding granularity; the encoding granularity may refer to the size of the smallest unit into which the text data is segmented. In practical applications, the preset segmentation granularity may be determined according to task requirements. It is understandable that the preset segmentation granularity may be different for different tasks.

[0061] In the embodiments of this specification, segmentation processing may refer to dividing a text into smaller units, such as words, vocabulary, sentences or paragraphs. The segmentation position may be a preset position in a subtext obtained after segmentation processing is performed on the text to be processed.

[0062] Step 206: inserting index codes for distinguishing different sub-texts at the plurality of segmentation positions to obtain encoded texts.

[0063] In the embodiments of this specification, the subtext may be a plurality of texts obtained after segmentation of the to-be-processed text. The size of the subtext may be a text that meets the preset segmentation granularity requirements. Specifically, the subtext may be a word, a sentence, or a paragraph.

[0064] The index code can be a character used to identify or distinguish different sub-texts. In the embodiment of the present specification, the method of inserting the index code to obtain each segmentation position is explicit position coding. The explicit introduction of the position information of each sub-text can improve the accuracy of the output position code, and the index code can be used to represent each sub-text. It can be understood that the index codes of different sub-texts are different. In practical applications, in order to avoid confusion between the index code and the characters in the text to be processed, the index code can be a character different from the characters included in the text to be processed, or it can be a character with a different form from the characters included in the text to be processed.

[0065] In practical applications, you can insert index codes at the segmentation positions by writing scripts.

[0066] It can be understood that the encoded text may include various sub-texts and index codes of various sub-texts.

[0067] Step 208: Input the encoded text into a text extraction model to obtain a target index code output by the text extraction model; the text extraction model is a pre-trained large model used to extract index codes of sub-texts that meet task requirements from the input text.

[0068] In the embodiments of the present specification, the text extraction model may be a large language model that can be used to perform at least one of the tasks such as information extraction (IE) and machine reading comprehension (MRC). Specifically, the text extraction model may be a large language model corresponding to the task requirements; if the task requirement is to extract low-quality sentences in the text to be processed, then the text extraction model may be a low-quality model for extracting low-quality sentences; if the task requirement is to extract high-quality sentences in the text to be processed, then the text extraction model may be a high-quality model for extracting high-quality sentences; if the task requirement is to extract generalized sentences in the text to be processed, then the text extraction model may be a model for extracting generalized sentences; if the task requirement is to extract sentences in the text to be processed that represent the central idea of ​​the text to be processed, then the text extraction model may be a model for extracting the central idea of ​​the text to be processed; if the task requirement is to extract entities in the text to be processed, then the text extraction model may be a model for extracting entities.

[0069] In practical applications, a text extraction model can be selected according to the correspondence between the task requirements and the text extraction model, so as to process the encoded text using the selected text extraction model. The correspondence between the task requirements and the text extraction model can be pre-constructed.

[0070] As an implementation method, the text extraction model may also be a model that can comprehensively handle various tasks. For example, the text extraction model may be a comprehensive model that can extract low-quality sentences, high-quality sentences, and entities.

[0071] The target index code can be the index code of the sub-text that meets the task requirements and is determined by the text extraction model. In the embodiment of the present specification, the method of inserting the index code at each segmentation position can be used as an explicit position coding method, and the position information of each sub-text is explicitly introduced to improve the accuracy of the output target position code. In the embodiment of the present specification, the text extraction model can output the index code of the sub-text that meets the task requirements, which can avoid the text extraction model from generating sub-texts that meet the task requirements, simplify the generation process of the text extraction model, and reduce the data processing amount of the text extraction model, thereby improving data processing efficiency.

[0072] Step 210: According to the target index code, determine the target subtext corresponding to the target index code in the text to be processed.

[0073] The target subtext is a subtext that meets the task requirements. In the embodiment of this specification, after the text extraction model outputs a target index number that meets the task requirements, the server can match the target index code to determine the target subtext corresponding to the target index code.

[0074] It should be understood that the order of some steps in the methods described in one or more embodiments of this specification can be interchanged according to actual needs, or some steps can be omitted or deleted.

[0075] Figure 2 The method in the invention is to segment the acquired text to be processed according to the preset segmentation granularity to obtain several segmentation positions and several subtexts; insert index codes for distinguishing different subtexts at several segmentation positions to obtain encoded text; input the encoded text into a pre-trained text extraction model, use the text extraction model to output a target index code that meets the task requirements, and determine the target subtext corresponding to the target index code in the text to be processed according to the target index code. The text extraction model can output the target index code of the subtext that meets the task requirements, and the text extraction model does not need to output the text itself, which simplifies the generation process of the text extraction model, and can also reduce the data processing amount of the text extraction model and improve the data processing efficiency.

[0076] On the other hand, the meaning of a sentence or word may be affected by the context, and the same sentence may have different meanings in different contexts. If the text to be processed includes multiple identical sentences or words, some of the multiple identical sentences or words may meet the task requirements, while others do not. In the embodiment of this specification, the text extraction model can output the index code of the subtext that meets the task requirements, so that the subtext that meets the task requirements can be clearly indicated, avoiding the inability to distinguish the subtext that meets the task requirements from other subtexts in the text to be processed that are the same as the subtext that meets the task requirements.

[0077] On the other hand, the text extraction model outputs the index encoding of the sub-text that meets the task requirements, which can also avoid the problem of inaccurate sub-text generated by the text extraction model, and thus avoid the abnormal situation caused by the inconsistency between the generated sub-text and the content in the text to be processed.

[0078] based on Figure 2 The method, the examples of this specification also provide some specific implementation plans of the method, which are described below.

[0079] Optionally, segmenting the text to be processed according to a preset segmentation granularity may specifically include:

[0080] The text to be processed is input into a text segmentation model for segmentation according to the preset segmentation granularity.

[0081] In the embodiments of this specification, a text segmentation model can be used to segment the text to be processed. The text segmentation model can be a model that can perform segmentation processing according to a preset segmentation granularity, or a model that can perform segmentation processing according to multiple segmentation granularities. In addition, the text segmentation model can be an existing natural language processing model, such as a rule-based segmentation model, a content-aware segmentation model, a semantic clustering segmentation model, etc.; it can also be a model obtained after data training, such as an LSTM model or a BERT model.

[0082] In the embodiment of the present specification, optionally, the preset segmentation granularity may include any one of representing words, sentences, and paragraphs.

[0083] In the embodiments of this specification, the preset segmentation granularity can be determined according to the task requirements. Specifically, if the task requirement is to extract paragraphs from the text to be processed, such as extracting low-quality paragraphs, high-quality paragraphs, summaries, etc. from the text to be processed, the preset segmentation granularity can be paragraphs. If the task requirement is to extract sentences from the text to be processed, such as extracting low-quality sentences, high-quality sentences, general sentences, sentences that express the central idea of ​​the text to be processed, etc. from the text to be processed, the preset segmentation granularity can be sentences. If the task requirement is to extract entities from the text to be processed, such as extracting names of people, places, organization names, etc. from the text to be processed, the preset segmentation granularity can be vocabulary.

[0084] In the embodiments of this specification, a specific scheme for segmenting the text to be processed and obtaining a plurality of segmentation positions is also provided and described.

[0085] Optionally, segmenting the text to be processed to obtain a plurality of segmentation positions may specifically include:

[0086] The text to be processed is segmented to obtain a plurality of sub-texts.

[0087] The head or tail position of each subtext is determined as the segmentation position.

[0088] The inserting of index codes for distinguishing different subtexts at the plurality of segmentation positions to obtain encoded texts may specifically include:

[0089] For each subtext among the plurality of subtexts, an index code is inserted at the segmentation position of each subtext to obtain each encoded subtext.

[0090] The encoded sub-texts are concatenated in sequence to obtain the encoded text.

[0091] In the embodiments of this specification, for the convenience of distinguishing each sub - text, the splitting position may include the position between two sub - texts. Specifically, the splitting position may be the head or the tail of each sub - text.

[0092] In practical applications, the splitting position may also be at other preset positions in the sub - text. For example, the splitting position may be after the first or several characters in the sub - text.

[0093] In the embodiments of this specification, each encoded sub - text may include a sub - text and an index code of this sub - text. The encoded text may be a text obtained by splicing each encoded sub - text in the front - to - back order of the sub - texts in the text to be processed.

[0094] In the embodiments of this specification, optionally, the index code may include characters in at least one form of numbers, letters, and symbols.

[0095] In practical applications, the index code may be numbers in any form, such as ①②③④⑤⑥⑦⑧⑨⑩ etc. or 12345678910 etc.; it may also be letters in any form, such as abcdefg etc., or ABCDEFG etc.; when the index code is a number or a letter, the index code may be continuous or discontinuous, and no specific limitation is made here. In addition, the index code may also be a symbol, such as!@&* / etc.

[0096] In practical applications, the index code may also be a combination of numbers and letters, or numbers and symbols, or letters and symbols, or a combination of numbers, letters, and symbols, and no specific limitation is made here.

[0097] In the embodiments of this specification, optionally, the index code includes at least one number and at least one symbol; different numbers and the same symbols may be included in the index codes corresponding to different sub - texts.

[0098] In the embodiments of this specification, the form of the numbers included in the index code may be in any form. Taking the number 1 as an example, the number 1 may be in the following forms: 1, ①, one. The symbol of the index code may be any symbol, such as the symbols "丨", "#", "*", etc. Among them, the symbols of the index codes of different sub - texts may be the same, but the numbers are different.

[0099] In the embodiments of this specification, the symbols included in the index codes corresponding to different sub - texts may be the same. Using symbols such as "丨", "#", "*" as delimiters can clearly separate different sub - texts, which can help the text extraction model better determine the sub - texts, avoid the text extraction model confusing different sub - texts, and thus improve the accuracy of the extracted text content.

[0100] As an implementation manner, the index coding may include numbers and the symbol "丨". "|" is the only symbol on the keyboard that has no meaning in written language. Using "number|" as the index coding can greatly reduce the possibility of the index coding appearing in the text to be processed, thereby improving the convenience of coding.

[0101] As an example, assume the text to be processed is as follows:

[0102] "Title: A shares are active, and B shares have a daily limit

[0103] Text: Go to the App to listen to the voice broadcast

[0104] Stock market news 10 minutes ago, on the morning of July 1st, A shares were active, B shares had a daily limit, and C shares and D shares rose more than 9%. Open the Jiemian News APP to view the original text

[0105] Jiemian Express The Jiemian editor has published 12,801 high-quality contents. Go and have a look.

[0106] Open Jiemian News to view more professional reports."

[0107] First, it is necessary to extract the text content of the body part. Assume that the sentence is used as the preset segmentation granularity for the above text to be processed, and several sub-texts can be obtained. For example, the above text can be segmented according to the paragraph number or the period as the identifier to obtain multiple sub-texts.

[0108] For example:

[0109] Open the App More information is waiting for you.

[0110] 10 minutes ago, on the morning of July 1st, A shares were active, B shares had a daily limit, and C shares and D shares rose more than 9%.

[0111] Open the M News APP to view more news. There are 239,586 high-quality contents in News N. Hurry up and have a look. Open M News to read more professional and authoritative news reports.

[0112] Then, insert the index coding at the beginning of each sub-text. The encoded sub-texts obtained can be as follows:

[0113] 1|Open the App More information is waiting for you.

[0114] 2|10 minutes ago, on the morning of July 1st, A shares were active, B shares had a daily limit, and C shares and D shares rose more than 9%.

[0115] 3|Open the M News APP to view more news. There are 239,586 high-quality contents in News N. Hurry up and have a look.

[0116] 4|Open M News to read more professional and authoritative news reports.

[0117] The encoded sub-texts may also be concatenated according to the order of the sub-texts in the text to be processed to obtain the encoded text, and the obtained encoded text may be as follows:

[0118] 1|Open the App for more information. 2|10 minutes ago on the morning of July 1, A shares were active, B shares hit the daily limit, C shares and D shares rose by more than 9%. 3|Open the M News APP to view more news. N Express has 239,586 high-quality content. Hurry up and read it. 4|Open M News to read more professional and authoritative news reports.

[0119] In practical applications, the relationship between words and sentences, words and paragraphs, and sentences and paragraphs in the text to be processed can also be considered. According to the relationship between words and sentences, words and paragraphs, and sentences and paragraphs in the text to be processed, each encoded sub-text is spliced ​​according to the order of each text in the text to be processed to obtain the encoded text. For example:

[0120] Title: A shares are active, B shares hit the daily limit

[0121] 1|Open the App for more information.

[0122] 2|10 minutes ago On the morning of July 1, A shares were active, B shares hit the daily limit, and C shares and D shares rose by more than 9%.

[0123] 3|Open M News APP to view more news. N Express has 239,586 high-quality content. Check it out. 4|Open M News to read more professional and authoritative news reports.

[0124] In practical applications, in the process of segmenting the text to be processed according to the preset segmentation granularity, it is not necessary to actually segment each sub-text, as long as each segmentation position can be obtained. The index code can be directly inserted at the determined segmentation position to obtain the encoded text. It is not necessary to obtain each sub-text, or it is also necessary to splice each sub-text again. In practical applications, if the segmentation position is the head of the text obtained after each segmentation (such as the sentence obtained after segmentation), the index code can be inserted into the head of the text obtained after the first segmentation and each segmentation position to obtain the encoded text; if the segmentation position is the tail of the text obtained after each segmentation, the index code can be inserted into the tail of the text obtained after the last segmentation and each segmentation position to obtain the encoded text.

[0125] In the embodiment of this specification, the index code corresponding to each sub-text may also be determined based on each sub-text.

[0126] Optionally, the method may further include:

[0127] By using a preset function, the sub-texts obtained by segmenting the text to be processed are subjected to function calculation to obtain character strings corresponding to the sub-texts.

[0128] The character string is determined as the index code corresponding to each subtext.

[0129] In the embodiments of the present specification, the preset function may be a hash function, and the hash function may be used to perform hash calculation on each sub-text. Specifically, the MD5 algorithm or the SHA algorithm may be used to perform hash calculation on each sub-text to obtain the hash value of each sub-text, and the hash value corresponding to each sub-text is used as the index code of each sub-text.

[0130] As an example, assuming that each subtext is a sentence, a hash function can be used to perform hash calculations on each sentence to obtain a hash value of each sentence, and the hash value corresponding to each sentence can be used as the index code of each sentence. In actual applications, if the text to be processed includes multiple identical sentences, a salted hash algorithm can be used to perform hash calculations on each sentence to make the index code of each subtext different.

[0131] For ease of understanding, the embodiments of this specification also provide specific content for determining the target subtext corresponding to the target index code.

[0132] Optionally, determining the target subtext corresponding to the target index code in the to-be-processed text according to the target index code may specifically include:

[0133] According to the target index code, a target subtext corresponding to the target index code is extracted from the encoded text.

[0134] Alternatively, the target subtext corresponding to the target index code is determined according to the target index code and the corresponding relationship between the index code and the subtext.

[0135] In the embodiment of this specification, the target subtext is a subtext in the text to be processed, and the target subtext is included in the text to be processed.

[0136] As an implementation method, the target subtext can be obtained based on the encoded text. After the text extraction model outputs the target index number, the server can extract the subtext corresponding to the target index code from the encoded text to obtain the target subtext based on the subtext corresponding to the target index code extracted from the encoded text. Specifically, the server can use a matching algorithm to match the target index code with the index code in the encoded text, and then obtain the target subtext based on the subtext of the index code that matches the target index code.

[0137] As another implementation, after the text to be processed is segmented, a correspondence between each sub-text obtained by segmentation and the corresponding index code can also be established. One sub-text can correspond to one index code, and different sub-texts correspond to different index codes. In the embodiment of this specification, the target index code can be one or more index codes. After the correspondence between each sub-text and the corresponding index code is established, the target sub-text corresponding to the target index code can be determined based on the correspondence between each sub-text and the corresponding index code.

[0138] In practical applications, the text to be processed may contain text in the same format as the index code. For example, the text to be processed may include the text "2|", while the index code is in the format of "1|", "2|", "3|", which may cause confusion. In order to avoid confusion between the index code and the text in the text to be processed, in the embodiments of this specification, further processing may be performed on the text to be processed.

[0139] Optionally, the method may further include:

[0140] If the text to be processed includes text in the same format as the index code, the text in the same format as the index code included in the text to be processed is escaped to obtain an escaped text;

[0141] The segmentation of the to-be-processed text may specifically include:

[0142] The escaped text is segmented.

[0143] In practical applications, the format of the index code may be fixed or known, or the index code may be set according to the text to be processed. For different texts to be processed, the format of the index code may be the same or different. The specific content and format of the index code are not limited here.

[0144] In the embodiments of the present specification, after determining the format of the index code used to encode the text to be processed, the text to be processed can also be escaped according to the determined index code or index code format. Specifically, if the text to be processed includes text in the same format as the index code, the text in the text to be processed in the same format as the index code can be processed, so that the text contained in the processed text to be processed is in a different format from the index code.

[0145] Among them, the text in the same format as the index code may include text with the same text content as the index code; for example, assuming that the index codes are "1|", "2|", "3|", if the text to be processed includes the text "2|", then the text "2|" is the same text as the index code text. Alternatively, the text in the same format as the index code may also include text with different text content but the same text structure as the index code, such as assuming that the index codes are "1|", "2|", "3|", if the text to be processed includes the text "5|", then "5|" is the text in the same format as the index code. In the embodiments of the present specification, the text in the text to be processed that is in the same format as the index code can be escaped so that the format of the text contained in the text to be processed is different from that of the index code.

[0146] Among them, escape processing can refer to converting special characters in the text into a format that can be safely parsed. In the embodiment of this specification, the escape processing for the text to be processed can represent the processing of converting at least part of the text in the text to be processed into another text format. Specifically, it can be to convert the text in the same format as the index code in the text to be processed into another form. For example, one or more specific characters can be added at the preset position of the text in the same format as the index code. Among them, the preset position can be set according to actual needs. Specifically, the preset position can be after the first character of the text in the same format as the index code, or after the second character, which is not specifically limited here. The specific character can be a character different from the character included in the index code, such as "\". For example, one or more "\" can be added after the first character of the text "2|" in the same format as the index code, and the escaped text "2|" can be "2\\|".

[0147] In the embodiment of this specification, the text to be processed after the escape process does not include text in the same format as the index code.

[0148] In practical applications, before performing the escape process, it can be determined based on the format of the index code whether the text to be processed contains text in the same format as the index code. If the determination result indicates that the text to be processed contains text in the same format as the index code, the escape process can be performed on the text to be processed that is in the same format as the index code. If the determination result indicates that the text to be processed does not contain text in the same format as the index code, there is no need to perform the escape process on the text to be processed.

[0149] As an implementation method, it is also possible not to determine whether the text to be processed contains text in the same format as the index code. Regardless of whether the text to be processed contains text in the same format as the index code, the escape processing step is performed, so that the escaped text obtained does not include text in the same format as the index code.

[0150] In the embodiment of this specification, the text to be processed can be firstly escaped, and then the escaped text obtained after the escape is segmented to obtain several subtexts. As another implementation, the text to be processed can be segmented to obtain several subtexts, and then each subtext can be escaped.

[0151] Optionally, after obtaining the plurality of subtexts, the following may also be included:

[0152] For any subtext among the several subtexts, if any subtext includes text in the same format as the index code, the text in the same format as the index code included in the any subtext is escaped to obtain the escaped subtext.

[0153] In the embodiment of the present description, before performing the escape processing, for any sub-text among several sub-texts, it can be determined whether the sub-text contains text in the same format as the index code. If the determination result indicates that the sub-text contains sub-text in the same format as the index code, the escape processing can be performed on the text in the same format as the index code contained in the sub-text, so that the escaped sub-text does not contain text in the same format as the index code.

[0154] In actual applications, it is not necessary to determine whether each sub-text contains text in the same format as the index code. Regardless of whether each sub-text contains text in the same format as the index code, the escape processing step is performed on each sub-text, so that the escaped sub-text does not contain text in the same format as the index code.

[0155] In the embodiment of the specification, the index code can be inserted at the segmentation position of the escaped subtext to obtain the encoded subtext; the encoded subtexts can also be spliced ​​according to the order of the subtexts in the text to be processed to obtain the encoded text. In this embodiment, the encoded text can include the subtext that has been escaped.

[0156] As an implementation method, after the index code is inserted into the segmentation position of each subtext, each encoded subtext is escaped to obtain the escaped encoded subtext. It can be understood that the encoded subtext may include the subtext part and the index code part.

[0157] Optionally, after obtaining the plurality of encoded subtexts, the following steps may also be included:

[0158] If any encoded subtext other than the index code includes text in the same format as the index code, the text in the same format as the index code included in the subtext other than the index code in any encoded subtext can be escaped to obtain the escaped encoded subtext.

[0159] As an implementation mode, before performing the escape processing, for any content other than the index code in any encoded subtext, it can be first determined whether the content other than the index code in the encoded subtext contains text in the same format as the index code. If the determination result indicates that the content other than the index code in the encoded subtext contains text in the same format as the index code, the escape processing can be performed on the text in the same format as the index code contained in the content other than the index code in the encoded subtext. If the determination result indicates that the content other than the index code in the encoded subtext does not contain text in the same format as the index code, it is not necessary to perform the escape processing on the above-mentioned other content in the encoded subtext.

[0160] In actual applications, it is also possible not to determine whether the other contents in each encoded sub-text except the index code contain text in the same format as the index code. Regardless of whether the other contents contained in each encoded sub-text contain text in the same format as the index code, escape processing is performed on the other contents contained in each encoded sub-text to obtain each escaped encoded sub-text, so that the obtained escaped encoded sub-text does not contain text in the same format as the index code in the other contents except the index code.

[0161] In the embodiments of this specification, the execution order of escape processing is not limited. The escape processing can be performed before the text to be processed is segmented; it can also be performed after the text to be processed is segmented to obtain several sub-texts; it can also be performed after the index code is inserted into the segmentation position to obtain each encoded sub-text; it only needs to be performed before the encoded text is input into the text extraction model.

[0162] In an embodiment of the present specification, if the text to be processed or the sub-text obtained by segmenting the text to be processed is subjected to escape processing, after the target index code is determined using the text extraction model, the sub-text corresponding to the target index code can also be subjected to reverse escape processing, and the target sub-text corresponding to the target index code in the text to be processed can be determined.

[0163] Optionally, determining the target subtext corresponding to the target index code in the to-be-processed text according to the target index code may specifically include:

[0164] An initial subtext corresponding to the target index code is extracted from the encoded text.

[0165] The initial subtext is subjected to a reverse semantics process to obtain the target subtext.

[0166] In the embodiment of this specification, the initial subtext may be the subtext corresponding to the target index code in the encoded text, and the initial subtext may be the subtext after escape processing. The target subtext is the subtext contained in the text to be processed. Escape processing may refer to the process of converting a string containing escape characters into its original form. In the embodiment of this specification, performing escape processing on the initial subtext may refer to converting the format of the escaped text in the initial subtext into the format of the corresponding text in the original text to be processed.

[0167] In practical applications, after the target index code is determined, the initial subtext can be extracted from the encoded text according to the target index code. It can be understood that one target index code corresponds to one initial subtext; if the target index code is multiple, the initial subtext is multiple; if the target index code is one, the initial subtext is one. After obtaining the initial subtext, the initial subtext can be reversed to obtain the target subtext. It can be understood that if the initial subtext includes escaped text, such as the text "2\\|", then reversed text can be performed on the initial subtext to obtain text with the same format as the text in the text to be processed, such as "2|". If the initial subtext does not include escaped text, that is, the initial subtext has the same format as the subtext contained in the text to be processed, then reversed text will not change the format of the text in the initial subtext.

[0168] In addition, after determining the target index code, it is also possible to first determine whether the initial sub-text has been escaped. If so, the escaped text included in the initial sub-text can be un-escaped to obtain the target sub-text; if not, the initial sub-text can be determined as the target sub-text without un-escaping it.

[0169] As an implementation method, after obtaining the target index code, the target subtext corresponding to the target index code can be determined according to the correspondence between the index code and each subtext in the text to be processed, without the need to perform a reverse semantics process on the initial subtext.

[0170] In practical applications, the encoded text may be firstly unsequenced to obtain the unsequenced encoded text, and the target sub-text may be determined based on the unsequenced encoded text.

[0171] Specifically, the index code part and the sub-text part in the encoded text can be marked in advance to distinguish the index code part from the sub-text part; the marked encoded sub-text can be reversed to obtain the reversed encoded text; according to the target index code, the target sub-text can be extracted from the reversed encoded text.

[0172] As another implementation, in order to avoid confusion between the index code and the characters in the text to be processed, in the embodiments of this specification, the format of the index code can also be selected based on the format of the text to be processed.

[0173] Optionally, the method may further include:

[0174] Determine whether the text to be processed contains numbers.

[0175] If the text to be processed contains numbers, a first number format to which the numbers belong is determined.

[0176] The index code is generated according to a second digital format; the second digital format is different from the first digital format.

[0177] In practical applications, numbers can usually be used as index codes. If the text to be processed also contains numbers, the index code may be confused with the numbers contained in the text to be processed, and it is impossible to distinguish whether the number is an index code or a number originally contained in the text to be processed. Based on this, before generating the index code, it can be determined whether the text to be processed contains numbers. If the text to be processed contains numbers, the digital format of the numbers can be further determined to facilitate the generation of an index code with a digital format different from the numbers in the text to be processed. For example, if the digital format of the numbers contained in the text to be processed is circled, such as ①②③④⑤⑥⑦⑧⑨⑩. Then the digital format of the index code can be without circles, such as 12345678910. It can be understood that if the text to be processed contains multiple numbers, and the multiple numbers are in different digital formats, the index code can be generated according to a digital format different from the multiple digital formats in the text to be processed.

[0178] As an implementation mode, if the text to be processed does not contain numbers, numbers in any digital format can be used as index codes as needed.

[0179] In practical applications, letters can also be used as index codes. In order to avoid confusion between the index codes and the letters in the text to be processed, before generating the index codes, it is possible to first determine whether the text to be processed contains letters. If the text to be processed contains letters, the form of the letters can be further determined to facilitate the generation of index codes that are different from the form of the letters in the text to be processed. For example, if the form of the letters contained in the text to be processed is uppercase, lowercase letters can be used as index codes. For another example, if the form of the letters contained in the text to be processed is a single letter, such as A, multiple superimposed and repeated letters, such as AA and AAA, can be used as index codes. The index codes of different sub-texts can be different superimposed and repeated letters, such as the index code of the first sub-text can be AA, and the index code of the second sub-text can be BB.

[0180] In practical applications, prompt words are often used to guide large models to generate specific responses.

[0181] Based on this, in the embodiments of this specification, optionally, the step of inputting the encoded text into a text extraction model may specifically include:

[0182] Get the prompt word template containing task requirement information.

[0183] The encoded text is added to the prompt word template to obtain a prompt word containing the encoded text.

[0184] The prompt word is provided to the text extraction model.

[0185] In the embodiments of the present specification, the task requirement information may be information related to the task that the text extraction model needs to perform. The prompt word template may be a structured text for generating prompt words, and prompt words may be generated based on the prompt word template. The prompt words may be information that instructs the text extraction model to generate relevant content according to the task requirement information.

[0186] In actual application, the prompt word template can be obtained first, and then the task requirement information can be inserted into the prompt word template, so as to obtain the prompt word template containing the task requirement information. In addition, the task requirement information can also be the content already in the prompt word template.

[0187] In the embodiment of this specification, a prompt word template can be obtained first, and then the encoded text is inserted into the prompt word template to obtain the prompt word. The prompt word can guide the text extraction model to better understand the task requirements, thereby facilitating the text extraction model to generate and output content that better meets the task requirements.

[0188] For ease of understanding, the task requirement information is specifically described in the embodiments of this specification.

[0189] Optionally, the task requirement information may include at least one of information for representing a model role, information for representing a task introduction, and information for representing a task analysis idea.

[0190] The information representing the model role is a key element in the formation of the prompt word. In the embodiment of this specification, the information representing the model role may refer to the profession or identity that the user wants the text extraction model to play when generating text or performing tasks. The information representing the model role can enable the text extraction model to better simulate the behavior and language habits of a specific profession or identity, thereby enhancing the professionalism and credibility of the content.

[0191] In practical applications, the role of the model can be specified through clear sentences. For example, a sentence like "You are an editor who is responsible for reviewing the quality of articles" not only provides role information for the text extraction model, but also provides the corresponding background for the text extraction model.

[0192] The information representing the task introduction is usually the core part of the prompt word, which can clearly and explicitly point out the specific task that the user wants the text extraction model to complete. It should be noted that the information representing the task introduction usually needs to be brief and clear, avoiding the use of lengthy or complex sentences to reduce the understanding cost of the text extraction model; in addition, the information representing the task introduction also needs to be specific and clear, avoiding the use of vague or ambiguous sentences.

[0193] Information indicating task analysis ideas can indicate information about task analysis ideas. Specifically, it can refer to instructions that can guide the text extraction model to think and analyze in a specific logical order or method. Information indicating task analysis ideas can help the model better understand task requirements, thereby generating more accurate and targeted content.

[0194] For ease of understanding, this specification provides an example of a prompt word containing task requirement information, as shown below:

[0195] You are an editor reviewing the quality of an article. The "number|" in front of each sentence in the article is the numerical number of the sentence fragment. Please output the article quality labels in order.

[0196] 1. First determine whether the article *main body* is *nonsense* or *dirty data*. If so, output the label and the review is completed.

[0197] 2. If not, further determine the *advertising*, *worthless*, and *nonsense* sentence fragments in the *main text*, and output the digital numbers or number ranges of the corresponding sentence fragments.

[0198] 3. Then determine the relevance between the title and the text. If they are completely unrelated, output “title and text are irrelevant”.

[0199] 4. Finally, determine whether the *title* and *text* are missing place names, company names, personal names, product names and other main bodies. If missing, output *Main body missing [title]* or *Main body missing [title, text]*.

[0200] 5. If there are no problems mentioned above, the output will be *high quality*.

[0201] Please review the following article. \nArticle *Title*: {title}\nArticle *Body*: {content}\nArticle quality label:

[0202] \nPlease review the following article. \nArticle *Title*: First-class Highway Practical Question Bank: Rich Questions to Meet Candidates with Different Needs\nArticle; *Text*: 1|As one of the required subjects for the first-class construction engineer examination, highway practice occupies an important position in the examination. 2|In order to successfully pass the first-class highway practical subject examination, in addition to systematically learning relevant knowledge, doing questions is also an indispensable link. 3|In order to help the majority of candidates prepare for the exam better, we have specially launched the first-class highway practical question bank, aiming to provide rich questions to meet the needs of candidates with different needs. 4|The first-class highway practical question bank contains a large number of real test questions, covering all aspects of highway engineering, such as engineering measurement, engineering construction, engineering quality, etc. 5|These questions not only test basic knowledge, but also test practical application ability, which can comprehensively test the candidates' mastery and test-taking ability. n6|In the question bank, we have classified the questions according to the difficulty of the questions and the importance of the test points. Candidates can choose the questions that suit them according to their actual situation to practice. 7|At the same time, we also provide detailed analysis and answers. Candidates can check for omissions and improve their problem-solving ability by comparing their own answers. 8|In order to meet the needs of different candidates, our question bank also provides a variety of practice modes. 9|Candidates can choose the sequential practice mode to practice in the order of the questions; they can also choose the random practice mode to randomly select questions for practice; they can also choose the wrong question practice mode to focus on the questions they have done wrong for intensive training. 10|These different practice modes can help candidates better consolidate their knowledge and improve their problem-solving ability. 11|In addition to providing a wealth of questions and a variety of practice modes, our question bank also provides real-time exam dynamics and score analysis. 12|Candidates can view the latest exam information through the question bank to understand the difficulty and changing trends of the exam; at the same time, they can also view their scores and rankings, adjust their study plans in time, and improve their preparation efficiency. 13|In short, the first-level construction highway practical real question bank is a good helper for the majority of candidates to prepare for the exam. 14|Through systematic practice and targeted training, candidates can better master knowledge, improve their problem-solving ability, and lay a solid foundation for successfully passing the first-level construction engineer examination. 15|It is recommended to use the mobile and online question-answering software such as the Global Online School First-Class Construction Engineer Quick Question Bank and the First-Class Construction Engineer Holy Question Bank, the First-Class Construction Engineer Excellent Question Bank, and the First-Class Construction Question Bank. 16|These software provide a wealth of questions and exam simulations, which can help candidates prepare for the exam better. 17|I especially recommend the Global Online School First-Class Construction Engineer Quick Question Bank, which not only provides a large number of real questions and simulation questions, but also detailed analysis and answers, which can help candidates improve their problem-solving ability in an all-round way. \n18|Whether you want to review knowledge systematically or want to conduct targeted training, the First-Class Construction Highway Practical Question Bank can meet your needs. 19|I believe that through our efforts and your persistence, you will be able to successfully pass the first-class construction engineer exam and achieve excellent results!

[0203] In the above example, "You are an editor and are reviewing the quality of the article. The "number|" in front of each sentence in the article is the digital number of the sentence fragment. Please output the article quality label in order. 1. First determine whether the *main body* of the article is *nonsense* or *dirty data*. If so, output the label and the review is complete. 2. If not, further determine the *advertising*, *worthless*, and *nonsense* sentence fragments in the *main body*, and output the digital number or number range of the corresponding sentence fragments. 3. Then determine the relevance of the *title* and *main body*. If they are completely irrelevant, output *irrelevant to the title and text*. 4. Finally, determine whether the *title* and *main body* are missing place names, company names, personal names, product names and other subjects. If missing, output *subject missing [title]* or *subject missing [title, text]*. 5. If there are no above problems, output *high quality*." indicates the task requirement information.

[0204] Among them, "You are an editor who is reviewing the quality of the article" can be information representing the model role.

[0205] "The 'number|' in front of each sentence in the text is the numerical number of the sentence fragment. Please output the article quality label in order" can be information representing the task introduction.

[0206] "1. First determine whether the *main body* of the article is *nonsense* or *dirty data*. If so, output the label and the review is completed. 2. If not, further determine the *advertising*, *worthless* and *nonsense* sentence fragments in the *main body*, and output the numerical numbers or number ranges of the corresponding sentence fragments. 3. Then determine the relevance of the *title* and *main body*. If they are completely irrelevant, output *title and text are irrelevant*. 4. Finally, determine whether the *title* and *main body* are missing the main body such as place names, company names, personal names, product names, etc. If missing, output *main body missing [title]* or *main body missing [title, text]*. 5. If there are no above problems, output *high quality*." can be information that represents the task analysis ideas.

[0207] The above content is only used as an illustration of prompt words, and different prompt words can be set according to different task requirements.

[0208] In practical applications, the prompt word may also include one or more reference examples to improve the accuracy of the text extraction model output, wherein the reference examples may include sample texts and the results of analyzing the sample texts according to the information representing the task analysis ideas.

[0209] In the embodiments of this specification, after the prompt words are obtained, sequence annotation can be performed on the prompt words, so as to convert the prompt words into a format that can be understood by the machine, so that the text output model can understand the natural language.

[0210] If the task requirement is to extract high-quality sentences from the article, the low-quality sentences can also be deleted in the embodiment of this specification to obtain high-quality text for business use, such as inputting high-quality sentences for users, analyzing text summaries, etc. Optionally, if the task requirement is to determine that the text to be processed contains low-quality sentences; the method can also include:

[0211] The target subtext is deleted from the text to be processed to obtain a deleted text.

[0212] In the embodiment of the present specification, the deleted text may be the text after deleting the target subtext corresponding to the target index code in the text to be processed.

[0213] As an implementation method, continue to use the example in the previous article, assuming that the encoded text obtained by processing the text to be processed is as follows:

[0214] 1|Open the App for more information.

[0215] 2|10 minutes ago On the morning of July 1, A shares were active, B shares hit the daily limit, and C shares and D shares rose by more than 9%.

[0216] 3|Open the M News APP to view more news. N Express has 239,586 high-quality contents. Come and read it.

[0217] 4|Open M News to read more professional and authoritative news reports.

[0218] If the target index codes output by the text extraction model are "[1|]", "[3|-4|]", it can be indicated that the target subtext "Open App for more information" corresponding to the target index code 1|, the target subtext "Open M News APP to view more news N Express has 239,586 high-quality content, hurry up and read" corresponding to the target index code 3| are the texts that need to be deleted. After deleting the above target subtexts in the text to be processed, the deleted text can be obtained.

[0219] In practical applications, if the task requirement is to extract high-quality sentences from an article, high-quality sentences can also be directly extracted through the text extraction model.

[0220] Continuing with the example in the previous article, assuming that the target index number output by the text extraction model is "[2|]", it can be indicated that the target subtext corresponding to the target index code 2| "10 minutes ago on the morning of July 1, A shares were active, B shares hit the daily limit, and C shares and D shares rose by more than 9%" is a high-quality sentence.

[0221] In the embodiment of the present specification, optionally, if the task requirement is a task requirement indicating determination of a target type of text contained in the to-be-processed text; the target type of text includes at least one of a text indicating an abstract, a text indicating a topic sentence, a text indicating a summary sentence, and a text indicating an entity attribute; the method may further include:

[0222] The target subtext is output.

[0223] In the embodiment of this specification, if the task requirement is to use the text extraction model to determine the abstract, topic sentence, summary sentence, or entity in the text to be processed. After determining the target subtext corresponding to the target index code, the target subtext can be displayed. Specifically, the target subtext can be displayed on the terminal device.

[0224] In practical applications, in order to facilitate the use of text extraction models to perform text extraction tasks, the text extraction models can also be trained in advance.

[0225] Optionally, the method may further include:

[0226] Acquire several training samples; one of the training samples may include a training text with a training index code and annotation information corresponding to the training text; the annotation information represents the expected model output information, including the training index code corresponding to the sub-training text in the training text that meets the task requirements.

[0227] At least one of the training samples is added to a prompt word template corresponding to the task requirement to obtain a training prompt word.

[0228] The large model is trained using the training prompt words to obtain the text extraction model.

[0229] In the embodiments of this specification, the qwen2.5-1.5B model can be used as a basic model, and the training sample can be used to fine-tune the basic model to obtain a trained text extraction model. The training sample can include training text and annotation information.

[0230] The training text may include a training index code and a sub-training text; the training index code may be used to distinguish different sub-training texts. It is understandable that different sub-training texts have different training index codes.

[0231] The annotation information may be information output by the expected text extraction model. In an embodiment of the present specification, the annotation information may include a training index code that meets the task requirements. Specifically, if the task requirement is to determine the low-quality text in the text, if the entire text in the training text is of low quality, the annotation information may be a preset character, such as -1; if multiple consecutive sentences in the training text are of low quality, the annotation information may be an index code interval of low-quality sentences such as [3|-4|]. Using relatively simple preset characters, it is indicated that the entire text is of low quality; and using the index code interval to indicate that multiple consecutive sentences in the interval are of low quality, if the entire text to be processed is of low quality or multiple consecutive sentences are of low quality, the text extraction model does not need to generate index codes for each low-quality sub-text separately, and the index code interval can be used to indicate that multiple sentences are of low quality, or the preset characters can be used to indicate that the entire text to be processed is of low quality, which can further reduce the number of characters generated by the text extraction model and improve the efficiency of text extraction. In addition, if individual sentences in the text are of low quality, the annotation information may be the index code of the low-quality sentence. In practical applications, the annotation information may include the index coding interval and the index coding of a single sentence.

[0232] In addition, during the use of the text extraction model, after the text extraction model outputs the target index number, the target index number can be parsed. For example, if the target index number is -1, it means that all sub-texts are low-quality texts; if the target index number is the index number of an interval, it means that the sub-text corresponding to the index number in the interval is a low-quality text.

[0233] For ease of understanding, we will continue to use the previous example to illustrate. Assume that the task requires identifying low-quality text in the training text, and the training text is as follows:

[0234] 1|Open the App for more information.

[0235] 2|10 minutes ago On the morning of July 1, A shares were active, B shares hit the daily limit, and C shares and D shares rose by more than 9%.

[0236] 3|Open the M News APP to view more news. N Express has 239,586 high-quality contents. Come and read it.

[0237] 4|Open M News to read more professional and authoritative news reports.

[0238] The annotation information can be “[1|], [3|-4|]”.

[0239] According to the title and content of the above training text, the main content of the above training text is about the rise and fall of stocks, while "1|Open the App for more information." and "3|Open the M News APP to see more news. N Express has 239,586 high-quality content. Check it out. 4|Open M News to read more professional and authoritative news reports." are not related to stocks. These sentences can be identified as low-quality texts, and the annotation information can be [1|], [3|-4|].

[0240] As an implementation method, the annotation information may also include thought chain information. For example, the annotation information may be "worthless [1|], advertising traffic [3|-4|]"

[0241] Among them, the chain of thought information can be used to explain the specific reasons why the sub-training text corresponding to the training index code in the annotation information meets the task requirements. Including the chain of thought information in the annotation information can improve the interpretability of the determined training index code that meets the task requirements, and further improve the interpretability of the target index code determined by the trained text extraction model, so that users can understand the specific reasons for the final determination of the target sub-text.

[0242] For ease of understanding, the embodiments of this specification also provide a process for generating training samples.

[0243] Optionally, obtaining a number of training samples may specifically include:

[0244] Get some training texts.

[0245] For any training text among the plurality of training texts, segmentation processing is performed on the training text according to the preset segmentation granularity.

[0246] A training index code for distinguishing different sub-training texts is inserted at the segmentation position of any training text to obtain a training text with a training index code; the encoding form of the training index code is the same as the form of the index code contained in the encoded text.

[0247] In the embodiments of this specification, the training text can be obtained in various ways. As an implementation method, the training text can be obtained through a web crawler.

[0248] In the embodiment of this specification, the model for segmenting the training text and the model for segmenting the to-be-processed text in the foregoing text may be the same type of model. Further, the model for segmenting the training text and the model for segmenting the to-be-processed text in the foregoing text may be the same.

[0249] In order to improve the accuracy of the text extraction model obtained through training, in the embodiments of the present specification, the segmentation position of the training text and the segmentation position of the text to be processed can be the same, for example, they can both be the beginning of the sub-text obtained after segmentation, or they can both be the end of the sub-text obtained after segmentation.

[0250] In addition, in order to improve the accuracy of the trained text extraction model, in the embodiments of this specification, the encoding form of the training index encoding and the index encoding can be the same, for example, the training index encoding and the index encoding are both in the form of numbers plus symbols.

[0251] Figure 3 A swim lane diagram of a text extraction method provided in an embodiment of this specification. Figure 3 As shown, the text extraction process may involve execution entities such as terminal devices and servers, and the process may include text extraction model training and text extraction stages.

[0252] In the text extraction model training stage, the execution subject can be a server or a terminal device. The text extraction model training stage can specifically include the following steps:

[0253] Step 302: Obtain a number of training samples.

[0254] In an embodiment of the present specification, one of the training samples may include a training text having a training index code and annotation information corresponding to the training text; the annotation information represents the expected model output information, including the training index code corresponding to the sub-training text in the training text that meets the task requirements.

[0255] Step 304: adding at least one training sample among the training samples to a prompt word template corresponding to the task requirement to obtain a training prompt word.

[0256] In the embodiments of this specification, the prompt template may be a structured text for generating prompts, and training prompts may be generated based on the prompt template. The training prompts may be information for instructing the text extraction model to generate relevant content according to task requirement information.

[0257] Step 306: Use the training prompt words to train the large model to obtain a text extraction model.

[0258] In the embodiments of this specification, the qwen2.5-1.5B model can be used as a basic model, and the training samples can be used to fine-tune the above basic model to obtain a trained text extraction model.

[0259] In the text extraction stage, the execution entity may include a server. The server in the text extraction stage and the server in the text extraction model training stage may be the same server or different servers, and no specific limitation is made here.

[0260] The text extraction stage can specifically include the following steps:

[0261] Step 308: Obtain the text to be processed.

[0262] The text to be processed may be a text that needs to be processed, and may be a text that only includes a plurality of phrases, a text that includes a plurality of sentences, or a text that includes a plurality of paragraphs.

[0263] In practical applications, the text to be processed can be text obtained in various ways. As an example, the text to be processed can be text extracted from an existing file, data set, or database; as another example, the text to be processed can also be text generated using a large model, which is not limited in the embodiments of this specification.

[0264] Step 310: segment the text to be processed according to a preset segmentation granularity to obtain a plurality of sub-texts.

[0265] In the embodiments of the present specification, a text segmentation model that performs segmentation according to a preset segmentation granularity may be used to perform segmentation processing on the text to be processed.

[0266] The preset segmentation granularity may be a preset encoding granularity; the encoding granularity may refer to the size of the smallest unit into which the text data is segmented. In practical applications, the preset segmentation granularity includes any one of representing vocabulary, sentence, and paragraph. In specific applications, the preset segmentation granularity may be determined according to task requirements. It is understandable that the preset segmentation granularity may be different for different tasks.

[0267] In the embodiments of this specification, segmentation processing may refer to dividing a text into smaller units, such as words, vocabulary, sentences or paragraphs.

[0268] Step 312: Determine the head or tail position of each subtext as the segmentation position.

[0269] The segmentation position may be a preset position in a subtext obtained after segmentation of the text to be processed. Specifically, the head or tail position of each subtext may be determined as the segmentation position.

[0270] Step 314: inserting index codes for distinguishing different sub-texts at a number of segmentation positions to obtain encoded texts.

[0271] In the embodiments of this specification, for each of several sub-texts, index codes can be inserted at the segmentation positions of each sub-text to obtain each encoded sub-text; then, the encoded sub-texts are concatenated in the order in the text to be processed, so as to obtain the encoded text.

[0272] In the embodiments of this specification, the index code includes one or more forms of characters among numbers, letters, and symbols. As an implementation, the index code can include at least one number and at least one symbol; the different numbers and the same symbols can be included in the index codes corresponding to different sub-texts. Specifically, the symbol included in the index code can be "丨".

[0273] Step 316: Obtain a prompt template containing task requirement information.

[0274] Among them, the task requirement information can be the relevant requirement information for indicating the tasks that the text extraction model needs to execute. The prompt template can be a structured text for generating prompt words.

[0275] Step 318: Add the encoded text to the prompt template to obtain a prompt word containing the encoded text.

[0276] In the embodiments of this specification, the prompt template can be obtained first, and then the encoded text is inserted into the prompt template to obtain the prompt word. The prompt word can guide the text extraction model to better understand the task requirements, so as to facilitate the text extraction model to generate and output content more in line with the task requirements.

[0277] Step 320: Provide the prompt word to the text extraction model, and the text extraction model outputs the target index code.

[0278] In the embodiments of this specification, the text extraction model can be a model deployed in the server, or a model that the server calls from other servers or terminals.

[0279] The text extraction model can be a large language model and can be used to execute at least one of tasks such as Information Extraction (IE) and Machine Reading Comprehension (MRC). Specifically, the text extraction model can be a large language model corresponding to the task requirements.

[0280] In actual applications, the text extraction model can be selected according to the correspondence between the task requirements and the text extraction model.

[0281] In addition, the text extraction model can also be a model that can comprehensively handle various tasks. For example, the text extraction model can be a comprehensive model that can extract low-quality sentences, high-quality sentences, and entities.

[0282] In the embodiment of this specification, the target index code may be an index code of a sub-text that meets the task requirements and is determined by a text extraction model.

[0283] Step 322: According to the target index code, extract the target subtext corresponding to the target index code from the encoded text.

[0284] It is understandable that in actual applications, the text extraction model training phase and the text extraction phase can be performed separately. During the text extraction phase, the trained text extraction model can be directly used. That is, there is no need to perform the text extraction model training process every time text extraction is performed.

[0285] Based on the same idea, the embodiments of this specification also provide a device corresponding to the above method. Figure 4 This is a schematic diagram of the structure of a text extraction device provided in an embodiment of this specification. Figure 4 As shown, the device may include:

[0286] An acquisition module 402 is used to acquire the text to be processed;

[0287] The segmentation processing module 404 is used to segment the text to be processed according to a preset segmentation granularity to obtain a plurality of segmentation positions;

[0288] An encoding module 406 is used to insert index codes for distinguishing different subtexts at the plurality of segmentation positions to obtain encoded texts;

[0289] An input module 408 is used to input the encoded text into a text extraction model to obtain a target index code output by the text extraction model; the text extraction model is a pre-trained large model used to extract the index code of the sub-text that meets the task requirements from the input text;

[0290] The extraction module 410 is used to determine the target subtext corresponding to the target index code in the text to be processed according to the target index code.

[0291] based on Figure 4 The present specification also provides some specific implementation schemes of the method, which are described below.

[0292] Optionally, the segmentation processing module 404 may be specifically used for:

[0293] The text to be processed is input into a text segmentation model for segmentation according to the preset segmentation granularity.

[0294] Optionally, the preset segmentation granularity includes representing any one of vocabulary, sentence, and paragraph.

[0295] Optionally, the segmentation processing module 404 may be specifically used for:

[0296] The text to be processed is segmented to obtain a plurality of sub-texts.

[0297] The head or tail position of each subtext is determined as the segmentation position.

[0298] The encoding module 406 may be specifically used for:

[0299] For each subtext among the plurality of subtexts, an index code is inserted at the segmentation position of each subtext to obtain each encoded subtext.

[0300] The encoded sub-texts are concatenated in sequence to obtain the encoded text.

[0301] Optionally, the index code includes characters in at least one form of numbers, letters, and symbols.

[0302] Optionally, the index code includes at least one number and at least one symbol; the index codes corresponding to different subtexts contain different numbers and the same symbols.

[0303] Optional, Figure 4 The device may further include:

[0304] The calculation module is used to use a preset function to perform function calculation on each sub-text obtained by segmenting the text to be processed, so as to obtain a character string corresponding to each sub-text.

[0305] The first determination module is used to determine the character string as the index code corresponding to each subtext.

[0306] Optionally, the extraction module 410 may be specifically used for:

[0307] According to the target index code, a target subtext corresponding to the target index code is extracted from the encoded text.

[0308] Alternatively, the target subtext corresponding to the target index code is determined according to the target index code and the corresponding relationship between the index code and the subtext.

[0309] Optional, Figure 4 The device may further include:

[0310] The first escape processing module is used to, if the text to be processed includes text in the same format as the index code, perform escape processing on the text in the same format as the index code included in the text to be processed to obtain escaped text.

[0311] The segmentation processing module 404 may be specifically used for:

[0312] The escaped text is segmented.

[0313] Optional, Figure 4 The device may further include:

[0314] The second escape module is used to escape the text in the same format as the index code included in any sub-text among the several sub-texts to obtain the escaped sub-text.

[0315] Optionally, the extraction module 410 may be specifically used for:

[0316] An initial subtext corresponding to the target index code is extracted from the encoded text.

[0317] The initial subtext is subjected to a reverse semantics process to obtain the target subtext.

[0318] Optionally, the input module 408 may be specifically used for:

[0319] Get the prompt word template containing task requirement information.

[0320] The encoded text is added to the prompt word template to obtain a prompt word containing the encoded text.

[0321] The prompt word is provided to the text extraction model.

[0322] Optionally, the task requirement information includes at least one of information for representing a model role, information for representing a task introduction, and information for representing a task analysis idea.

[0323] Optionally, if the task requirement is a task requirement indicating that the text to be processed contains low-quality sentences; Figure 4 The device may also include:

[0324] The deletion module is used to delete the target subtext from the text to be processed to obtain a deleted text.

[0325] Optionally, if the task requirement is a task requirement for determining a target type of text contained in the to-be-processed text; the target type of text includes at least one of a text representing an abstract, a text representing a topic sentence, a text representing a summary sentence, and a text representing an entity attribute; Figure 4 The device may also include:

[0326] The target subtext is output.

[0327] Optional, Figure 4 The device may further include:

[0328] A training sample acquisition module is used to acquire a number of training samples; one of the training samples includes a training text with a training index code and annotation information corresponding to the training text; the annotation information represents the expected model output information, including the training index code corresponding to the sub-training text in the training text that meets the task requirements.

[0329] The prompt word acquisition module is used to add at least one training sample among the training samples to the prompt word template corresponding to the task requirement to obtain a training prompt word.

[0330] The training module is used to train the large model using the training prompt words to obtain the text extraction model.

[0331] Optional, training sample acquisition module, which can be used for:

[0332] Get some training texts.

[0333] For any training text among the plurality of training texts, segmentation processing is performed on the training text according to the preset segmentation granularity.

[0334] A training index code for distinguishing different sub-training texts is inserted at the segmentation position of any training text to obtain a training text with a training index code; the encoding form of the training index code is the same as the form of the index code contained in the encoded text.

[0335] Based on the same idea, the embodiments of this specification also provide a device corresponding to the above method.

[0336] Figure 5 This is a schematic diagram of the structure of a text extraction device provided in an embodiment of this specification. Figure 5 As shown, the device 500 may include:

[0337] at least one processor 510; and,

[0338] A memory 530 in communication with the at least one processor; wherein,

[0339] The memory 530 stores instructions 520 executable by the at least one processor 510. The instructions are executed by the at least one processor 510 to enable the at least one processor 510 to:

[0340] Get the text to be processed.

[0341] The text to be processed is segmented according to a preset segmentation granularity to obtain a plurality of segmentation positions.

[0342] Index codes for distinguishing different sub-texts are inserted at the plurality of segmentation positions to obtain encoded texts.

[0343] The encoded text is input into a text extraction model to obtain a target index code output by the text extraction model; the text extraction model is a pre-trained large model used to extract the index code of the sub-text that meets the task requirements from the input text.

[0344] According to the target index code, a target subtext corresponding to the target index code in the text to be processed is determined.

[0345] Based on the same idea, the embodiment of this specification also provides a computer-readable medium corresponding to the above method. The computer-readable medium stores computer-readable instructions, which can be executed by a processor to implement the above text extraction method:

[0346] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. Figure 5 As for the device shown, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0347] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards in the relevant regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0348] In the 1990s, improvements to a technology could be clearly distinguished as hardware improvements (for example, improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the method flow). However, with the development of technology, many improvements to the method flow today can be regarded as direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement in a method flow cannot be implemented using a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to ask a chip manufacturer to design and produce a dedicated integrated circuit chip. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages ​​and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.

[0349] The controller can be implemented in any appropriate manner, for example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (such as software or firmware) that can be executed by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in a purely computer-readable program code manner, the controller can be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, this controller can be considered as a hardware component, and the devices included therein for implementing various functions can also be regarded as structures within the hardware component. Or even, the devices for implementing various functions can be regarded as both software modules for implementing the method and structures within the hardware component.

[0350] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0351] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0352] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0353] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0354] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0355] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0356] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0357] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0358] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0359] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0360] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0361] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0362] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A text extraction method, comprising: Get the text to be processed; According to a preset segmentation granularity, the text to be processed is segmented to obtain a plurality of segmentation positions; Inserting index codes for distinguishing different subtexts at the plurality of segmentation positions to obtain encoded texts; Inputting the encoded text into a text extraction model to obtain a target index code output by the text extraction model; the text extraction model is a pre-trained large model used to extract the index code of the sub-text that meets the task requirements from the input text; According to the target index code, a target subtext corresponding to the target index code in the text to be processed is determined.

2. The method according to claim 1, wherein the segmentation of the to-be-processed text according to a preset segmentation granularity specifically comprises: The text to be processed is input into a text segmentation model for segmentation according to the preset segmentation granularity.

3. According to the method of claim 1, the preset segmentation granularity includes representing any one of vocabulary, sentence, and paragraph.

4. The method according to claim 1, wherein the segmenting of the text to be processed to obtain a plurality of segmentation positions specifically comprises: Segmenting the text to be processed to obtain a plurality of sub-texts; Determine the head or tail position of each subtext as the segmentation position; The inserting of index codes for distinguishing different sub-texts at the plurality of segmentation positions to obtain the encoded text specifically includes: For each subtext among the plurality of subtexts, insert an index code at the segmentation position of each subtext to obtain each encoded subtext; The encoded sub-texts are concatenated in sequence to obtain the encoded text.

5. The method as claimed in claim 1, wherein the index code includes characters in at least one form of numbers, letters, and symbols.

6. The method as claimed in claim 5, wherein the index code includes at least one number and at least one symbol; the index codes corresponding to different subtexts contain different numbers and the same symbols.

7. The method of claim 1, further comprising: Using a preset function, performing function calculation on each sub-text obtained by segmenting the text to be processed to obtain a character string corresponding to each sub-text; The character string is determined as the index code corresponding to each subtext.

8. The method according to claim 1, wherein determining the target subtext corresponding to the target index code in the to-be-processed text according to the target index code specifically comprises: According to the target index code, extracting a target subtext corresponding to the target index code from the encoded text; Alternatively, the target subtext corresponding to the target index code is determined according to the target index code and the corresponding relationship between the index code and the subtext.

9. The method of claim 1, further comprising: If the text to be processed includes text in the same format as the index code, the text in the same format as the index code included in the text to be processed is escaped to obtain an escaped text; The segmentation of the to-be-processed text specifically includes: The escaped text is segmented.

10. The method according to claim 4, after obtaining the plurality of subtexts, further comprising: For any subtext among the several subtexts, if any subtext includes text in the same format as the index code, the text in the same format as the index code included in any subtext is escaped to obtain an escaped subtext.

11. The method according to claim 9 or 10, wherein determining the target subtext corresponding to the target index code in the to-be-processed text according to the target index code specifically comprises: Extracting an initial subtext corresponding to the target index code from the encoded text; The initial subtext is subjected to a reverse semantics process to obtain the target subtext.

12. The method according to claim 1, wherein inputting the encoded text into a text extraction model comprises: Obtain a prompt word template containing task requirement information; Adding the encoded text to the prompt word template to obtain a prompt word containing the encoded text; The prompt word is provided to the text extraction model.

13. The method according to claim 12, wherein the task requirement information comprises at least one of information for representing a model role, information for representing a task introduction, and information for representing a task analysis idea.

14. The method according to claim 1, wherein if the task requirement is a task requirement indicating that the text to be processed contains low-quality sentences; the method further comprises: The target subtext is deleted from the text to be processed to obtain a deleted text.

15. The method according to claim 1, wherein if the task requirement is a task requirement for determining a target type of text contained in the to-be-processed text; the target type of text includes at least one of a text representing an abstract, a text representing a topic sentence, a text representing a summary sentence, and a text representing an entity attribute; the method further comprises: The target subtext is output.

16. The method of claim 1, further comprising: Acquire a plurality of training samples; one of the training samples includes a training text having a training index code and annotation information corresponding to the training text; The annotation information represents the expected model output information, including the training index codes corresponding to the sub-training texts in the training text that meet the task requirements; adding at least one of the training samples to a prompt word template corresponding to the task requirement to obtain a training prompt word; The large model is trained using the training prompt words to obtain the text extraction model.

17. The method according to claim 16, wherein obtaining a plurality of training samples comprises: Acquire a plurality of training texts; and for any training text among the plurality of training texts, segment the any training text according to the preset segmentation granularity; A training index code for distinguishing different sub-training texts is inserted at the segmentation position of any training text to obtain a training text with a training index code; the encoding form of the training index code is the same as the form of the index code contained in the encoded text.

18. A text extraction device, comprising: An acquisition module is used to obtain the text to be processed; A segmentation processing module is used to segment the text to be processed according to a preset segmentation granularity to obtain a plurality of segmentation positions; An encoding module, used for inserting index codes for distinguishing different subtexts at the plurality of segmentation positions to obtain encoded texts; An input module, used to input the encoded text into a text extraction model to obtain a target index code output by the text extraction model; The text extraction model is a pre-trained large model, which is used to extract the index encoding of the sub-text that meets the task requirements from the input text; The extraction module is used to determine the target subtext corresponding to the target index code in the text to be processed according to the target index code.

19. A text extraction device, comprising: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to: Get the text to be processed; According to a preset segmentation granularity, the text to be processed is segmented to obtain a plurality of segmentation positions; Inserting index codes for distinguishing different subtexts at the plurality of segmentation positions to obtain encoded texts; Inputting the encoded text into a text extraction model to obtain a target index code output by the text extraction model; the text extraction model is a pre-trained large model used to extract the index code of the sub-text that meets the task requirements from the input text; According to the target index code, a target subtext corresponding to the target index code in the text to be processed is determined.

20. A computer-readable medium having computer-readable instructions stored thereon, wherein the computer-readable instructions can be executed by a processor to implement the text extraction method according to any one of claims 1 to 17.