Text processing method, text generation method, device, medium and program product

By performing hierarchical analysis of text and threshold segmentation of word elements, the problem of information loss in text splitting is solved, and the information integrity and query retrieval effect of text blocks are improved.

CN120387446APending Publication Date: 2025-07-29KE COM (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510353380.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing text splitting method cannot effectively retain the document's structural information, resulting in incomplete text block information after splitting, affecting the subsequent query retrieval effect.

Method used

By performing hierarchical analysis of the target text, identifying the parent-child relationship between hierarchical elements, and dividing it according to the number of word elements and the number of word elements of hierarchical elements, ensuring that the text content of the same hierarchical element is in the same text block to avoid information loss.

Benefits of technology

The structure information of the document is retained, the information integrity of the text block is ensured, and the effect of subsequent query retrieval is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387446A_ABST
    Figure CN120387446A_ABST
Patent Text Reader

Abstract

The invention provides a text processing method, a text generation method, an electronic device, a readable storage medium and a computer program product.The text processing method comprises the steps that firstly, a target text is analyzed to obtain hierarchical information of the target text, and the hierarchical information is used for describing the set membership between hierarchical elements of the target text; then determining the number of text lexical elements contained in the plurality of hierarchical elements in the target text, then segmenting the target text according to a specified lexical element number threshold and the number of lexical elements of the hierarchical elements to obtain a plurality of text blocks, and connecting the positions of the texts in the same text block obtained through segmentation in the target text; and for the target hierarchy element with the lowest hierarchy in the belonging highest hierarchy elements, the text contents belonging to the same target hierarchy element are segmented into the same text block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical fields such as computers, and specifically relates to a text processing method, a text generation method, an electronic device, a readable storage medium, and a computer program product. Background Art

[0002] In some scenarios, in order to better perform text processing and analysis, the text can be first split, and the split text blocks can be used to implement functions such as building a knowledge base.

[0003] Currently, there are two common ways to split text. One is to split the text according to a fixed text length, and the other is to split the text according to the smallest parsable element. However, since the text may have a clear structure, such as the text in a document has structures such as headings, tables of contents, chapters, and tables, both splitting by a fixed length and splitting by the smallest parsable element have the defect that the text blocks obtained after splitting lack text structure information, resulting in incomplete text block information. This makes the subsequent text processing such as querying and retrieving using the text blocks less effective. Summary of the Invention

[0004] The present disclosure provides a text processing method, a text generation method, an electronic device, a readable storage medium, and a computer program product.

[0005] A first aspect of the present disclosure provides a text processing method, including: parsing a target text to obtain hierarchical information of the target text, where the hierarchical information is used to describe the parent-child relationship between hierarchical elements of the target text; determining the number of tokens of the text included in multiple hierarchical elements in the target text; and segmenting the target text into multiple text blocks according to a specified token number threshold and the token number of the hierarchical elements, where the text in the same text block obtained by the segmentation is connected in position in the target text, and for a target hierarchical element with the lowest level in the highest-level hierarchical element to which it belongs, the text content belonging to the same target hierarchical element is segmented into the same text block.

[0006] According to some embodiments of the present disclosure, before parsing the target text, the method further includes: converting the document format of a target document storing the target text into a specified format, where the specified format is suitable for text parsing.

[0007] According to some embodiments of the present disclosure, parsing a target text to obtain hierarchical information of the target text includes: identifying one or more of the font format and paragraph format of the text content in the target text; and determining the hierarchical structure of the target text according to the identified format information to obtain the hierarchical information of the target text.

[0008] According to some embodiments of the present disclosure, the hierarchical information is further used to describe the types of hierarchical elements included in the target text, and the types of hierarchical elements include title elements and table of contents elements.

[0009] According to some embodiments of the present disclosure, before determining the number of tokens of the text included in multiple hierarchical elements in the target text, the method further includes: deleting the text belonging to the table of contents element from the target text.

[0010] According to some embodiments of the present disclosure, determining the number of tokens of the text included in multiple hierarchical elements in the target text specifically means: determining the number of tokens of the text included in the highest-level hierarchical element in the target text, and the text included in the highest-level hierarchical element includes the text included in the lower-level hierarchical elements of the highest-level hierarchical element and the text belonging to the highest-level hierarchical element and not belonging to the lower-level hierarchical elements, and the lower-level hierarchical elements are the hierarchical elements with the highest-level hierarchical element as the highest level.

[0011] According to some embodiments of the present disclosure, segmenting the target text into multiple text blocks based on a specified token number threshold and the token number of the hierarchical element includes: if the token number of the highest-level hierarchical element is less than or equal to the specified token number threshold, then segmenting the text included in the highest-level hierarchical element into one text block.

[0012] According to some embodiments of the present disclosure, segmenting the target text into multiple text blocks based on a specified token number threshold and the token number of the hierarchical element includes: if the token number of the highest-level hierarchical element is greater than the specified token number threshold, then taking the next-level hierarchical element of the highest-level hierarchical element as the new highest-level hierarchical element, and performing the step of determining the number of tokens of the text included in the highest-level hierarchical element in the target text until the segmentation of all the text content of the target text is completed.

[0013] According to some embodiments of the present disclosure, if the token number of the highest-level hierarchical element is greater than the specified token number threshold, then before taking the next-level hierarchical element of the highest-level hierarchical element as the new highest-level hierarchical element, first merge the text belonging to the highest-level hierarchical element and not belonging to the lower-level hierarchical elements with the text included in the first next-level hierarchical element of the highest-level hierarchical element.

[0014] According to some embodiments of the present disclosure, in the case where the token number of the highest-level hierarchical element is greater than the specified token number threshold, if the number of hierarchical elements of the lower-level hierarchical elements of the highest-level hierarchical element is less than or equal to the specified element number threshold, then segment the text included in the highest-level hierarchical element into one text block.

[0015] After obtaining the multiple text blocks, the method further includes: if there is a target text block among the multiple text blocks whose number of included tokens is less than the specified lower limit of the number of tokens, and the sum of the number of tokens of the target text block and the adjacent text block of the target text block is less than the token number threshold, then merge the target text block and the adjacent text block, where the adjacent text block is a text block adjacent to the target text block in the hierarchical structure of the target text.

[0016] A second aspect of the present disclosure provides a text generation method, including: segmenting a target text according to the text processing method described in any of the above embodiments to obtain multiple text blocks; storing the multiple text blocks; in response to receiving a query request, performing recall based on the similarity between the indication text included in the query request and the stored text blocks to obtain a recall result including at least one text block; and inputting the indication text and a prompt word formed based on the recall result into a language model to obtain a response text output by the language model for replying to the indication text.

[0017] According to some embodiments of the present disclosure, storing the multiple text blocks includes: generating vectors for each of at least some of the multiple text blocks; and storing the multiple text blocks and the vectors of the text blocks.

[0018] According to some embodiments of the present disclosure, performing recall based on the similarity between the indication text included in the query request and the stored text blocks includes: generating a target vector of the indication text included in the query request; determining multiple vectors with the highest similarity to the target vector from the vectors of the stored text blocks; and generating a recall result including the text blocks corresponding to the multiple vectors.

[0019] According to some embodiments of the present disclosure, inputting the indication text and a prompt word formed based on the recall result into a language model includes: inputting the indication text and the recall result into a re-ranking model to obtain a re-ranked recall result output by the re-ranking model; and inputting the indication text and a prompt word formed based on the re-ranked recall result into the language model.

[0020] A third aspect of the present disclosure provides an electronic device, including: a memory storing execution instructions; and a processor that executes the execution instructions stored in the memory, so that the processor executes the method described in any of the above embodiments.

[0021] A fourth aspect of the present disclosure provides a readable storage medium storing a computer program, which when executed by a processor is used to implement the method described in any of the above embodiments.

[0022] A fifth aspect of the present disclosure provides a computer program product including a computer program, which when executed by a processor is used to implement the method described in any of the above embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The drawings illustrate exemplary embodiments of the present disclosure and, together with the description thereof, are used to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure and are incorporated in this specification and form a part of this specification.

[0024] Figure 1 FIG. shows a schematic overall flow diagram of a text processing method M100 according to some embodiments of the present disclosure.

[0025] Figure 2 FIG. shows a schematic hierarchical structure diagram of hierarchical elements according to some embodiments of the present disclosure.

[0026] Figure 3 FIG. shows a schematic diagram of the text content of hierarchical elements according to some embodiments of the present disclosure.

[0027] Figures 4 - 5 FIG. shows a schematic overall flow diagram of a text processing method M100 according to some other embodiments of the present disclosure.

[0028] Figure 6 FIG. shows the process of Figure 2 performing text segmentation on the shown hierarchical structure to form text blocks.

[0029] Figure 7 FIG. shows a schematic overall flow diagram of a text processing method M100 according to some other embodiments of the present disclosure.

[0030] Figure 8 FIG. shows a schematic application scenario diagram of a text generation method according to some embodiments of the present disclosure.

[0031] Figure 9 FIG. shows a schematic overall flow diagram of a text generation method M200 according to some embodiments of the present disclosure.

[0032] Figure 10 FIG. shows a schematic overall flow diagram of a text generation method M200 according to some other embodiments of the present disclosure.

[0033] Figure 11 is a schematic block diagram of the structure of a text processing device according to an embodiment of the present disclosure.

[0034] Figure 12 It is a schematic block diagram of the text generation device according to an embodiment of the present disclosure.

[0035] Figure 13 It is a schematic block diagram of the electronic device 1000 according to an embodiment of the present disclosure.

[0036] Figure 14 It is a schematic block diagram of the electronic device 2000 according to an embodiment of the present disclosure. Detailed Embodiments

[0037] The present disclosure will be further described in detail below in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the relevant content and do not limit the present disclosure. Additionally, it should be noted that for the sake of convenience of description, only the parts related to the present disclosure are shown in the drawings.

[0038] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The technical solutions of the present disclosure will be described in detail below with reference to the accompanying drawings and embodiments.

[0039] For texts with a clear structure, such as documents for system usage instructions or documents recording various regulations, they will contain document structure information such as titles, prefaces, tables of contents, chapters, etc. When performing text splitting (i.e., text slicing), if the document is split into text blocks (text segments) of equal length according to a preset fixed length (character length or token number), multiple text blocks of the same length will be obtained; or if split according to the smallest parsable elements, multiple text blocks each containing one parsable element will be obtained. When splitting by length, it may split in the middle of a sentence or paragraph. When splitting by the smallest parsable elements, a sentence or a cell can be used as a text block, and the context information of each text block is too little.

[0040] It can be seen that both of the above splitting methods will lose the document structure information and context information during the splitting process, resulting in incomplete information in the text blocks. If the text blocks are used to build a knowledge base and perform query retrieval, the above defects will lead to poor retrieval effects and it is difficult to correctly retrieve the target that meets the user's requirements.

[0041] To this end, the present disclosure proposes a text processing method. A user can operate on a separate terminal device to utilize the text processing method for text splitting. The text processing method can be implemented in application scenarios on a server or a terminal device (such as a mobile phone) alone, or application scenarios where operations are performed on a terminal device alone, etc.

[0042] Figure 1 FIG. 4 shows an overall flowchart of a text processing method M100 according to some embodiments of the present disclosure. As Figure 1 shown, the method includes step S110, step S120, and step S130. Among them, the method can be executed by an electronic device such as a computer.

[0043] S110, parse the target text to obtain the hierarchical information of the target text. Among them, the hierarchical information is used to describe the parent-child relationship between the hierarchical elements of the target text.

[0044] The target text can be obtained from a document, a web page, or other carriers. The document storing the target text can be a local document or an online document. After obtaining the target text or the target document, the hierarchical structure of the target text is identified through text parsing. For example, the target text on the web page can be copied to an online platform or application for text parsing, or the target document storing the target text can be directly uploaded to the online platform or application, thereby realizing the parsing of the target text.

[0045] The hierarchical element is an element representing the text structure existing in the text. For example, it can be a structural element such as a title, a table of contents, or a chapter existing in the text. Different hierarchical elements may have different formats or styles in the text. There may be a parent-child relationship or a superior-subordinate relationship between different hierarchical elements. For example, there may be a parent-child relationship between different titles or between different chapters.

[0046] Figure 2 FIG. 5 shows a schematic diagram of the hierarchical structure of hierarchical elements according to some embodiments of the present disclosure. Refer to Figure 2 , Figure 2Three hierarchical elements are shown, namely the first-level element (the highest level), the second-level element, and the third-level element. Each hierarchical element is represented by an Arabic numeral and represents a kind of title, and xxx represents the title name. Among them, 1, 2, 3, and 4 are the first-level elements, 1.1, 1.2, 2.1, 2.2, 4.1, and 4.2 are the second-level elements, and 1.1.1, 1.1.2, 1.2.1, 1.2.2, 1.2.3, 2.1.1, 2.1.2, 2.2.1, and 2.2.2 are the third-level elements. The hierarchical element 1 is the parent level (parent node) or the upper level of the hierarchical elements 1.1 and 1.2, and the hierarchical elements 1.1 and 1.2 are the sub-levels (sub-nodes) or the lower levels of the hierarchical element 1. The hierarchical element 1.1 is the parent level of the hierarchical elements 1.1.1 and 1.1.2, and the hierarchical elements 1.1.1 and 1.1.2 are the sub-levels of the hierarchical element 1.1. The lower-level hierarchical elements (descendant hierarchical elements) of the hierarchical element 1 include 1.1, 1.2, 1.1.1, 1.1.2, 1.2.1, 1.2.2, and 1.2.3, and these hierarchical elements have the hierarchical element 1 as the upper-level hierarchical element (ancestor hierarchical element). The parent-child relationships between other hierarchical elements can be deduced by analogy. The hierarchical element 3 does not contain sub-levels.

[0047] The formats or styles of different hierarchical elements can be different. For example, Figure 2 in it, the indentation amounts of the first-level element, the second-level element, and the third-level element increase in sequence, and the font or other formats can also be different from the main text. For example, the font of all first-level elements can be a bold font with a larger font size, and the font of all third-level elements can be an italic font with a normal font size. When parsing the target text, the text can be identified as a hierarchical element and the hierarchical structure relationship between the identified hierarchical elements can be recognized through different formats.

[0048] Through step S110, the quantity, type, and parent-child relationship of the above hierarchical elements can be obtained. It can be understood that for the convenience of representing the format differences between different hierarchical elements, Figure 2 the text content included in each hierarchical element is not shown in it.

[0049] S120, determine the number of lemmas of the text included in multiple hierarchical elements in the target text.

[0050] A token is the smallest information unit of text. The number of tokens can be the number of characters or the number of tokens (where a token can be obtained by tokenizing the text). Tokenization can be performed by the tokenizer of the model. The model can be a BERT (Bidirectional Encoder Representations from Transformers) model or other models.

[0051] Figure 3 A schematic diagram showing the text content of the hierarchical elements of some embodiments of the present disclosure. Refer to Figure 3 , a1 represents the text content that belongs only to the first-level element 2 and does not belong to any lower-level element (2.1, 2.2, 2.1.1, 2.1.2, 2.2.1, 2.2.2) of the hierarchical element 2, a2 represents the text content that belongs only to the second-level element 2.1 and does not belong to any lower-level element (2.1.1, 2.1.2) of the hierarchical element 2.1, a3 represents the text content of the third-level element 2.1.1, a4 represents the text content of the third-level element 2.1.2, a5 represents the text content that belongs only to the second-level element 2.2 and does not belong to any lower-level element (2.2.1, 2.2.2) of the hierarchical element 2.2, a6 represents the text content of the third-level element 2.2.1, and a7 represents the text content of the third-level element 2.2.2.

[0052] It can be understood that Figure 3 the dashed boxes in are only used to label the text content of different hierarchical elements in cooperation with the reference numerals, and there is no need to display the dashed box labels when the target text is displayed on the client.

[0053] The text included in a hierarchical element can refer to the text that belongs only to this hierarchical element and does not belong to the descendant hierarchical elements of this hierarchical element. In this case, the number of tokens of the text included in the first-level element 2 is the number of tokens of a1, the number of tokens of the text included in the second-level element 2.1 is the number of tokens of a2, the number of tokens of the text included in the second-level element 2.2 is the number of tokens of a5, the number of tokens of the text included in the third-level element 2.1.1 is the number of tokens of a3, and so on for other hierarchical elements. The number of tokens of the seven hierarchical elements are the number of tokens of a1 to a7 respectively.

[0054] The tokens included in a hierarchical element can refer to both the tokens that only belong to this hierarchical element and the tokens included in all lower-level hierarchical elements with this hierarchical element as the upper-level hierarchical element. In this case, the number of tokens in the text included in the first-level element 2 is the sum of the number of tokens from a1 to a7, the number of tokens in the text included in the second-level element 2.1 is the sum of the number of tokens from a2 to a4, the number of tokens in the text included in the second-level element 2.2 is the sum of the number of tokens from a5 to a7, the number of tokens in the text included in the third-level element 2.1.1 is the number of tokens of a3, and so on for other hierarchical elements.

[0055] S130, segment the target text into multiple text blocks according to the specified token quantity threshold and the token quantity of the hierarchical element. Among them, the texts in the same text block obtained by the above segmentation are connected in position in the target text. For the target hierarchical element with the lowest level in the highest-level element it belongs to, the text content belonging to the same target hierarchical element is segmented into the same text block.

[0056] To avoid splitting too much text into the same text block, which may lead to an excessive amount of text in the text block and affect the processing effect when processing the text block subsequently, using the preset token quantity threshold as the basis for text splitting (segmentation) can keep the text volume of each segmented text block at a reasonable level.

[0057] When segmenting the target text, the text belonging to the same lowest-level element will not be split into two different segments, but will be completely segmented into the same text block (text slice) to avoid situations such as semantic loss and incomplete context in the text block. Take Figure 2 as an example, among the hierarchical elements 1.1, 1.1.1, 1.1.2, 1.2, 1.2.1, 1.2.2, 1.2.3 with the hierarchical element 1 as the highest-level element, the target hierarchical elements with the lowest level include 1.1.1, 1.1.2, 1.2.1, 1.2.2, 1.2.3, a total of five. The text content of any one of these five target hierarchical elements will be completely segmented into the same text block. For example, according to the relationship between their token quantity and the token quantity threshold, the texts of these five target hierarchical elements are respectively segmented into five different text blocks, or the texts of the target hierarchical elements 1.1.1 and 1.1.2 are segmented into one text block, and the texts of the target hierarchical elements 1.2.1, 1.2.2, and 1.2.3 are segmented into another text block. Whether it is segmented into a single text block or jointly segmented into a text block with the text content of other hierarchical elements, its text content is complete and will not damage the integrity of the context with relevance in content due to segmentation.

[0058] Similarly, among the hierarchical elements 4.1 and 4.2 with the hierarchical element 4 as the highest hierarchical element, there are two target hierarchical elements with the lowest level, namely 4.1 and 4.2. Therefore, according to the magnitude relationship between the number of tokens of the two and the token number threshold, the texts of the hierarchical elements 4.1 and 4.2 can be respectively segmented into two different text blocks or segmented into the same text block.

[0059] In addition, the texts in the same text block are connected in position in the target text. Specifically, the segmented text block is segmented from the target text as a whole. That is to say, the text arrangement order of the text content contained in each text block is exactly the same as that of these text contents in the target text, and the texts that are not connected in position in the hierarchical structure will not be segmented into the same text block. For example, the positions of the above target hierarchical elements 1.1.1 and 1.1.2 in the hierarchical structure of the target text are connected, and their text contents are connected in the target text. Therefore, their texts can be segmented into a text block as a whole when being segmented. The above target hierarchical elements 1.2.1, 1.2.2, and 1.2.3 are also connected in position in sequence. Therefore, their texts can also be segmented into a text block. However, there will be no situation where the texts of 1.2.1 and 1.2.3 are in the same text block and 1.2.2 is in another text block. Otherwise, the texts of these three will not be segmented as a whole. Because for the hierarchical elements of the same level that are adjacent in position, their text contents are also strongly correlated. Therefore, this can avoid the discontinuity of the content in the text block and ensure that the integrity of the context with relevance in content is not damaged.

[0060] When using the segmented text blocks for query retrieval or other purposes subsequently, it can be ensured that the text blocks can support these purposes well.

[0061] The text processing method proposed according to the embodiments of the present disclosure can be used to split the text of a document containing a hierarchical structure such as a title, and utilize the text content relevance reflected by the positions of hierarchical elements such as the title. Combining with the token number threshold that limits the text volume of the text block, taking hierarchical elements such as the title in the text as the minimum unit of segmentation, the text volume of each text block is reasonably configured through the threshold, and the text of each lowest-level hierarchical element can be fully segmented into the same text block instead of being segmented into different text blocks. Compared with splitting by sentence, more context information can be retained in the text block, semantic loss can be avoided, information integrity can be ensured, and it is beneficial for subsequent text processing using the text block. Compared with splitting by a fixed length, it can avoid the loss of context information and semantic loss caused by the text of the same lowest-level hierarchical element being segmented into different text blocks, ensure information integrity, and is beneficial for subsequent text processing using the text block.

[0062] Figure 4 The overall flowchart of the text processing method M100 according to other embodiments of the present disclosure is shown. Referring to Figure 4 , before performing step S110, step S101 may be performed first.

[0063] S101, convert the document format of the target document storing the target text into a specified format. The above-specified format is suitable for text parsing.

[0064] When parsing the target text, a possible scenario is that the user obtains the target document at the user side, and then manipulates the user side to parse the text in the target document as the target text. When parsing the target document, the program or tool for text parsing may only be able to parse documents in a specific format. Therefore, after obtaining the target document, if the file format of the target document is not the specified format that can be parsed, the format conversion needs to be performed first through step S101.

[0065] For example, the target document may be a document within an enterprise. Different business departments of the enterprise may use different file formats when creating the target document. Therefore, the format of the target document obtained by the user is not unique. Suppose PyMuPDF (a high-performance Python library for data extraction, analysis, conversion, and manipulation of PDF documents) is used for text parsing or assisting in text parsing, and the obtained target document is a doc-format word document. Then, the doc format can be first converted to the pdf (Portable Document Format) format through unoconv (Universal Office Converter, a command-line tool for file format conversion), and then the target document in pdf format is parsed.

[0066] It can be understood that when performing document parsing, pdfplumber (a Python library for extracting text, tables, etc. from pdf documents) or pypdf2 (a Python library for performing operations such as creating, merging, and splitting pdf documents) can also be used for parsing or assisting in parsing.

[0067] Continuing to refer to Figure 4 , step S110 may specifically include the following steps S111 and S112.

[0068] S111, identify one or more of the font format and paragraph format of the text content in the target text.

[0069] S112. Determine the hierarchical structure of the target text based on the recognized format information to obtain the hierarchical information of the target text.

[0070] When parsing the target text using a script or other tool capable of text parsing, the font or paragraph format of the text can be used to identify hierarchical elements such as headings. For example, in the writing specification of the target document, it is required that the font size of some levels of headings be larger than that of the body text, some levels of headings are required to use bold or italic fonts, and the indentation amounts of different levels of headings are different, etc. When parsing, the above-known writing specifications can be used to implement the corresponding parsing functions. By recognizing format information such as the font size, whether the font is bold, whether the font is italic, and the indentation amount of the line where it is located, hierarchical elements can be identified, and the specific level of the hierarchical element can be determined. For example, if the style of a line of text in the target text is bold, the font size is the first size, and the paragraph format is centered, then this line of text can be recognized as the title of the first-level element; if the style of a line of text is bold, the font size is the second size, and the paragraph format is left-aligned and indented at the first line, then this line of text can be recognized as the second-level element; if the style of a paragraph of text is not bold, the font size is the fourth size, and the paragraph format is left-aligned and indented at the first line, then this paragraph of text can be recognized as the body text (lexical unit) subordinate to a certain hierarchical element.

[0071] Sort all the obtained hierarchical elements according to the size of the specific level, and construct the parent-child relationship between the hierarchical elements, so as to generate hierarchical information and obtain the hierarchical tree of the document.

[0072] It can be understood that the parsing tool can also be used to obtain the document element information of the target text, including information such as text content, text type, and picture link. Adjacent multi-sentence and multi-paragraph text content can be merged to form a complete text content and be subordinate to a certain hierarchical element.

[0073] Exemplarily, the hierarchical information obtained through step S110 can also be used to describe the types of hierarchical elements included in the target text. The types of hierarchical elements can include heading elements and table of contents elements.

[0074] In some documents that need to be parsed, the hierarchical elements not only include heading elements or chapter elements, but may also include some other elements different from the body text, such as table of contents elements. However, the semantics of the table of contents is lower than that of the body text, and the content of the table of contents does not have the value of forming a text block.

[0075] Based on this, before performing step S120 (determining the number of tokens of the text included in multiple hierarchical elements in the target text), the text belonging to the table of contents element can be deleted from the target text first. By deleting the table of contents, the table of contents will not participate in text segmentation and will not appear in the text blocks. It can be understood that when deleting the table of contents, all the text before the table of contents can be deleted, because the content that appears before the table of contents in the document is generally information such as the document title, date, author, etc., and these contents also do not have the value of forming text blocks, so they can be deleted together.

[0076] Figure 5 Fig. shows a schematic overall flow diagram of a text processing method M100 according to some other embodiments of the present disclosure. Refer to Figure 5 , step S120 may include the following steps S121.

[0077] S121, determining the number of tokens of the text included in the highest-level element in the target text. Among them, the text included in the highest-level element includes: the text included in the lower-level elements of the highest-level element and the text that belongs to the highest-level element but does not belong to the lower-level elements. The lower-level element is a hierarchical element with the highest-level element as the highest level.

[0078] Figure 6 Fig. shows Figure 2 a schematic diagram of the process of text segmentation of the hierarchical structure shown into text blocks. Refer to Figure 6 , which shows a total of four highest-level elements 1, 2, 3, and 4, and N1, N2, N3, and N4 are the number of tokens of the highest-level elements 1, 2, 3, and 4 in sequence. The lower-level elements of the highest-level element 1 include hierarchical elements 1.1, 1.1.1, 1.1.2, 1.2, 1.2.1, 1.2.2, and 1.2.3, and the number of tokens N1 includes the number of tokens of these hierarchical elements. The number of tokens N1 also includes the number of tokens of the text that belongs to the highest-level element 1 but does not belong to the above-mentioned lower-level elements. Specifically, refer to Figure 3 the text content a1 in, where a1 is the text that belongs to the highest-level element 2 but does not belong to the lower-level elements of the hierarchical element 2. The composition of the number of tokens N2 and N4 of the highest-level elements 2 and 4 can be obtained in the same way. The highest-level element 3 does not include lower-level elements, so N3 only includes the number of tokens of the text directly belonging to the highest-level element 3. The values of N1, N2, N3, and N4 can be obtained through a text parsing tool.

[0079] Correspondingly, step S130 may include the following steps S131.

[0080] S131, if the number of tokens of the highest-level element is less than or equal to the specified token number threshold, then the text included in the highest-level element is segmented into one text block.

[0081] Continue to refer to Figure 6 and compare the number of tokens N1 to N4 of all the top-level elements with the token number threshold q1 in sequence. If it is determined that N1 > q1, N2 > q1, N3 < q1, and N4 > q1, then the text content of the top-level element 3 is divided into a text block to obtain the text block T1. Figure 6 The text blocks divided in are represented by dotted lines. The text block T1 contains all the text content of Chapter 3 in the target document. It can be understood that if it is determined that N4 < q1, then the text of the hierarchical element 4 including the text of the hierarchical elements 4.1 and 4.2 will be divided into a text block, and this text block contains all the text content of Chapter 4 in the target document, including the text of Section 1 and Section 2 of Chapter 4.

[0082] Step S130 may further include the following step S132. Step S131 and step S132 can start synchronously and be executed in parallel, or can be executed sequentially in any order.

[0083] S132, if the number of tokens of the top-level element is greater than the specified token number threshold, then use the next-level element of the top-level element as the new top-level element and execute the above step S121 (the step of determining the number of tokens of the text included in the top-level element in the target text) until the segmentation of all the text content of the target text is completed.

[0084] Continue to refer to Figure 6 and since it is determined that N1 > q1, N2 > q1, and N4 > q1, taking the top-level element 2 as an example, for the top-level element 2, use the next-level element 2.1 of the top-level element 2 as the top-level element and substitute it into step S121 to determine the number of tokens, and then compare the determined number of tokens with the token number threshold. Also, use the next-level element 2.2 of the top-level element 2 as the top-level element and substitute it into step S121 to determine the number of tokens, and then compare the determined number of tokens with the token number threshold.

[0085] Let n3 and n4 be the number of tokens of the text included in the new top-level elements 2.1 and 2.2 respectively, and n3 + n4 = N2. If it is determined through step S121 that n3 < q1, n4 < q1, and assuming that the hierarchical element 2 does not contain text directly belonging to 2 itself, then according to step S131, the text of the hierarchical element 2.1 including the text of the hierarchical elements 2.1.1 and 2.1.2 is divided into a text block to obtain the text block T2. Also, according to step S131, the text of the hierarchical element 2.2 including the text of the hierarchical elements 2.2.1 and 2.2.2 is divided into a text block to obtain the text block T3.

[0086] The hierarchical elements 1 and 4 can be analogized in turn according to the above steps. n1 and n2 are the number of word tokens of the text contained in the new highest-level elements 1.1 and 1.2 in turn, and n1 + n2 = N1. If it is determined through step S121 that n1 > q1 and n2 > q1, and assuming that the hierarchical element 1 does not contain text directly belonging to 1 itself, then the next-level element 1.1.1 of the hierarchical element 1.1 is used as the highest-level element and substituted into step S121 to determine the number of word tokens, and then the determined number of word tokens is compared with the word token threshold. Also, the hierarchical elements 1.1.2, 1.2.1, 1.2.2, and 1.2.3 are respectively substituted into step S121 according to the operation steps of 1.1.1.

[0087] n5 and n6 are the number of word tokens of the text contained in the new highest-level elements 4.1 and 4.2 in turn, and n5 + n6 = N4. If it is determined through step S121 that n3 < q1 and n4 < q1, and assuming that the hierarchical element 4 does not contain text directly belonging to 4 itself, then the text of the hierarchical element 4.1 is segmented into a text block according to step S131 to obtain the text block T4. Also, the text of the hierarchical element 4.2 is segmented into a text block according to step S131 to obtain the text block T5.

[0088] At this time, after a new round of iteration, a new round of determination of the number of word tokens and segmentation of text blocks is performed, and the segmentation of Chapters 2, 3, and 4 is completed, obtaining a total of five text blocks T1 to T5, and the text content of Chapter 1 remains unsegmented.

[0089] After that, the next round of iteration is started to segment the text content of Chapter 1. n11, n12, n13, n14, and n15 are the number of word tokens of the text contained in the new highest-level elements 1.1.1, 1.1.2, 1.2.1, 1.2.2, and 1.2.3 in turn. If it is determined through step S121 that n11 < Q1, n12 < Q1, n13 < Q1, n14 < Q1, and n15 < Q1, and assuming that the hierarchical elements 1.1 and 1.2 do not contain text directly belonging to 1.1 and 1.2 themselves, then the text of the hierarchical element 1.1.1 is segmented into a text block according to step S131 to obtain the text block T6. The texts of the hierarchical elements 1.1.2, 1.2.1, 1.2.2, and 1.2.3 are segmented in the same way to obtain the text blocks T7, T8, T9, and T10 respectively. Thus, the segmentation of the target text is completed, and the 4 chapters in the target text are segmented into ten text blocks T1 to T10.

[0090] Figure 7 The overall flowchart of the text processing method M100 according to some other embodiments of the present disclosure is shown. Refer to Figure 7, in step S132, if the number of tokens of the highest-level element is greater than the specified token number threshold, before taking the next-level element of the highest-level element as the new highest-level element, first merge the text that belongs to the highest-level element but not to the lower-level elements with the text contained in the first next-level element of the highest-level element.

[0091] Taking the highest-level element 4 as an example, through step S132, it is determined that the number of tokens N4 of the highest-level element 4 > q1. Since the highest-level element 4 may contain text directly belonging to itself, that is, text that belongs to the highest-level element 4 but not to the lower-level elements 4.1 and 4.2, then this part of the text is merged with the text of the hierarchical element 4.1, which is the first next-level element. Through Figure 6 It can be seen that the number of tokens n5 of the hierarchical element 4.1 includes the number of tokens of the text of 4 because Figure 6 the implementation method in it adopts text merging.

[0092] The same applies to other highest-level elements. For example, when after one round of iteration, at this time the hierarchical element 2.1 is the latest highest-level element, and its number of tokens n3 is the number of tokens of the text content directly belonging to the hierarchical element 2, that is, the number of tokens of Figure 3 the text content a1 in it. At this time, n3 is the sum of the number of tokens of the text content a1, a2, a3, and a4. Since n4 is not the first next-level element of the highest-level element 2, n4 does not include the number of tokens of the text content a1, but is only the sum of the number of tokens of the text content a5, a6, and a7.

[0093] Continue to refer to Figure 7 , in step S132, when the number of tokens of the highest-level element is greater than the specified token number threshold, if the number of hierarchical elements of the lower-level elements (descendant hierarchical elements) of the highest-level element is less than or equal to the specified element number threshold, then the text contained in the highest-level element can be split into a text block.

[0094] When it is determined through step S132 that the number of tokens of the current highest-level element is greater than the token number threshold q1, instead of directly taking the next-level element of the current highest-level element as the new highest-level element and starting a new round of iteration, first judge whether the number of elements of all lower-level elements of the highest-level element is less than or equal to the element number threshold q2. If it is greater than q2, start a new round of iteration normally. If it is less than or equal to q2, text splitting can be directly performed.

[0095] For example, when splitting Chapter 4 in the target text, the number of tokens N4 of the top-level element 4 > q1. At this time, first obtain the number of lower-level elements of the top-level element 4, and confirm that the top-level element 4 includes a total of 2 lower-level elements, namely 4.1 and 4.2. The element number threshold q2 = 2. At this time, 4.1 and 4.2 can be not used as new top-level elements for a new round of iteration, but the content of Chapter 4 can be directly split into a single text block. This can split the text content of lower-level elements with a lower depth and breadth of the hierarchical structure into the same text block, rather than splitting the content of 4.1 and 4.2 into text blocks T4 and T5 respectively, so that texts with closely related content can be located in the same text block, while taking into account both the text volume of the text block and the relevance of the text content.

[0096] For another example, when splitting Chapter 1 in the target text, the number of tokens N1 of the top-level element 1 > q1. Assume that the hierarchical elements 1.1 and 1.2 contain text directly belonging to 1.1 and 1.2 themselves. Then obtain the number of lower-level elements of the top-level element 1, and confirm that the top-level element 1 includes a total of 7 lower-level elements, namely 1.1, 1.1.1, 1.1.2, 1.2, 1.2.1, 1.2.2, 1.2.3. Since the number of lower-level elements contained in the top-level element 1 > q2, a new round of iteration will be started.

[0097] After obtaining multiple text blocks through step S130, a text block merging step can also be performed. The text block merging step includes: if there is a target text block among multiple text blocks whose contained number of tokens is less than the specified lower limit of the number of tokens, and the sum of the number of tokens of the target text block and the adjacent text block of the target text block is less than the number of tokens threshold, then merge the target text block and the adjacent text block. Among them, the adjacent text block is the text block adjacent to the target text block in the hierarchical structure of the target text.

[0098] After obtaining the above 10 text blocks T1 to T10, compare the number of tokens of each text block with the lower limit of the number of tokens, and determine the text block with the number of tokens less than the lower limit of the number of tokens as the target text block. Taking Figure 6 the hierarchical element 1.2 as an example, assume that the number of tokens n2 of the hierarchical element 1.2 is 520, the number of tokens threshold q1 is 400, and n2 > q1. Therefore, the hierarchical element 1.2 will be used as a new top-level element for iteration. During the iteration process, it is determined that the number of tokens n13, n14, and n15 of the hierarchical elements 1.2.1, 1.2.2, and 1.2.3 are 300, 120, and 100 in sequence. Since n13, n14, and n15 are all less than q1, they are split into T8, T9, and T10 in sequence.

[0099] After such segmentation, the amount of text in T9 and T10 is actually small, and the knowledge content contained is less. If they are used as text blocks alone, it may reduce the effect of subsequent text processing. Therefore, the lower limit of the number of tokens is used for judgment. Assuming that the lower limit of the number of tokens is 150, after comparing all text blocks, it is determined that the number of tokens in T9 and T10 is lower than the lower limit of the number of tokens, then T9 and T10 are the target text blocks. At this time, they can be merged in order according to the positions of the target text blocks in the target text. Since T9 is before T10, it is possible to first determine whether the adjacent text block of T9 is also a target text block.

[0100] Specifically, the adjacent text block refers to the text block adjacent to the target text block in the text order of the target text among the text blocks obtained by each segmentation. For example, in the hierarchical structure of the target text, the adjacent text blocks of T9 are T8 and T10. Since T10 is the target text block, then determine whether the sum of the number of tokens in T9 and T10 is less than the token number threshold q1. n14 + n15 < q1 = 400. Therefore, it is determined that the sum of the number of tokens in T9 and T10 is less than the token number threshold q1, so T9 and T10 are merged. After merging, the text blocks of the target text become 9 text blocks. That is, 10 text blocks are obtained by segmentation, and 9 text blocks are obtained after continuous merging. Merging can further ensure the information integrity of a single text block and avoid text completion during subsequent text processing. It can be understood that the above text block merging steps can be repeatedly executed multiple times until the number of tokens in each text block is greater than the lower limit of the number of tokens.

[0101] Figure 8 The figure shows a schematic diagram of the application scenario of the text generation method according to some embodiments of the present disclosure. In this application scenario, it may include a service side 10, a server side 20, and a client side 30. The server side 20 can communicate with the service side 10 and the server side 20 respectively to perform data transmission and reception or instruction transmission and reception. In the present disclosure, the service side 10, the server side 20, and the client side 30 each include at least one processor and at least one memory.

[0102] Exemplarily, the business terminal 10 may be a computer device operable by business personnel, the client terminal 30 may be a terminal device held by a user, and the server 20 may include a server. The business terminal 10 may store multiple target documents regarding different businesses. The server 20 obtains the target documents from the business terminal 10, performs text segmentation on the target documents by implementing the above-mentioned text processing method M100 to obtain multiple text blocks, and then by implementing the above-mentioned text generation method M200, generates vectors for each text block, and stores each text block and its vector in the server configured by itself, so as to construct a knowledge base in the server or update the knowledge base in the server. The user may initiate a query request through the client terminal 30. The client terminal 30 sends the query request to the server 20. The server 20 recalls multiple text blocks related to the user query information included in the query request in the knowledge base by comparing the vector similarity of the text, and then forms a prompt word with the recalled text blocks as the knowledge background, inputs the prompt word and the user query information into the language model together, obtains the reply of the language model, and sends the reply to the client terminal 30 for display to the user as an answer to the user's query.

[0103] In Figure 1 the shapes and structures of the business terminal 10, the server 20, and the client terminal 30 shown should not be construed as limiting the scope of protection of the present disclosure. In the present disclosure, the "terminal device" may be different types of electronic devices. For example, the terminal device may be a mobile phone, a tablet computer, a laptop computer, a desktop computer, or a self-service machine, etc. Additionally, the server of the server 20 may be a server with a physical form or a cloud server, and the present disclosure does not limit the type of the server.

[0104] Figure 9 shows an overall flowchart of the text generation method M200 according to some embodiments of the present disclosure. As Figure 9 shown, the method includes step S210, step S220, step S230, and step S240. Among them, the method may be executed by an electronic device such as a computer.

[0105] S210, segment the target text according to the text processing method to obtain multiple text blocks.

[0106] The server can receive the target document sent by the business side or actively obtain the target document from the business side. The target document can be an internal system instruction manual, a rules and regulations document, or other types of documents within the enterprise, and these documents are set with a chapter structure. The target text in the target document is segmented according to the text processing method M100 described in the above implementation manner to obtain h text blocks. By using the text processing method M100 for text segmentation, the text volume of the obtained text blocks is relatively moderate, avoiding the problem that the text block is too large to be well matched and queried, and avoiding the problem that the text block is too small to have sufficient context to generate an answer.

[0107] S220. Store the above-mentioned multiple text blocks.

[0108] The server can be configured with a server, and the obtained h text blocks can be stored in the server, thereby constructing a knowledge base as knowledge content, or storing the h text blocks in the knowledge base of the server to increase the knowledge content stored in the knowledge base. The knowledge content in the knowledge base is used as background knowledge to be provided to the language model to assist in generating answers during the subsequent process of answering user questions.

[0109] S230. In response to receiving a query request, recall is performed based on the similarity between the indication text included in the query request and the stored text blocks to obtain a recall result including at least one text block.

[0110] The user can initiate a question through the APP on the client side, and the APP answers the user's question. The user inputs the indication text in the APP, and the indication text is used to express the question asked (such as consulting how to operate a certain program) or the instruction given (such as listing the operable environment of a certain program). The client generates a query request based on the indication text input by the user and sends the query request to the server. After receiving the query request, the server extracts the indication text in the query request, and then finds one or more text blocks with the closest semantic similarity to the indication text from the text blocks in the knowledge base as the recall result. The text blocks in the recall result will be used as supplementary background information for answering the indication text.

[0111] S240. Input the indication text and the prompt words formed based on the recall result into the language model to obtain a reply text output by the language model for replying to the indication text.

[0112] The server constructs a prompt based on the text-based recall result to enable the generative language model to better understand the intent of the instruction text and answer more accurately. Then, the instruction text and the prompt are input into the language model. The language model can be a generative pre-trained model. The prompt is used to instruct the language model to generate a response text for the instruction text based on the background knowledge (i.e., the recall result) contained in the prompt.

[0113] For example, the instruction text is: "How to operate Program A to generate a form", and the recall result includes multiple text blocks about the operation steps of Program A and multiple text blocks about the operation precautions of Program A. The prompt contains the recall result, enabling the large model to reply to the user's question based on the knowledge content about the operation of Program A provided in the recall result and generate a response text. Then, the server feeds the response text back to the client to answer the user's question.

[0114] According to the text generation method proposed in the embodiments of the present disclosure, when splitting the target text for constructing the knowledge base, it is possible to split the document containing hierarchical structures such as titles, and utilize the text content relevance reflected by the positions of hierarchical elements such as titles. Combining with the token number threshold that limits the text volume of text blocks, taking hierarchical elements such as titles in the text as the minimum unit of segmentation, the text volume of each text block is reasonably configured through the threshold, and the text of each lowest-level hierarchical element can be fully segmented into the same text block instead of being segmented into different text blocks.

[0115] Compared with splitting by sentence, more context information can be retained in the text block, avoiding semantic loss, ensuring information integrity, and having higher efficiency when using the generative language model to answer. There is no need to supplement information in real time when inputting to the model.

[0116] Compared with splitting by fixed length, it can avoid the loss of context information and semantic loss caused by the text of the same lowest-level element being segmented into different text blocks, ensuring information integrity.

[0117] Therefore, when the language model generates response text using text chunks as knowledge content, the RAG (Retrieval-Augmented Generation) technology can be adopted to combine retrieval and generative language models. This allows direct reference to an external knowledge base when generating answers, eliminating the need to fine-tune the model every time new knowledge is introduced, reducing the model training cost, avoiding the generation of inaccurate content (hallucinations) due to the lack of relevant domain knowledge in the model, and also avoiding the risk of data leakage when enterprises upload private domain data to third-party platforms for model training. Instead, the new knowledge is directly provided to the model by the prompt words, enabling the model to provide more accurate, up-to-date, and contextually relevant answers, improving the accuracy and knowledgeability of the model's answers, avoiding real-time completion, and enhancing the retrieval efficiency.

[0118] Figure 10 FIG. shows a schematic overall flow diagram of a text generation method M200 according to some other embodiments of the present disclosure. Refer to Figure 10 , in step S220, the manner of storing the above-mentioned multiple text chunks may include the following steps S221 and S222.

[0119] S221, generating vectors for each of at least some of the above-mentioned multiple text chunks.

[0120] S222, storing the above-mentioned multiple text chunks and the vectors of the text chunks.

[0121] The manner of generating text chunk vectors can be to generate embedding multi-dimensional vectors of text chunks. Embedding is the process of mapping high-dimensional input data to a low-dimensional vector space; it can also be through models based on Transformer such as BERT. The Transformer model usually includes an encoder and a decoder and is suitable for processing sequence data. The Transformer model has a fixed input sequence length. Therefore, after the target text is segmented into text chunks, it is more suitable for embedding search, improving the accuracy of recall. The generated multi-dimensional vectors are used to implement semantic search of text chunks. When storing text chunks and vectors, they can be stored in the form of key-value pairs, with the vector as the index of the text chunk and the text chunk as the knowledge text for facilitating fast search.

[0122] Continue to refer to Figure 10 , in step S230, the manner of recalling the retrieval request by the similarity between the indication text included in the query request and the stored text chunks may include the following steps S231, S232, and S233.

[0123] S231, generating a target vector of the indication text included in the query request.

[0124] S232. Determine multiple vectors with the highest similarity to the target vector from the vectors of the stored text blocks.

[0125] S233. Generate a recall result including the text blocks corresponding to the multiple vectors.

[0126] When generating the target vector of the instruction text, the model used can be the same as the model used to generate the vectors of the text blocks. For example, the generation of the target vector is implemented through the Transformer model. Determine the similarity between the target vector and the stored vectors by calculating the similarity between the vectors. Specifically, the cosine similarity, Euclidean distance, etc. of the vectors can be calculated to determine the similarity between two vectors. The similarity between each stored vector and the target vector can be calculated, and then the similarities are sorted from large to small to determine the k stored vectors with the highest similarity, so as to determine the k stored text blocks with the highest similarity to the instruction text and obtain the recall result.

[0127] Continue to refer to Figure 10 In step S240, the way of inputting the instruction text and the prompt words formed based on the recall result into the language model may include the following steps S241 and S242.

[0128] S241. Input the instruction text and the recall result into the re-ranking model to obtain the re-ranked recall result output by the re-ranking model.

[0129] S242. Input the instruction text and the prompt words formed based on the re-ranked recall result into the language model.

[0130] The re-ranking model is used to perform a secondary ranking on the recall result to optimize the finally presented order and ensure that the most relevant or most useful results are ranked in the front. The re-ranking model can adopt the Rerank model. For multiple text blocks in the recall result, the re-ranking model can score the text relevance between the text blocks and the knowledge text, the text quality (such as length and complexity) of the text blocks, etc., and then re-rank the text blocks according to the scores to obtain multiple re-ranked text blocks. The number of these re-ranked text blocks can be less than or equal to the number of text blocks in the recall result, that is, after re-ranking, some text blocks may be filtered out. The training of the Rerank model can be supervised training according to the training data set formed by the labeled data.

[0131] Based on any of the above embodiments, the present disclosure also provides a text processing device. Figure 11 It is a structural schematic diagram of the text processing device according to an embodiment of the present disclosure. As Figure 11 shown, the text processing device includes: a text parsing module 110, a token number determination module 120, and a text segmentation module 130.

[0132] The text parsing module 110 is used to parse the target text to obtain the hierarchical information of the target text. Among them, the hierarchical information is used to describe the parent-child relationship between the hierarchical elements of the target text.

[0133] The token number determination module 120 is used to determine the number of tokens in the text included in multiple hierarchical elements in the target text.

[0134] The text segmentation module 130 is used to segment the target text according to the specified token number threshold and the number of tokens of the hierarchical elements to obtain multiple text blocks. Among them, the texts in the same text block obtained by the above segmentation are connected in position in the target text. For the target hierarchical element with the lowest level in the highest hierarchical element to which it belongs, the text content belonging to the same target hierarchical element is segmented into the same text block.

[0135] The above text processing device may be in the form of computer software, and each module of the above text processing device may be implemented by computer software modules. The implementation processes of the functions and roles of each module in the above text processing device are specifically described in the implementation processes of the corresponding steps in the above method, and will not be elaborated here.

[0136] Based on any one of the above embodiments, the present disclosure also provides a text generation device. Figure 12 It is a structural schematic diagram of a text generation device according to an embodiment of the present disclosure. As Figure 12 shown, the text generation device includes: a text processing device 210, a text storage module 220, a recall module 230, and a text generation module 240.

[0137] The text processing device 210 is used to segment the target text according to the text processing method to obtain multiple text blocks. The text processing device 210 is the text processing device described in any one of the above embodiments, and includes a text parsing module, a token number determination module, and a text segmentation module.

[0138] The text storage module 220 is used to store the above multiple text blocks.

[0139] The recall module 230 is used to respond to a received query request, and recall by the similarity between the indication text included in the query request and the stored text blocks to obtain a recall result including at least one text block.

[0140] The response text generation module 240 is used to input the indication text and the prompt words formed based on the recall result into the language model to obtain a response text output by the language model for replying to the indication text.

[0141] The above-mentioned text generation device may be in the form of computer software, and each module of the above-mentioned text generation device may be implemented by computer software modules. The implementation processes of the functions and roles of each module in the above-mentioned text generation device are specifically described in the implementation processes of the corresponding steps in the above-mentioned method, and will not be elaborated here.

[0142] The execution subject of the text processing method in the specific implementation manner of the present disclosure may be an electronic device such as a computer or a server.

[0143] Therefore, based on any of the above embodiments, the present disclosure further provides an electronic device, which can execute the text processing method of any of the above embodiments described in the present disclosure.

[0144] Figure 13 It is a structural schematic diagram of an electronic device 1000 according to an embodiment of the present disclosure. The hardware structure of the electronic device 1000 can be implemented using a bus architecture. The bus architecture can include any number of interconnected buses and bridges, depending on the specific application of the hardware and the overall design constraints. The bus 1100 connects various circuits including one or more processors 1200, a memory 1300, and / or hardware modules together. The bus 1100 can also connect various other circuits 1400 such as peripheral devices, voltage regulators, power management circuits, external antennas, etc.

[0145] The execution subject of the text generation method in the specific implementation manner of the present disclosure may be an electronic device such as a computer or a server.

[0146] Therefore, based on any of the above embodiments, the present disclosure further provides an electronic device, which can execute the text generation method of any of the above embodiments described in the present disclosure.

[0147] Figure 14 It is a structural schematic diagram of an electronic device 2000 according to an embodiment of the present disclosure. The hardware structure of the electronic device 2000 can be implemented using a bus architecture. The bus architecture can include any number of interconnected buses and bridges, depending on the specific application of the hardware and the overall design constraints. The bus 2100 connects various circuits including one or more processors 2200, a memory 2300, and / or hardware modules together. The bus 2100 can also connect various other circuits 2400 such as peripheral devices, voltage regulators, power management circuits, external antennas, etc.

[0148] The present disclosure also provides a readable storage medium storing a computer program which, when executed by a processor, is used to implement the above method. The "readable storage medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples of the readable storage medium include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer diskette case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable read-only memory (CDROM), etc.

[0149] The present disclosure also provides a computer program product. The method of the present disclosure can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed, the processes or functions of the present disclosure are executed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device, a core network device, an OAM, or other programmable devices.

[0150] The computer program or instructions can be stored in a readable storage medium, or transmitted from one readable storage medium to another. For example, the computer program or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The readable storage medium can be any accessible available medium or a data storage device such as a server or data center integrating one or more available mediums. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it can also be an optical medium, such as a digital video disc; or it can be a semiconductor medium, such as a solid-state drive. The computer-readable storage medium can be a volatile or non-volatile storage medium, or can include both volatile and non-volatile types of storage media.

[0151] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, a system, or a computer program product. Therefore, the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0152] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present disclosure. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more of the blocks.

[0153] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction means that implements the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more of the blocks.

[0154] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more of the blocks.

[0155] In the description of this specification, the description with reference to terms such as "one embodiment / way", "some embodiments / ways", "example", "specific example", or "some examples", etc. means that the specific features, structures, or characteristics described in connection with the embodiment / way or example are included in at least one embodiment / way or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment / way or example. Moreover, the specific features, structures, or characteristics described can be combined in a suitable manner in any one or more embodiments / ways or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments / ways or examples described in this specification and the features of different embodiments / ways or examples.

[0156] Those skilled in the art should understand that the above embodiments are only for clearly illustrating the present disclosure, rather than limiting the scope of the present disclosure. For those skilled in the art, other changes or variations can be made based on the above disclosure, and these changes or variations are still within the scope of the present disclosure.

Claims

1. A text processing method, characterized in that, Including: Parsing the target text to obtain the hierarchical information of the target text, where the hierarchical information is used to describe the parent-child relationship between the hierarchical elements of the target text; Determining the number of tokens of the text included in multiple hierarchical elements in the target text; And Dividing the target text into multiple text blocks according to a specified token number threshold and the token number of the hierarchical elements. The text in the same text block obtained by the division is connected in position in the target text. For the target hierarchical element with the lowest level in the highest-level element to which it belongs, the text content belonging to the same target hierarchical element is divided into the same text block.

2. The text processing method according to claim 1, wherein Before parsing the target text, the method further includes: Converting the document format of the target document storing the target text into a specified format, where the specified format is suitable for text parsing.

3. The text processing method according to claim 1 or 2, characterized in that, Parsing the target text to obtain the hierarchical information of the target text, including: Identifying one or more of the font format and paragraph format of the text content in the target text; and Determining the hierarchical structure of the target text according to the identified format information to obtain the hierarchical information of the target text.

4. The text processing method according to claim 1, wherein The hierarchical information is further used to describe the types of hierarchical elements included in the target text, and the types of hierarchical elements include title elements and table of contents elements.

5. The text processing method according to claim 4, wherein Before determining the number of tokens of the text included in multiple hierarchical elements in the target text, the method further includes: Deleting the text belonging to the table of contents element from the target text.

6. The text processing method according to claim 1, characterized in that Determining the number of tokens of the text included in multiple hierarchical elements in the target text specifically includes: Determining the number of tokens of the text included in the highest-level element in the target text. The text included in the highest-level element includes the text included in the lower-level elements of the highest-level element and the text belonging to the highest-level element and not belonging to the lower-level elements, where the lower-level elements are the hierarchical elements with the highest-level element as the highest level; Optionally, dividing the target text into multiple text blocks according to a specified token number threshold and the token number of the hierarchical elements, including: If the number of tokens of the highest-level element is less than or equal to the specified token number threshold, dividing the text included in the highest-level element into one text block; Optionally, dividing the target text into multiple text blocks according to a specified token number threshold and the token number of the hierarchical elements, including: If the number of tokens of the highest-level element is greater than the specified token number threshold, taking the next-level element of the highest-level element as the new highest-level element, and performing the step of determining the number of tokens of the text included in the highest-level element in the target text until the division of all the text content of the target text is completed; Optionally, if the number of tokens of the highest-level element is greater than a specified token number threshold, before taking the next-level element of the highest-level element as the new highest-level element, first merge the text that belongs to the highest-level element and does not belong to the lower-level element with the text included in the first next-level element of the highest-level element; Optionally, when the number of tokens of the highest-level element is greater than the specified token number threshold, if the number of hierarchical elements of the next-level element of the highest-level element is less than or equal to the specified element number threshold, divide the text included in the highest-level element into a text block; Optionally, after obtaining the multiple text blocks, the method further includes: If there is a target text block among the multiple text blocks whose included token number is less than the specified lower limit of the token number, and the sum of the token numbers of the target text block and the adjacent text block of the target text block is less than the token number threshold, then merge the target text block and the adjacent text block, where the adjacent text block is the text block adjacent to the target text block in the hierarchical structure of the target text.

7. A text generation method, characterized in that, Including: Segment the target text according to the text processing method according to any one of claims 1-6 to obtain multiple text blocks; Store the multiple text blocks; In response to receiving a query request, recall through the similarity between the indication text included in the query request and the stored text blocks to obtain a recall result including at least one text block; And Input the indication text and the prompt word formed based on the recall result into a language model to obtain a reply text output by the language model for replying to the indication text; Optionally, storing the multiple text blocks includes: Generating vectors for each of at least some of the multiple text blocks; and Storing the multiple text blocks and the vectors of the text blocks; Optionally, recalling through the similarity between the indication text included in the query request and the stored text blocks includes: Generating a target vector for the indication text included in the query request; Determining multiple vectors with the highest similarity to the target vector from the stored vectors of the text blocks; and Generating a recall result including the text blocks corresponding to the multiple vectors; Optionally, inputting the indication text and the prompt word formed based on the recall result into a language model includes: Inputting the indication text and the recall result into a re-ranking model to obtain a re-ranked recall result output by the re-ranking model; and Inputting the indication text and the prompt word formed based on the re-ranked recall result into a language model.

8. An electronic device, characterized in that, Including: A memory that stores execution instructions; And A processor that executes the execution instructions stored in the memory, so that the processor executes the method according to any one of claims 1 to 7.

9. A readable storage medium, characterized in that, A computer program is stored in the readable storage medium, and when the computer program is executed by a processor, it is used to implement the method according to any one of claims 1 to 7.

10. A computer program product, characterized in that, The computer program product includes a computer program which, when executed by a processor, is used to implement the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Image-based test question monomer segmentation method and device, equipment and storage medium

    CN122290130A