Text processing method and related device

By generating and scoring the second text of the associated relationship, the final reply text is filtered out from the first reply text set of the large language model, which solves the problem of low answer accuracy of the large language model in open domain question answering and achieves higher reply text matching and accuracy.

CN120687554APending Publication Date: 2025-09-23MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510566690.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing large language models cannot effectively judge the correctness of retrieved document information in open-domain question answering, resulting in low answer accuracy.

Method used

By generating a second text corresponding to each first reply text in the first reply text set, the second text contains information representing the association relationship between the first reply text and the first text, and filtering out the final reply text from the first reply text set based on the text score.

Benefits of technology

Improved the accuracy of the final response text, ensuring that the determined response text closely matches the original text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687554A_ABST
    Figure CN120687554A_ABST
Patent Text Reader

Abstract

The invention provides a text processing method and a related device. The method comprises the steps of determining a first reply text set of a first text based on a first reference text corresponding to the first text; performing text processing on the first reference text and a preset example text to obtain a second text corresponding to each first reply text in the first reply text set; the second text comprises first information used for representing an association relationship between the first reply text and the first text; determining a text score of the second text based on the first information; and determining reply text of the first text from the first reply text set based on the text score. Through the method and the device, the accuracy of the reply text of the finally determined first text can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of Internet technology, and in particular to a text processing method and related devices. Background Art

[0002] As a text processing and information retrieval tool, question-answering systems (Q&A) provide an efficient and convenient way to acquire knowledge. Using natural language processing technology, Q&A systems accurately understand user questions and search for answers within vast amounts of data, significantly improving the efficiency and accuracy of knowledge acquisition.

[0003] With the development of large language model technology, open-domain question answering is now commonly performed using large language models such as the Generative Pre-trained Transformer 4 (GPT4). This approach retrieves multiple documents related to the question from an external database, and directly obtains the answer based on these documents. However, this approach fails to consider the accuracy of the information in the retrieved documents, which can lead to inaccurate answers. Summary of the Invention

[0004] The embodiments of the present application provide a text processing method and related devices, which can improve the accuracy of the reply text finally determined.

[0005] The technical solution of the embodiment of the present application is implemented as follows:

[0006] An embodiment of the present application provides a text processing method, which includes: determining a first reply text set of the first text based on a first reference text corresponding to the first text; performing text processing on the first reference text and a preset example text to obtain a second text corresponding to each first reply text in the first reply text set; the second text includes first information for characterizing the association relationship between the first reply text and the first text; based on the first information, determining a text score of the second text; and based on the text score, determining a reply text of the first text from the first reply text set.

[0007] An embodiment of the present application provides a text processing device, including: a determination module, used to determine a first reply text set of the first text based on a first reference text corresponding to the first text; a text processing module, used to perform text processing on the first reference text and a preset example text to obtain a second text corresponding to each first reply text in the first reply text set; the second text includes first information for characterizing the association relationship between the first reply text and the first text; the determination module is also used to determine a text score of the second text based on the first information; the determination module is also used to determine a reply text of the first text from the first reply text set based on the text score.

[0008] An embodiment of the present application provides an electronic device, comprising: a memory for storing computer-executable instructions; and a processor for implementing the text processing method provided in the embodiment of the present application when executing the computer-executable instructions stored in the memory.

[0009] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions for implementing the text processing method provided in the embodiment of the present application when executed by a processor.

[0010] An embodiment of the present application provides a computer program product, which includes computer-executable instructions, which are stored in a computer-readable storage medium; wherein, when a processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, the text processing method provided in the embodiment of the present application is implemented.

[0011] The embodiments of the present application have the following beneficial effects:

[0012] When performing text processing, first determine a first reply text set containing different first reply texts, and perform text processing on the first reference text and the preset example text to obtain a second text corresponding to each first reply text in the first reply text set, wherein the second text includes first information for characterizing the association relationship between the first reply text and the first text, and the first information explains the association relationship between the first reply text and the first text, that is, the second text can provide an explanation of the association relationship between the first reply text and the first text; then, determine the text score of the second text, and thereby determine the final reply text of the first text from the first reply text set according to the text score of the second text, that is, filter out the final reply text of the first text from the first reply text set according to the text score of the second text. In this way, when screening the final reply text, the relevant explanatory content given by the second text regarding the association relationship between the first reply text and the first text is taken into consideration. That is to say, the embodiment of the present application will combine the score of the relevant explanatory content of the association relationship between the first reply text and the first text (i.e., the text score) to screen the final reply text of the first text from the first reply text set. In this way, it can be ensured that the determined final reply text is the reply text that best matches the first text, thereby improving the accuracy of the final reply text. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 This is an architectural diagram of a text processing system provided by an embodiment of the present application;

[0014] Figure 2 This is an optional flowchart of the text processing method provided in the embodiment of the present application;

[0015] Figure 3 This is a schematic diagram of an optional implementation flow of determining a first reply text set for a first text provided in an embodiment of the present application;

[0016] Figure 4 This is another optional implementation flow diagram of determining a first reply text set for a first text provided in an embodiment of the present application;

[0017] Figure 5 This is a schematic diagram of the implementation process of generating a second text provided in an embodiment of the present application;

[0018] Figure 6 This is a schematic diagram of the implementation process of determining a text score provided in an embodiment of the present application;

[0019] Figure 7 This is another optional flowchart of the text processing method provided in the embodiment of the present application;

[0020] Figure 8This is a model block diagram of the text processing method provided in the embodiment of the present application;

[0021] Figure 9 is a structural diagram of a text processing device provided in an embodiment of the present application;

[0022] Figure 10 It is a schematic diagram of the composition structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0024] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0025] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0026] The relevant data collection and processing in the embodiments of this application should be strictly in accordance with the requirements of relevant national laws and regulations when applied in examples, and the informed consent or separate consent of the personal information subject should be obtained. Subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.

[0027] Before further explaining the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are first explained. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.

[0028] 1) Large Language Model (LLM): This refers to a natural language processing model with a large number of parameters. This model typically requires extensive computing resources and training data to handle a variety of complex natural language tasks, including language understanding, generation, and translation.

[0029] 2) Zero-Shot Learning: This refers to the ability of a model to make effective predictions and decisions on new tasks or categories without ever having seen training data for that specific task or category, greatly expanding the applicability of large models.

[0030] With the development of large language model technology, large language models such as GPT4 and Qwen-14B are currently commonly used for open domain question answering. The main process of question answering is as follows: First, based on the question to be answered, several documents that are most relevant to the question to be answered are retrieved from an external database as reference documents for the large model. Then, multiple reference documents and question and answer prompts are input into the large model, and the answers to the questions to be answered are obtained through the large model. This method has the advantages of not requiring fine-tuning of the large language model and is simple; however, when the retrieved external information contains conflicting information, the large model cannot determine the correctness of the document information, resulting in incorrect answers.

[0031] Based on at least one problem existing in the related art, in order to solve the problem that there is conflicting information in multiple recalled reference documents and the large model cannot judge the correctness of the information in the reference, resulting in a low answer accuracy, the embodiment of the present application proposes an open domain text processing method based on large model conflict information, that is, a text processing method. When performing text processing, first determine a first reply text set containing different first reply texts, and based on the first reference text corresponding to the first text, generate a second text corresponding to each first reply text in the first reply text set, the second text includes first information for characterizing the association relationship between the first reply text and the first text, and the second text is used to give the relevant explanation content of taking the first reply text as the final reply text of the first text, and then, by determining the text score of the second text, the final reply text of the first text is determined from the first reply text set according to the text score of the second text, that is, the final reply text of the first text is filtered out from the first reply text set according to the text score of the second text. In this way, when screening the final reply text, the relevant explanatory content given by the second text is taken into consideration. That is to say, the embodiment of the present application will combine the score of the reason for each first reply text as the final reply text of the first text (i.e., the text score), and screen the final reply text of the first text from the first reply text set. In this way, it can be ensured that the determined final reply text is the reply text that best matches the first text, thereby improving the accuracy of the final reply text.

[0032] See also Figure 1 , Figure 1 This is an architectural diagram of a text processing system 100 provided in an embodiment of the present application. To support a text processing application, a terminal 400 is connected to a server 200 via a network 300. The network 300 may be a wide area network or a local area network, or a combination of the two.

[0033] Terminal 400 is used to send a first text to server 200. Server 200 is used to respond to the first text sent by terminal 400 and determine a first reply text set for the first text based on a first reference text corresponding to the first text; then, text processing is performed on the first reference text and a preset sample text to generate a second text corresponding to each first reply text in the first reply text set; the second text includes first information for characterizing the association relationship between the first reply text and the first text; then, a text score of the second text based on the first information is determined; finally, based on the text score, a reply text for the first text is determined from the first reply text set. After obtaining the reply text for the first text, server 200 sends the reply text for the first text to terminal 400, so that the reply text for the first text is output on terminal 400.

[0034] In some embodiments, the server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (CDNs), and big data and artificial intelligence platforms. The terminal 400 can be any device such as a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch or car terminal that can install a question-and-answer application or provide a question-and-answer program to implement the text processing method, but the terminal 400 is not limited to the listed devices. The terminal 400 and the server 200 can be directly or indirectly connected by wired or wireless communication, which is not limited in the embodiments of the present application.

[0035] The text processing methods provided in each embodiment of the present application can be executed by an electronic device, wherein the electronic device can be a server or a terminal, that is, the text processing methods in each embodiment of the present application can be executed by a server, or by a terminal, or by interaction between a server and a terminal.

[0036] See also Figure 2 , Figure 2 This is an optional flow chart of the text processing method provided in the embodiment of the present application, which will be combined with Figure 2 The steps shown are described below, and the text processing method is described by taking the execution subject as a server as an example. The method includes the following steps S101 to S104:

[0037] Step S101: determining a first reply text set for the first text based on a first reference text corresponding to the first text.

[0038] The text processing method in the embodiment of the present application can be a question-and-answer method. The process of the server performing text processing on the first text can be understood as performing question-and-answer processing on the first text. The first text can be a question input by the user. When performing question-and-answer processing, the first reference text corresponding to the question can be obtained first, and multiple first reply texts of the question can be generated based on the first reference text to obtain a first reply text set. Then, the most accurate reply text is screened out from the first reply text set as the final answer to the question. In the implementation process, the first reply text in the first reply text set can be a candidate answer to the question. The first reply text set can be understood as a candidate answer set. The final answer to the question can be screened out from the candidate answer set, that is, the final reply text of the first text is obtained.

[0039] Here, the first reference text may be any one or more types of documents, and the type of the first reference text may include at least one of the following: text documents, structured data documents, multimedia-related documents, professional field documents, social media content, and other document types.

[0040] Text documents include web pages, news articles, blog posts, academic papers, and e-books. For web pages, search engines can extract text content from web pages, including titles, body text, and meta tags. News articles can be found on news websites or in news databases, providing the latest information and event coverage. Blog posts can be content from personal or professional blogs, covering a variety of topics and perspectives. Academic papers can be research papers in academic databases, used for academic research and knowledge dissemination. E-books can be complete books, including novels and non-fiction. Structured data documents include database records and tabular data. Database records can include product information (such as product titles, descriptions, and prices on e-commerce websites), user information, and order records. Tabular data can include Excel spreadsheets or database tables, which are used to search for specific data fields. Multimedia-related documents include video scripts and audio transcripts. Video scripts are text transcripts of video content, used for video retrieval and content matching. Audio transcripts are text transcriptions of audio files, such as podcasts and lectures. Professional documents include medical records, legal documents, and technical documentation. Medical records include medical records, diagnostic reports, and other medical records; legal documents include legal clauses, case studies, and contracts; and technical documents include software development documentation, user manuals, and technical standards. Social media content includes documents such as social media posts and forum posts. Social media posts include user-generated content on platforms like Weibo and Twitter; forum posts include discussion topics and replies within forums. Other document types include email content and chat logs. Email content includes the subject line and body of the email; chat logs are chat logs within instant messaging software.

[0041] In this embodiment of the present application, the number of first reference texts obtained is one or more, and the number of first reference texts to be obtained can be preset. After obtaining multiple first reference texts corresponding to the first text, a first reply text set for the first text can be determined based on all the first reference texts. The first reply text set includes multiple first reply texts.

[0042] When determining a first set of reply texts for the first text based on a first reference text corresponding to the first text, semantic analysis can be performed on the first text to extract key information of the first text, such as keywords and text type (e.g., "what" and "how"). Simultaneously, text feature extraction can be performed, i.e., features related to the first text, such as keywords, key sentences, and paragraphs, can be extracted from the first reference text. In a specific implementation, semantic analysis can be performed on the question input by the user to extract key information of the question, and text feature extraction can be performed, i.e., features related to the question can be extracted from the first reference text.

[0043] In an embodiment of the present application, the first reference text can be vectorized using methods such as TF-IDF (Term Frequency-Inverse Document Frequency) and Word2Vec (Word to Vector) to extract the semantic features of the first reference text. After extracting the semantic features of the first reference text, the content that matches the keywords of the first text can be located in the first reference text to determine the paragraphs or sentences that may contain the first reply text. The semantic similarity between the first text and the first reference text fragments can be calculated using a semantic matching algorithm (such as cosine similarity, text representation based on a bidirectional encoder of Transformer (BERT, Bidirectional Encoder Representations from Transformers), etc.), and the fragments that are most relevant to the semantics of the first text can be screened out. Afterwards, based on the semantic matching results, a fragment that is highly relevant to the first text is extracted from the first reference text as the first reply text, and the extracted fragments are reorganized and optimized to make them semantically complete and clearly expressed. Natural language generation technology (such as a generative pre-trained transformer (GPT, Generative Pre-trained Transformer), etc.) can be used to polish the fragments to obtain multiple first reply texts.

[0044] In some embodiments, see Figure 3 , Figure 3 It is shown that step S101 can be implemented by the following steps S1011 to S1012:

[0045] Step S1011 : Retrieve first reference texts corresponding to the first text from a first database to obtain a first reference text set.

[0046] In an embodiment of the present application, multiple search paths can be used to retrieve first reference texts corresponding to the first text from the first database to obtain a first reference text set. The first database can include multiple types of documents, for example, text documents, structured data documents, multimedia-related documents, professional domain documents, social media content, and documents of other document types. The first database can also include multiple data sub-databases, each of which includes a type of document.

[0047] Here, the term "search path" refers to different search methods, including but not limited to at least one of the following: full-text search, index-based search, Boolean search, natural language search, graph database query, data mining, and fuzzy matching. Alternatively, methods such as BM25 and vector search can be used. In full-text search, relevant documents can be quickly located by associating each term in the documents in the first database with a list of documents containing that term. Algorithms such as TF-IDF and BM25 are then used to score the query results, ensuring that the most relevant documents are ranked first. Full-text search can utilize a distributed architecture, support horizontal scalability, and handle massive amounts of data. In index-based search, a B-tree structure can be created to quickly locate documents that meet the criteria. A hash function can also be used to map key values ​​to storage locations for faster search speeds. Alternatively, a full-text index can be used, which is a specialized index for text data and can quickly retrieve documents containing specific keywords. In Boolean search, logical operations, such as AND, OR, and NOT, can be used to combine multiple keywords for more precise search results. When performing natural language retrieval, users can be allowed to ask questions in natural language, and the system will automatically analyze the questions and return relevant results. When querying a graph database, data is stored in the form of nodes (entities) and edges (relationships), which is suitable for data modeling and querying of complex relationships. By traversing the graph structure, relevant nodes and relationships can be efficiently queried, making it suitable for retrieval in scenarios such as social networks and recommendation systems. When performing data mining, large amounts of data can be analyzed to discover hidden patterns, regularities, and association rules, thereby obtaining relevant first reference texts. When performing fuzzy matching, users are allowed to enter keywords that do not fully match, but records similar to the keywords will be returned. It should be noted that these listed retrieval methods can be selected and combined according to different application scenarios and needs to improve the efficiency and accuracy of retrieval.

[0048] In the embodiment of the present application, each search path can be used to retrieve multiple first reference texts corresponding to the first text from the first database. The first reference text set includes first reference texts retrieved by multiple search paths. The first reference texts in the first reference text set can be grouped according to the different search paths, with each group of first reference texts corresponding to a search path.

[0049] Step S1012: Perform text mapping on the first text and the first reference text in the first reference text set to obtain a first reply text set of the first text.

[0050] During implementation, the first text and the first reference text in the first reference text set can be input into a pre-trained large language model. The large language model then performs text mapping on the first text and the first reference text in the first reference text set to obtain a first reply text set for the first text. In other words, the large language model generates a first reply text set for the first text based on the first text and the first reference text.

[0051] For example, the pre-trained large language model can be a Qwen-72B-chat model. The first text and a first reference text from the first reference text set can be combined and input into the Qwen-72B-chat model, and the Qwen-72B-chat model can be used to generate a first reply text to the first text. During implementation, a prompt word can also be generated and input into the large language model along with the first text and the first reference text from the first reference text set.

[0052] In an embodiment of the present application, a first reply text can be generated by combining the retrieved first reference text, and the retrieved first reference text can be input into the large language model as context. The large language model then generates the first reply text based on these first reference texts. Since multiple first reference texts are retrieved, the contents of the multiple first reference texts can be merged and input into the large language model, or after generating paragraphs of the first reply text based on each first reference text, the paragraphs of the first reply text can be merged to obtain the first reply text.

[0053] In some embodiments, the large language model can be configured to generate multiple different first reply texts by adjusting generation parameters of the large language model (e.g., temperature, sampling strategy, etc.). After generating multiple first reply texts, a sorter can be used to sort the generated multiple first reply texts to ensure that the most relevant first reply texts are arranged at the front of the sequence. A preset first number of first reply texts are then selected from the sequence as the first reply texts in the set of first reply texts.

[0054] In an embodiment of the present application, the first reference text set obtained by retrieval provides the large language model with rich contextual information, which helps the large language model to understand the background and semantics of the first text more accurately, thereby generating a more accurate first reply text, and multiple retrieval methods can cover different types of documents and information sources, ensuring that the large language model can obtain knowledge related to the first text from multiple perspectives. The first reference text set contains multiple related documents, and the large language model can extract different viewpoints and information from the first reference text to generate a variety of first reply texts. The large language model can also generate multiple different first reply texts by adjusting generation parameters (such as temperature, sampling strategy, etc.) to meet the needs of different users. In addition, the use of multiple retrieval methods can reduce the deviations or omissions that may be caused by a single retrieval method, ensuring that even if some retrieval methods fail to find the ideal document, other methods can supplement it, and even if some of the retrieved documents have low relevance to the first text, the large language model can filter out useful information and generate high-quality first reply texts through its powerful semantic understanding and generation capabilities. Multiple search methods can also quickly locate the documents most relevant to the first text, reducing the amount of irrelevant information that the large language model needs to process and improving overall processing efficiency. Through precise search, the large language model only needs to process documents that are highly relevant to the first text, reducing the computational cost of generating the first reply text. Multiple search methods can also adapt to different types of data sources and document formats, allowing the text processing system to be flexibly expanded to new fields and application scenarios. Retrieved documents can also change dynamically based on updates to the first database, ensuring that the text processing system can obtain the latest information in a timely manner. By combining the advantages of retrieval and generation, the text processing system can provide higher-quality first reply text that better meets user needs, thereby improving user satisfaction.

[0055] In other embodiments, see Figure 4 , Figure 4 It shows that step S101 can also be implemented by the following steps S1013 to S1016:

[0056] Step S1013: Retrieve first reference texts corresponding to the first text from the first database to obtain a first reference text set.

[0057] Step S1014: Deduplication of the first reference texts in the first reference text set is performed to obtain a second reference text set.

[0058] Here, deduplication of the first reference text in the first reference text set can be achieved by using any of the following deduplication methods: hash-based deduplication, feature set-based deduplication, deep learning-based deduplication, and sentence fingerprint-based deduplication.

[0059] Hash-based deduplication can be achieved through the MinHash algorithm, SimHash algorithm or Locality Sensitive Hashing (LSH). The MinHash algorithm refers to representing the first reference text as an n-gram set, and then generating a signature of the first reference text through the MinHash algorithm. The first reference texts are judged to be duplicated by comparing the similarity of the signatures. The SimHash algorithm generates a fixed-length signature (such as 64 bits) for each first reference text, and judges whether the first reference texts are similar by calculating the Hamming distance between the signatures. Locality sensitive hashing is combined with the MinHash algorithm to map the first reference text to multiple buckets. Similar first reference texts are more likely to be mapped to the same bucket, thereby accelerating deduplication.

[0060] Deduplication based on feature sets can be achieved through the K-Shingle algorithm or n-gram and Jaccard similarity. The K-Shingle algorithm involves dividing the first reference text into continuous segments (K-shingles) of length K, then using these segments as feature sets for the first reference text. The similarity between the K-shingle sets of two first reference texts is then calculated to determine whether the two first reference texts are duplicates. The n-gram and Jaccard similarity algorithms extract n-gram features from the first reference text and then calculate similarity based on the extracted n-gram features. If the similarity exceeds a set similarity threshold, the first reference text is considered duplicated.

[0061] Deep learning-based deduplication can be achieved through text embedding and similarity calculation or a distributed deduplication framework. Text embedding and similarity calculation refers to using a deep learning model (such as BERT) to embed the first reference text into a vector space, and then determining whether the first reference text is duplicated by calculating the cosine similarity between the vectors. A distributed deduplication framework utilizes a distributed multi-node mechanism to process the text data of the first reference text in parallel, efficiently deduplicating text through techniques such as generating hash values ​​and union-find sets.

[0062] Sentence fingerprint-based deduplication can be implemented using the KSentence algorithm. The KSentence algorithm assumes that the longest K sentences of two duplicate first reference texts are identical. Duplicate first reference texts are filtered out by concatenating the longest K sentences and calculating the MD5 value as the text fingerprint of the first reference text.

[0063] Step S1015 : Based on the fourth information of the second reference texts in the second reference text set, the second reference texts in the second reference text set are merged to obtain a third reference text set.

[0064] Here, the fourth information refers to attribute information of the second reference text. For example, the fourth information may include information such as the upload time of the second reference text, a timestamp recorded in the second reference text, or the subject of the second reference text. In other words, some of the second reference texts in the second reference text set may be merged based on information such as the subject, upload time, and timestamp recorded in the text.

[0065] In the case where the fourth information is the upload time of the second reference text, the second reference texts of the second reference text set are merged based on the fourth information of the second reference text of the second reference text set. This can be achieved in the following way: first, the server extracts the upload time of each second reference text from the metadata of the second reference text; then, the server sorts the second reference texts in the second reference text set according to the extracted upload time to ensure that the second reference texts are arranged in chronological order; finally, the server checks the sorted second reference texts and merges the second reference texts uploaded at adjacent time points (for example, the time interval is within a certain time threshold range). When merging, the contents of these second reference texts can be spliced ​​together, or the common themes or key information of these second reference texts can be extracted for integration. After generating the merged reference text, that is, after generating the third reference text in the third reference text set, the merged context information (such as the time range of the merger) is recorded.

[0066] In the case where the fourth information is the timestamp recorded in the second reference text, the second reference texts of the second reference text set are merged based on the fourth information of the second reference text of the second reference text set. This can be achieved in the following way: first, the server extracts the recorded timestamp from the content of the second reference text (for example, the time when the specific event mentioned in the second reference text occurred); then, the server sorts the second reference text set according to the extracted timestamp; thereafter, the server checks the sorted second reference text sequence and merges the second reference texts with similar timestamps (for example, the time when the event occurred on the same day or the same time period). When merging, the contents of these second reference texts can be spliced ​​together, or common events or key information of these second reference texts can be extracted for integration. After generating the merged reference text, that is, after generating the third reference text in the third reference text set, the merged context information (such as the merged timestamp range) is recorded.

[0067] In the case where the fourth information is the subject of the second reference text, the second reference texts of the second reference text set are merged based on the fourth information of the second reference text of the second reference text set. This can be achieved in the following way: first, the server performs a topic analysis on the second reference texts and extracts the subject or keyword of each second reference text. Natural language processing technology (such as TF-IDF, LDA, BERT, etc.) can be used here to identify the subject of the second reference text; then, the server groups the second reference texts in the second reference text set according to the extracted subject, and groups texts with the same or similar subjects into one group; thereafter, the server merges the second reference texts on the same subject. When merging, the contents of these second reference texts can be spliced ​​together, or the common views or key information of these second reference texts can be extracted and integrated. If the content of the second reference text is large, the key information can be extracted using summary technology. Finally, a merged third reference text is generated. After the server reads the third reference text, it records the merged context information (such as the merged subject).

[0068] Through the above three processing methods, the server can effectively merge the second reference text according to the upload time, recorded timestamp or subject, thereby reducing redundant information and improving the efficiency and quality of subsequent processing.

[0069] Step S1016: Perform text mapping on the first text and the third reference text in the third reference text set to obtain a first reply text set for the first text.

[0070] Here, the first text and the third reference text in the third reference text set can be combined and input into the large language model, and the first reply text of the first text can be generated by the large language model. During the implementation process, a prompt word can also be generated and input into the large language model together with the prompt word, the first text and the third reference text in the third reference text set.

[0071] In this embodiment of the present application, deduplication processing allows the server to remove duplicate first reference texts, preventing the large language model from being interfered with by redundant information when generating the first reply text, thereby improving the accuracy and relevance of the generated first reply text. Merging processing can integrate similar or related second reference texts, enabling the large language model to obtain more comprehensive and coherent contextual information, thereby generating a higher-quality first reply text.

[0072] Step S102: Perform text processing on the first reference text and the preset sample text to obtain a second text corresponding to each first reply text in the first reply text set.

[0073] In one implementation, based on the above steps S1011 to S1012, a first reply text set of the first text is generated based on the first text and the first reference text. In step S102, a second text corresponding to each first reply text in the first reply text set can be generated based on the first reference text and the preset example text.

[0074] In another implementation, since the first reference texts in the first reference text set are deduplicated and merged in steps S1013 to S1016 to obtain a third reference text set, and a first reply text set is generated based on the first text and the third reference texts in the third reference text set, in step S102, a second text corresponding to each first reply text in the first reply text set can also be generated based on the third reference text and a preset example text.

[0075] In some embodiments, see Figure 5 , Figure 5 It is shown that step S102 can be implemented by the following steps S1021 to S1023:

[0076] Step S1021: parse the first sample content in the sample text to obtain the information type of the second information.

[0077] The second information here refers to information such as the number of references contained in the first example content of the example text, the information source of the references, and the update time of the documents. The information type of the second information refers to the type of information angle corresponding to the second information, such as the number of documents, information source, and time type.

[0078] In one implementation, the first example content in the example text may be an example reason, and the example reason in the example text may be parsed to obtain the information type of the second information.

[0079] Step S1022: Determine the third information of the first reference text under the information type.

[0080] In this embodiment of the present application, third information of this information type can be determined for the first reference text, namely, third information of the second reference text based on information such as the number of reference texts included, the information source of the reference texts, and the update time of the document. The third information includes information such as the number of reference texts included in the first reference text, the information source of the reference texts, and the update time of the document.

[0081] Step S1023: Generate a second text including the first information based on each first reply text, the third information and the example text in the first reply text set.

[0082] In some embodiments, based on the first reply text, the third information and the example text in the first reply text set, a second text including the first information is generated. This can be achieved in the following way: for each first reply text in the first reply text set, the first reply text, the third information in the first reference text and the example text are input into a large language model; through the large language model, the first reply text, the third information and the example text are text-mapped to obtain a second text including the first information. That is, through the large language model, based on the first reply text, the third information in the first reference text and the example text, a second text including the first information is generated. In this way, for each first reply text, the first reply text, the third information and the example text can be input into the large language model, and then the large language model can be used to generate a second text corresponding to each first reply text.

[0083] In the embodiment of the present application, by describing the number and source of reference texts, the reliability and authority of each first reply text can be intuitively understood, the update time of the reference text can be clearly stated, and whether the information on which the first reply text is based is clear, thereby enhancing the credibility of the first reply text. Generating a second text can help quickly understand the basis of subsequent reply texts when selecting subsequent reply texts, reduce the time for further inquiries, and describe the reasons from multiple angles, which can meet the requirements for information integrity and depth of the second text. That is, describing the reasons from the perspectives of the number of reference documents, information sources, update time, etc., provides comprehensive background information for subsequent reply text selection, helps to select a more accurate reply text, and the detailed reason description can enhance the persuasiveness of the selected final reply text.

[0084] Step S103: Determine a text score of the second text based on the first information.

[0085] In some embodiments, see Figure 6 , Figure 6 It shows that step S103 can be implemented by the following steps S1031 to S1035:

[0086] Step S1031: Identify the first content of each sentence corresponding to the first information in the second text.

[0087] Here, the content corresponding to the first information in the second text may be divided into a plurality of sentences, and then the first content in each sentence may be identified.

[0088] Here, the purpose of performing sentence segmentation on the content corresponding to the first information in the second text is to divide the continuous content into multiple independent sentences. Sentence segmentation on the content corresponding to the first information in the second text can be performed using any of the following methods: punctuation-based segmentation, rule-based segmentation, machine learning-based segmentation, deep learning-based segmentation, language model-based segmentation, etc.

[0089] Punctuation-based segmentation methods divide sentences by identifying punctuation marks at the end of sentences (such as periods, question marks, and exclamation points). Punctuation-based segmentation methods are simple to implement, fast, and work well for well-formatted text (such as news articles and academic papers). Rule-based segmentation methods use complex rules to handle special cases, such as identifying abbreviations, ellipsis, and brackets. Rule-based segmentation methods can handle some special cases, improve segmentation accuracy, and the rules can be flexibly adjusted according to specific needs. Machine learning-based segmentation methods use machine learning models (such as hidden Markov models, conditional random fields, and neural networks) to automatically learn the rules for sentence segmentation. Machine learning-based segmentation methods can automatically learn patterns in text, adapt to different types of text, and work well for complex text (such as spoken text). Deep learning-based segmentation methods use deep learning models (such as the Transformer architecture) to further improve the accuracy of sentence segmentation. Deep learning-based segmentation methods can capture more complex language patterns and work well for large-scale text data. Language model-based segmentation methods use the contextual understanding capabilities of pre-trained language models (such as GPT and BERT) to segment sentences. Language model-based segmentation methods can leverage the powerful contextual understanding capabilities of language models and are particularly effective for complex and colloquial texts.

[0090] Identifying the first content in each sentence may be identifying the core content in each sentence, thereby ignoring non-key descriptive content. Text recognition technology may be used to identify key information in each sentence, extract the key information, and obtain the first content in the sentence.

[0091] Step S1032: rewrite the first content according to a preset format to obtain a third text corresponding to each sentence.

[0092] In the embodiment of the present application, the first content of the sentence is rewritten by at least one of the following methods: synonym replacement and introduction of related concepts, query reorganization and structured expression, and LLM-based rewriting method.

[0093] Synonym replacement and related concept introduction involves replacing keywords in the first content of a sentence with synonyms or near-synonyms while simultaneously introducing concepts or themes related to the first content, further enriching the query content and producing a rewritten third text. Query reorganization and structured expression involves reorganizing the vocabulary and phrases in the first content of a complex sentence to form a structured and clearly defined expression. For example, multiple keywords in the first content can be recombined into more meaningful phrases or sentences to produce a rewritten third text. LLM-based rewriting methods utilize the natural language understanding and generation capabilities of a large language model to analyze the context and intent of the first content of a sentence in the second text, generating a more accurate and natural rewriting result. Implementation methods may include: providing a small number of examples to allow the large language model to learn how to rewrite the first content into the third text; or, through multiple rewriting iterations, iteratively optimizing the first content based on each rewriting result, gradually adjusting the third text to improve retrieval performance; or, using the large language model to generate a third text related to the first content. For example, when rewriting the first piece of content, synonym rewriting can be performed, using different words to express the same meaning, such as rewriting "improve efficiency" to "enhance effectiveness." Sample prompts can also be added, such as by providing multiple example texts so that the large language model can learn how to rewrite based on these examples. Abstract rewriting can also be performed, such as distilling the core intent of the first piece of content to simplify the search process.

[0094] In an embodiment of the present application, the third text obtained after rewriting has a preset format. The preset format here can be a predefined format, and the format of the third text output by the large language model can be informed by inputting a prompt word into the large language model. For example, the preset format can be to write the first content before rewriting in the first paragraph, and to write the third text generated after rewriting in the second paragraph, and to add special characters (such as: ##) at the beginning of the second paragraph and the end of the third text, and output it in this two-paragraph format with special characters.

[0095] Step S1033: Determine a text segment corresponding to the third text from the second database.

[0096] Here, for the third text of each sentence, a plurality of text segments most relevant to the third text may be retrieved from the database.

[0097] Here, let's take a customer service Q&A scenario as an example. Suppose a customer service representative receives a question: "My order status has been showing 'Processing' for three days. What's going on?" The server first separates the second text of the first reply into sentences based on punctuation. For example, the second text might read: "The order status showing 'Processing' may be due to system delays. We recommend checking the status again later. If the problem persists, please contact customer support." After sentence segmentation, the following results are obtained: Sentence 1: The order status showing 'Processing' may be due to system delays. Sentence 2: We recommend checking the status again later. Sentence 3: If the problem persists, please contact customer support. The core content of these three sentences can then be rewritten into a third text using a large language model. The server inputs each sentence into the large language model, rewriting it into a third text suitable for retrieval. An example of the rewritten third text is as follows: Original Sentence 1: The order status showing 'Processing' may be due to system delays. Rewritten Third Text: The order status is processing, the reason for the system delay. Original Sentence 2: We recommend checking the status again later. Rewritten Third Text: The order status update time. Original Sentence 3: If the problem persists, please contact customer support. Rewritten third text: Order status is being processed, please contact customer service.

[0098] Step S1034: Determine a first score between the third text and the text segment.

[0099] In some embodiments, determining the first score between the third text and the text fragment can be achieved in the following manner: first, inputting the third text and the text fragment into a first model; determining the relevance between the text fragment and the third text through the first model; and then determining the first score between the third text and the text fragment based on the relevance.

[0100] In an embodiment of the present application, the first model can determine the relevance between the third text and the text fragment through semantic matching technology. The relevance between the third text and the text fragment can be determined by vector similarity calculation. Specifically, the third text and each text fragment can be encoded as vectors respectively, and the relevance is determined by calculating the cosine similarity between the vectors to determine whether the text fragment supports the third text. If the relevance is greater than a preset relevance threshold, it is determined that the text fragment supports the third text. If the relevance is less than or equal to the relevance threshold, it is determined that the text fragment does not support the third text. Alternatively, a pre-trained language model (such as BERT, SentenceBERT) can be used as the first model to encode the third text and the text fragment, and then the relevance between the third text and the text fragment is determined by the output of the first model.

[0101] Continuing with the above-mentioned customer service question-and-answer scenario as an example, assuming that 5 text fragments corresponding to each third text are retrieved, for the third text of sentence 1, after the first model determines the relevance of the contents of the 5 text fragments to the third text of sentence 1, the support results obtained based on the relevance are: [support, support, support, support, support], then based on the support result, the first score between the third text and the text fragment is determined to be 1; for the third text of sentence 2, after the first model determines the relevance of the contents of the 5 text fragments to the third text of sentence 2, the support results obtained based on the relevance are: [support, not support, support, support, support], then based on the support result, the first score between the third text and the text fragment is determined to be 0; for the third text of sentence 3, after the first model determines the relevance of the contents of the 5 text fragments to the third text of sentence 1, the support results obtained based on the relevance are: [support, support, not support, not support, support], then based on the support result, the first score between the third text and the text fragment is determined to be 0. That is to say, for the third text of each sentence, if all text fragments corresponding to the third text support the third text, the value of the first score between the third text and the text fragment is 1; if there is at least one text fragment among all text fragments corresponding to the third text that does not support the third text, the value of the first score between the third text and the text fragment is 0.

[0102] Step S1035 : Determine the text score of the second text based on the first score of each third text.

[0103] In some embodiments, determining the text score of the second text based on the first score of each third text can be achieved by: first, determining the first total score of the second text based on the first score of each third text; and determining the number of sentences in the second text; then, determining the text score of the second text based on the number and the first total score.

[0104] During implementation, the first scores of the third text for each sentence of the second text may be summed to obtain a first total score, and then the first total score may be divided by the number of sentences into which the second text is divided to obtain the text score of the second text.

[0105] Step S104: determining a reply text of the first text from the first reply text set based on the text score.

[0106] Here, the first reply text having the maximum text score may be selected as the final reply text to the first text.

[0107] In an embodiment of the present application, by retrieving multiple text fragments that are most relevant to each third text from the second database and determining the relevance between the text fragments and the third text through the first model, it is possible to ensure that the basis of the first reply text comes from multiple reliable sources, thereby improving the credibility of the first reply text, and the first model judges the relevance of the third text and the text fragments to ensure that only truly relevant text fragments are scored, thereby further improving the accuracy of the first reply text. By calculating the support score of the second text corresponding to each first reply text (that is, the first total score divided by the number of sentences divided by the second text), the server can quantitatively evaluate the quality of each first reply text, thereby more objectively selecting the final reply text.

[0108] The text processing method provided in the embodiments of the present application can be applied to various question-and-answer scenarios, such as customer service consultation scenarios, smart home control scenarios, educational counseling scenarios, medical and health consultation scenarios, and legal consultation scenarios, etc., as illustrated below.

[0109] In the customer service consultation scenario, the text processing method provided by the embodiment of the present application can be used in the enterprise customer service system to quickly respond to common user questions, such as product consultation, after-sales service, etc. For example, in the customer service consultation scenario, the user can enter the question "What are the advantages and disadvantages of this product A?" on the customer service consultation platform, and the customer service consultation platform can use the text processing method provided by the embodiment of the present application to determine the answer to the question and output the answer to the user. In the smart home control scenario, the text processing method provided by the embodiment of the present application can realize voice control of home appliances through predefined instruction rules. In the education and tutoring scenario, the text processing method provided by the embodiment of the present application can provide students with standardized answers to questions, such as mathematical formulas, grammatical analysis, etc. In the medical and health consultation scenario, the text processing method provided by the embodiment of the present application can provide standardized suggestions for common diseases. In the legal consultation scenario, the text processing method provided by the embodiment of the present application can provide relevant legal terms and cases as answer output based on the user's questions.

[0110] The text processing method provided in the embodiment of this application will be explained below using the above-mentioned customer service consultation scenario as an example. Figure 7 This is another optional flow chart of the text processing method provided in the embodiment of the present application, such as Figure 7 As shown, the method includes the following steps S201 to S222:

[0111] Step S201: The terminal obtains a first text input by a user.

[0112] The first text may be a question entered by the user.

[0113] Step S202: The terminal sends the first text to the server.

[0114] Here, the first text can be encapsulated into a question-and-answer request, and the terminal can send the question-and-answer request to the server. In some embodiments, the terminal can send the question-and-answer request using a protocol such as Hypertext Transfer Protocol (HTTP) or Web Socket. After receiving the question-and-answer request, the server parses the question-and-answer request to obtain the user's first text.

[0115] In step S203 , the server uses multiple search paths to search for first reference texts corresponding to the first text from the first database to obtain a first reference text set.

[0116] Step S204: The server removes duplicates from the first reference texts in the first reference text set to obtain a second reference text set.

[0117] Step S205 : The server merges the second reference texts in the second reference text set based on the fourth information of the second reference texts in the second reference text set to obtain a third reference text set.

[0118] In step S206 , the server inputs the first text and the third reference text in the third reference text set into the second model.

[0119] In step S207 , the server performs text mapping on the first text and the third reference text in the third reference text set using the second model to generate a first reply text set for the first text.

[0120] Here, the second model may be a pre-trained large language model. The second model may be the same model as the first model or a different model.

[0121] Step S208: The server performs text parsing on the first example content in the example text to obtain the information type of the second information.

[0122] Step S209: The server determines the third information of the first reference text under the information type.

[0123] In step S210 , the server inputs the first reply text, the third information and the sample text into the third model for each first reply text in the first reply text set.

[0124] In step S211, the server generates a second text containing at least the first information based on the first reply text, the third information and the sample text through a third model.

[0125] Here, the third model may also be a pre-trained large language model. The third model may be the same model as the first model and the second model, or a different model.

[0126] Step S212: The server identifies the first content of each sentence corresponding to the first information in the second text.

[0127] In step S213 , the server rewrites the first content according to a preset format to obtain a third text corresponding to the sentence.

[0128] Step S214: The server retrieves a text segment corresponding to the third text from the second database.

[0129] In step S215 , the server inputs the third text and the text segment into the first model.

[0130] In step S216 , the server determines the relevance between the text segment and the third text using the first model.

[0131] In step S217 , the server determines a first score between the third text and the text segment based on the relevance.

[0132] Step S218: The server determines a first total score of the second text based on the first score of each third text.

[0133] In step S219 , the server determines the number of sentences in the second text, and determines a text score of the second text based on the number and the first total score.

[0134] In step S220 , the server determines a reply text of the first text from the first reply text set based on the text score.

[0135] Step S221: The server sends a reply text of the first text to the terminal.

[0136] Step S222: The terminal displays the reply text of the first text on the current interface.

[0137] The text processing method provided by the embodiment of the present application, when performing text processing, first determines a first reply text set containing different first reply texts, and based on the first reference text corresponding to the first text, generates a second text corresponding to each first reply text in the first reply text set, and gives the reason for taking the first reply text as the final reply text of the first text through the second text, and then, by determining the text score of the second text, the final reply text of the first text is determined from the first reply text set according to the text score of the second text, that is, the final reply text of the first text is filtered out from the first reply text set according to the text score of the second text. In this way, when filtering the final reply text, the reasons given by the second text are taken into account, that is, the embodiment of the present application will combine the score (that is, the text score) of the reason for each first reply text being the final reply text of the first text, and filter the final reply text of the first text from the first reply text set. In this way, it can be ensured that the determined final reply text is the reply text that best matches the first text, thereby improving the accuracy of the final reply text.

[0138] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.

[0139] In order to solve the problem that there is conflicting information in multiple recalled reference texts, and the large language model is unable to judge the correctness of the reference information, resulting in low accuracy of the final reply text, an embodiment of the present application provides an open domain text processing method based on large model conflict information, namely the above-mentioned text processing method. Figure 8 This is a model block diagram of the text processing method provided in the embodiment of the present application, see Figure 8 The text processing method may be executed by an electronic device, which may be a terminal or a server. This embodiment of the application takes the text processing method executed by a server as an example. The method includes the following steps S301 to S308:

[0140] In step S301, based on the question to be answered (i.e., the first text), J search methods (J can be 3) (such as BM25, vector search, etc., i.e., the search path) are used to retrieve M (M can be 10) reference documents (i.e., the first reference text) that are most relevant to the question from an external database (i.e., the first database).

[0141] Step S302 : multiple reference documents retrieved by various search methods are combined to remove duplications and obtain L reference documents (ie, the third reference text set).

[0142] Step S303: Input the question and L reference documents into the large language model, and let the large language model generate K (K can be 3) candidate answers (i.e., the first reply text mentioned above).

[0143] For example, the generated prompt words can be:

[0144] "Below are L reference documents related to the question. Please generate 3 correct answers for the question based on the reference documents. Each answer should be in the following form: (a) xx, (b) yy, (c) zz.

[0145] Reference document 1: {content of reference document 1};

[0146] Reference document 2: {content of reference document 2};

[0147] Reference document L: {content of reference document L};

[0148] Question: {question}.

[0149] Please note that each answer should not exceed 5 words.

[0150] In step S304, K candidate answers and L reference documents are input into the large language model respectively. The large language model is asked to imitate a person when processing conflicting information, and explain the reasons for obtaining the candidate answer from the perspectives of the number of reference documents containing the candidate answer, the information source of the reference documents, the update time of the reference documents, etc. There are a total of K candidate reason texts (i.e., the second text mentioned above).

[0151] In this embodiment of the application, in order to improve the standardization of the second text format, S (S can be 5) sample texts are also provided to the large language model. Each sample text consists of a question, a list of fragments of reference documents, an answer, and the reason for the answer. The generated prompt words can be:

[0152] "You are a natural language processing expert. Based on the question and example text, please provide reasons for your answer from the perspectives of the number of reference documents, the authority of the reference documents' information sources, and the update time of the reference documents. Be careful not to make up stories.

[0153] The reference examples are as follows:

[0154] {sample text 1};

[0155] {sample text 2};

[0156] {Sample Text S}.

[0157] question:{question};

[0158] Reference document 1: {content of reference document 1};

[0159] Reference document 2: {content of reference document 2};

[0160] Reference document L: {content of reference document L};

[0161] Answer: {Answer}

[0162] explain:".

[0163] In step S305, the reasons for the candidate answers (i.e., the second text) are divided into sentences according to punctuation marks, and each sentence is rewritten into a third text (i.e., the search query) using a large language model. Assume that there are a total of P third texts (i.e., the reasons for the candidate answers are divided into P sentences).

[0164] For example, the generated third text may be:

[0165] "You are a natural language processing expert. Please rewrite the input sentence, identify the core content of the sentence, ignore non-critical descriptions, and make the rewritten question concise and clear.

[0166] The reference examples are as follows:

[0167] Input sentence: {input sentence}

[0168] Output in the following format:

[0169] -Rewritten question-: <rewritten question>".

[0170] Step S306 : Retrieve Q (Q may be 5) text segments that are most relevant to each third text from the database (ie the second database).

[0171] In step S307, the third text and the retrieved Q text segments are input into the large language model (i.e., the first model mentioned above). The large language model determines whether the third text is supported by the content of the Q text segments, and obtains a support score (i.e., the first score mentioned above). If supported, the score is 1; if not, the score is 0.

[0172] Step S308: Count the support scores of the reasons corresponding to each candidate answer, and calculate the total support score of the reason (i.e., the first total score mentioned above) by dividing the number of support points of the reason clauses by the number of clauses P of the reason. The total support score can be calculated using the following formula (1):

[0173]

[0174] Among them, answer[i] represents the model output result supported by the content of Q text fragments output by the large model regarding whether the i-th third text is retrieved.

[0175] After obtaining the total support score of each reason, the candidate answer with the highest total support score for the corresponding reason is selected as the final answer.

[0176] The text processing method provided in the embodiment of the present application first allows the large language model to generate multiple candidate answers based on the recalled reference documents, and then allows the large language model to explain the reasons for obtaining the candidate answers from the perspectives of the number of reference documents containing the candidate answers, the authority of the information source of the reference documents, the update time of the reference documents, etc., so that at least the following beneficial effects can be achieved: (1) Compared with generating only one answer, the large language model generates multiple candidate answers, which can prevent the large model from answering the wrong question due to information conflict, and retain the possibility of the final answer being correct; (2) When processing conflicting information, the large language model imitates human thinking to give reasons for the candidate answers, which lowers the threshold for judging whether the answer is correct and improves the accuracy of the final answer. In addition, the text processing method provided in the embodiment of the present application also divides the reasons corresponding to the candidate answers according to punctuation, recalls the relevant information of each sentence in the reason text (i.e., the retrieved Q text fragments), and allows the large language model to judge the support of the text fragments for each sentence one by one, thereby reducing the difficulty of judgment, improving the accuracy of the judgment results, and ultimately improving the accuracy of the answer.

[0177] The following continues to describe the exemplary structure of the text processing device 455 provided in the embodiment of the present application as a software module. The text processing device 455 may be a software module in an electronic device. In some embodiments, Figure 9 As shown, the software modules in the text processing device 455 may include: a determination module 4551, used to determine a first reply text set of the first text based on a first reference text corresponding to the first text; a text processing module 4552, used to perform text processing on the first reference text and a preset example text to obtain a second text corresponding to each first reply text in the first reply text set; the second text includes first information for characterizing the association relationship between the first reply text and the first text; the determination module 4551 is also used to determine the text score of the second text based on the first information; the determination module 4551 is also used to determine the reply text of the first text from the first reply text set based on the text score.

[0178] In some embodiments, the text processing module 4552 is also used to: perform text parsing on the first example content in the example text to obtain the information type of the second information; determine the third information of the first reference text under the information type; and generate a second text including the first information based on each first reply text, the third information and the example text in the first reply text set.

[0179] In some embodiments, the text processing module 4552 is also used to: for each first reply text in the first reply text set, perform text mapping on the first reply text, the third information and the example text to obtain a second text including the first information.

[0180] In some embodiments, the determination module 4551 is also used to: identify the first content of each sentence corresponding to the first information in the second text; rewrite the first content according to a preset format to obtain a third text corresponding to each sentence; determine the text fragment corresponding to the third text from the second database; determine the first score between the third text and the text fragment; and determine the text score of the second text based on the first score of each third text.

[0181] In some embodiments, the determination module 4551 is further used to: input the third text and the text fragment into a first model; determine the relevance between the text fragment and the third text through the first model; and determine a first score between the third text and the text fragment based on the relevance.

[0182] In some embodiments, the determination module 4551 is further used to: determine the first total score of the second text based on the first score of each third text; determine the number of sentences in the second text; and determine the text score of the second text based on the number and the first total score.

[0183] In some embodiments, the determination module 4551 is also used to: retrieve a first reference text corresponding to the first text from a first database to obtain a first reference text set; perform text mapping on the first text and the first reference text in the first reference text set to obtain a first reply text set for the first text.

[0184] In some embodiments, the text processing module 4552 is also used to: deduplicate the first reference text in the first reference text set to obtain a second reference text set; merge the second reference text in the second reference text set based on the fourth information of the second reference text in the second reference text set to obtain a third reference text set; the determination module 4551 is also used to: perform text mapping on the first text and the third reference text in the third reference text set.

[0185] It should be noted that the description of the device embodiment of the present application is similar to the description of the method embodiment described above, and has similar beneficial effects as the method embodiment, so it will not be repeated. For technical details not disclosed in the device embodiment, please refer to the description of the method embodiment of the present application for understanding.

[0186] Correspondingly, an embodiment of the present application provides an electronic device, Figure 10 This is a schematic diagram of the structure of the electronic device provided in the embodiment of the present application, such as Figure 10 As shown, the electronic device 1200 includes at least: a processor 1201, a communication interface 1202, and a storage medium 1203 configured to store executable instructions, wherein: the processor 1201 generally controls the overall operation of the electronic device 1200. The communication interface 1202 enables the electronic device to communicate with other terminals or servers via a network. The storage medium 1203 is configured to store instructions and applications executable by the processor 1201, and can also cache data to be processed or processed by the processor 1201 and various modules in the electronic device 1200, which can be implemented using flash memory (FLASH) or random access memory (RAM).

[0187] The embodiment of the present application provides a computer-readable storage medium in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the processor will execute the text processing method provided by the embodiment of the present application, for example, Figure 2 The text processing method shown.

[0188] An embodiment of the present application provides a computer program product including computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the text processing method described above in the embodiment of the present application.

[0189] In some embodiments, the computer-readable storage medium may be a memory such as RAM, read-only memory (ROM), flash memory, magnetic surface memory, optical disk, or compact disc read-only memory (CD-ROM); or it may be various devices including one or any combination of the above memories.

[0190] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0191] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).

[0192] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.

[0193] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.

Claims

1. A text processing method, characterized in that: The method comprises: Determining a first reply text set for the first text based on a first reference text corresponding to the first text; Performing text processing on the first reference text and a preset sample text to obtain a second text corresponding to each first reply text in the first reply text set; the second text includes first information for characterizing an association relationship between the first reply text and the first text; determining a text score of the second text based on the first information; Based on the text score, a reply text to the first text is determined from the first reply text set.

2. The method according to claim 1, characterized in that The performing text processing on the first reference text and the preset sample text to obtain a second text corresponding to each first reply text in the first reply text set includes: Performing text parsing on the first example content in the example text to obtain the information type of the second information; Determine third information of the first reference text under the information type; A second text including the first information is generated based on each first reply text in the first reply text set, the third information, and the sample text.

3. The method according to claim 2, characterized in that The step of generating a second text including the first information based on each first reply text in the first reply text set, the third information, and the sample text comprises: For each first reply text in the first reply text set, text mapping is performed on the first reply text, the third information and the example text to obtain a second text including the first information.

4. The method according to claim 1, wherein The determining the text score of the second text based on the first information includes: identifying a first content of each sentence corresponding to the first information in the second text; rewriting the first content according to a preset format to obtain a third text corresponding to each sentence; determining a text segment corresponding to the third text from a second database; determining a first score between the third text and the text segment; Based on the first score of each third text, a text score of the second text is determined.

5. The method according to claim 4, characterized in that Determining a first score between the third text and the text segment includes: inputting the third text and the text segment into a first model; Determining the relevance between the text segment and the third text using the first model; Based on the relevance, a first score between the third text and the text segment is determined.

6. The method according to claim 4, characterized in that The determining the text score of the second text based on the first score of each third text includes: determining a first total score of the second text based on the first score of each third text; determining the number of sentences in the second text; A text score for the second text is determined based on the quantity and the first total score.

7. The method according to any one of claims 1 to 6, characterized in that The determining, based on a first reference text corresponding to the first text, a first reply text set for the first text includes: Retrieving a first reference text corresponding to the first text from a first database to obtain a first reference text set; Text mapping is performed on the first text and a first reference text in the first reference text set to obtain a first reply text set for the first text.

8. The method according to claim 7, characterized in that After obtaining the first reference text set, the method further includes: Deduplicating first reference texts in the first reference text set to obtain a second reference text set; Based on the fourth information of the second reference texts in the second reference text set, merging the second reference texts in the second reference text set to obtain a third reference text set; The performing text mapping on the first text and the first reference text in the first reference text set includes: Text mapping is performed on the first text and a third reference text in the third reference text set.

9. An electronic device, characterized in that: include: a memory for storing computer-executable instructions; The processor is configured to implement the text processing method according to any one of claims 1 to 8 when executing the computer-executable instructions stored in the memory.

10. A computer program product, characterized in that The computer program product includes computer-executable instructions stored in a computer-readable storage medium; Wherein, when the processor of the electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, the text processing method according to any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Knowledge graph query method and device based on multi-language model result fusion

    CN121636662A