Oral text rewriting method and device based on counterfactual reasoning
Through a multi-level rewriting method based on counterfactual reasoning, the language model's intention understanding and answer accuracy problems when dealing with colloquial problems are solved, and more accurate and reliable professional text generation is achieved, which improves the practical application effect of the model.
Patent Information
- Application Number
- CN202510115591.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-01-24
AI Technical Summary
When dealing with colloquialization problems, it is difficult for existing language models to accurately understand the intention of colloquial expression, resulting in the generated answers that do not match the facts or provide outdated knowledge, which affects the reliability and practical application effect of the model.
The colloquial text rewriting method based on counterfactual reasoning is adopted, and by constructing colloquial data sets and counterfactual reasoning models, using multi-level counterfactual modeling and loss function optimization, colloquial text is gradually rewritten, making it close to the semantics and style of professional texts.
It improves the accuracy and reliability of rewriting spoken texts, ensures that the generated professional text can more accurately reflect the user's intentions and knowledge in professional fields, and enhances the effectiveness of the model in practical applications.
Smart Images

Figure CN119558270B_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to the field of artificial intelligence technology, and more specifically, to a method and device for rewriting colloquial text based on counterfactual reasoning. Background Art
[0002] In recent years, with the rapid progress of deep learning technology, large language models (LLMs) such as ChatGPT, LLaMa, Qwen, etc. have gradually become the stars in the field of artificial intelligence. These models, through pre-training on a vast amount of text data, have not only learned rich language representations and semantic understanding capabilities but also achieved remarkable results in various tasks of natural language processing (NLP). They can understand and generate natural language, perform various complex tasks such as machine translation, text summarization, sentiment analysis, etc., greatly promoting the development of artificial intelligence technology.
[0003] However, although LLMs perform excellently in handling language tasks, they still rely on large-scale high-quality corpora during the training process. Although these corpora provide rich language information, they lack direct perception of the real world. Especially in specific domain question-and-answer tasks, LLMs may generate hallucinations, that is, generate information inconsistent with facts or provide outdated knowledge, which seriously affects the reliability of the model and its effectiveness in practical applications. To overcome these problems, researchers have proposed a new method - Retrieval-Augmented Generation (RAG).
[0004] The core of the RAG method is that when a user asks a question, the system first retrieves in an external knowledge base to find information fragments related to the question, and then aggregates these fragments for the LLM to refer to in order to generate an answer. In this process, the quality of the retrieval results directly determines the accuracy of the LLM's answer. Retrieval methods usually include keyword matching, semantic matching, or a combination of both. However, the questions asked by users are often non-normalized and colloquial, which have significant semantic differences from the professional and normalized texts in the knowledge base, posing a major challenge to the accuracy of retrieval.
[0005] In order to improve the accuracy of retrieval, researchers have proposed a variety of text rewriting technologies. These technologies aim to convert users' colloquial questions into a format that is closer to the professional text in the knowledge base, so that the retrieval system can match relevant information more accurately. However, most existing text rewriting methods focus on rewriting regular texts. They assume that the distribution of the text before and after the rewriting is similar, and do not take into account the particularity of colloquial texts. It can be seen that there are still great challenges in rewriting colloquial texts. Summary of the invention
[0006] According to an embodiment of the present invention, a colloquial text rewriting solution based on counterfactual reasoning is provided. This solution enables the rewriting of colloquial text to understand the intention of the colloquial expression and convert it into a language that can be understood by a professional vertical field.
[0007] In a first aspect of the present invention, a method for rewriting colloquial text based on counterfactual reasoning is provided. The method comprises:
[0008] Constructing a spoken language dataset, wherein the spoken language dataset includes professional texts and spoken language texts;
[0009] Constructing a counterfactual reasoning model, using a spoken language dataset, and training the counterfactual reasoning model according to a loss function of the counterfactual reasoning model to obtain a trained counterfactual reasoning model;
[0010] Input the colloquial text into the trained counterfactual reasoning model to obtain the rewritten professional text;
[0011] The counterfactual reasoning model includes at least two generation models arranged in sequence; and the structure of each generation model of the counterfactual reasoning model is the same; wherein the first generation model is used to rewrite the spoken text once and output the rewritten text; the remaining generation models in the counterfactual reasoning model are used to rewrite the rewritten text again and output the rewritten text.
[0012] Furthermore, if the spoken data set does not include context information, each generative model in the counterfactual reasoning model is a first encoding-decoding network or a first single decoding network.
[0013] Furthermore, the first encoding and decoding network includes:
[0014] The first embedding layer is used to embed words and position encode the input data and output first feature information;
[0015] The encoding layer is used to receive the first feature information and output the second feature information;
[0016] The second embedding layer is used to perform word embedding and positional encoding on the input data and output the encoded information;
[0017] The first decoding layer is used to receive the second feature information of the encoding layer and the encoded information of the second embedding layer for decoding and output the rewritten text.
[0018] Further, the first single decoding network includes:
[0019] The fourth embedding layer is used to perform word embedding and positional encoding on the input data and output the fifth feature information;
[0020] The second decoding layer is used to receive the fifth feature information for decoding and output the rewritten text.
[0021] Further, if the colloquial dataset includes context information, each generation model in the counterfactual reasoning model is a second encoder-decoder network or a second single decoding network.
[0022] Further, the second encoder-decoder network includes:
[0023] The first embedding layer is used to perform word embedding and positional encoding on the input data and output the first feature information;
[0024] The third embedding layer is used to obtain context information for word embedding and positional encoding and output the third feature information;
[0025] The encoding layer is used to receive the fused feature information of the first feature information and the third feature information and output the fourth feature information;
[0026] The second embedding layer is used to perform word embedding and positional encoding on the input data and output the encoded information;
[0027] The first decoding layer is used to receive the fourth feature information of the encoding layer and the encoded information of the second embedding layer for decoding and output the rewritten text.
[0028] Further, the second single decoding network includes:
[0029] The fourth embedding layer is used to perform word embedding and positional encoding on the input data and output the fifth feature information;
[0030] The fifth embedding layer is used to obtain context information for word embedding and positional encoding and output the sixth feature information;
[0031] The second decoding layer is used to receive the fused feature information of the fifth feature information and the sixth feature information for decoding and output the rewritten text.
[0032] Further, the loss function of the counterfactual reasoning model is:
[0033] ;
[0034] Among them, is the first hyperparameter; is the second hyperparameter; is the third hyperparameter; is the loss function of the counterfactual reasoning model; is the contrastive learning loss; is the counterfactual consistency loss; is the counterfactual contrast loss.
[0035] Furthermore, the contrastive learning loss is:
[0036] ;
[0037] Among them, represents the cosine similarity function; represents performing a pooling operation on the output layer of the decoder to obtain a text embedding vector; represents the rewritten text output by the first generation model; represents the rewritten specialized text output by the counterfactual reasoning model; represents the specialized text;
[0038] The counterfactual consistency loss is:
[0039] ;
[0040] Among them, represents the cross-entropy loss;
[0041] The counterfactual contrast loss is:
[0042] ;
[0043] Among them, represents the boundary factor, and the boundary factor is dynamically updated during training and .
[0044] In the second aspect of the present invention, an electronic device is provided. The electronic device includes at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method of the first aspect of the present invention.
[0045] It should be understood that the content described in the Summary of the Invention section is not intended to define the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present invention will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:
[0047] Figure 1 FIG. shows a flowchart of a method for rewriting colloquial text based on counterfactual reasoning according to an embodiment of the present invention;
[0048] Figure 2 FIG. shows a schematic diagram of the principle of a method for rewriting colloquial text based on counterfactual reasoning according to an embodiment of the present invention;
[0049] Figure 3 FIG. shows a schematic diagram of an encoding and decoding network architecture according to an embodiment of the present invention;
[0050] Figure 4 FIG. shows a schematic diagram of a single decoding network architecture according to an embodiment of the present invention;
[0051] Figure 5 FIG. shows a block diagram of an exemplary electronic device capable of implementing the embodiments of the present invention;
[0052] Wherein, 500 is an electronic device, 501 is a computing unit, 502 is a ROM, 503 is a RAM, 504 is a bus, 505 is an I / O interface, 506 is an input unit, 507 is an output unit, 508 is a storage unit, and 509 is a communication unit. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0054] In addition, the term "and / or" in this document is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after.
[0055] Figure 1 The flowchart of the method for rewriting colloquial text based on counterfactual reasoning according to an embodiment of the present invention is shown.
[0056] The method includes:
[0057] S101. Construct a colloquial dataset.
[0058] The colloquial dataset includes at least specialized text and colloquial text.
[0059] On the Internet, colloquial data is very common. For example, in technical forums such as Zhihu and CSDN, there are special sections for users to ask their own questions or answer others' questions. These questions are questions raised by real users and can be classified as colloquial data, and this part of the data can be easily collected. However, in vertical fields such as education, government affairs, and healthcare, either the data volume is small or it involves user privacy and is not easily obtained. Therefore, constructing a colloquial dataset for vertical fields is a very important task.
[0060] In some embodiments, the colloquial dataset includes not only specialized text and colloquial text, but also context information.
[0061] In this embodiment, the original data is data scraped from message data and forum discussion areas publicly available on the Internet, including questions and messages for normal daily inquiries, colloquial data that conforms to daily life, etc. Generally, the scraped data will be cleaned, and only the questions raised by users will be retained. Using the Few-Shot method (small sample), a prompt template is constructed. The constraint condition can be based on colloquial examples and the given specialized text, allowing the LLM to learn the colloquial questioning method and output the colloquial text corresponding to the given specialized text to achieve the construction of the colloquial dataset. Then, the output of the model is parsed according to the constraint condition, and multiple colloquial texts in the format of imitation examples can be obtained.
[0062] In intelligent question answering in vertical fields, questions are often answered around a certain field. For example, in an intelligent question answering system in the education field, the questions raised by users are usually related to education knowledge. Therefore, introducing domain background knowledge (context information) into the model can provide additional information and constraints, reduce the ambiguity or irrelevant content that may exist in the generated text, and make the generated text better consistent with the context and closer to the target specialized text. The domain knowledge data can be the knowledge base in the intelligent question answering system, and the data type is not limited, such as knowledge graphs and document libraries. Mixed retrieval is performed through keywords and semantics, and the retrieved results are used as context information.
[0063] When the colloquial dataset includes professional texts, colloquial texts, and context information, it can be specifically expressed in a triple format.
[0064] S102. Construct a counterfactual reasoning model, use the colloquial dataset, and train the counterfactual reasoning model according to the loss function of the counterfactual reasoning model to obtain a trained counterfactual reasoning model.
[0065] In this embodiment, the counterfactual reasoning model includes at least two sequentially arranged generation models, for example, it can be three generation models or five generation models, etc.
[0066] As Figure 2 shown, in this embodiment, the structures of each generation model of the counterfactual reasoning model are the same; the difference between different generation models is only that, due to the different position orders in the counterfactual reasoning model, the input / output data are different. For example, the first generation model as the starting point obtains colloquial text, and the colloquial text is rewritten once through the first generation model, and the rewritten text is output; the remaining generation models in the counterfactual reasoning model are used to rewrite the rewritten text again and output the rewritten text again. If the generation model is the last generation model, the rewritten professional text needs to be output.
[0067] In this embodiment, since the colloquial dataset has two situations: including context information and not including context information, the counterfactual reasoning model needs to be discussed in different cases.
[0068] As an implementation manner in this embodiment, as Figure 3 and Figure 4 shown, if the colloquial dataset does not include context information, each generation model in the counterfactual reasoning model is a first encoder-decoder network or a first single-decoder network.
[0069] Specifically, the first part decoding network includes a first Embedding layer, an Encoder layer, a second Embedding layer, and a first Decoder layer. Among them, the first Embedding layer is used to perform word embedding and positional encoding on the input data and output the first feature information. The input data of the model is represented by "input text", that is, colloquial text. The Encoder layer is used to receive the first feature information and output the second feature information; the second Embedding layer is used to perform word embedding and positional encoding on the input data and output the encoded information. In the training stage, the input data of the second Embedding layer is professional text, while in the inference stage, the input data of the second Embedding layer is a start symbol, indicating the start of the sequence. The first Decoder layer is used to receive the second feature information of the Encoder layer and the encoded information of the second Embedding layer for decoding and output the rewritten text. If the rewritten text is intermediate text, it is recorded as "text_rewrite1", and then "text_rewrite1" is used as the new input and input into the next-level encoding and decoding network until the finally rewritten professional text "text_rewrite2" is output.
[0070] The parameters between the Embedding layers in each network are shared, and the parameters between the Decoder layers are shared.
[0071] Specifically, the first single-decoding network includes a fourth Embedding layer and a second Decoder layer. Among them, the fourth Embedding layer is used to perform word embedding and positional encoding on the input data and output the fifth feature information. The input data of the fourth Embedding layer is generally colloquial text; the second Decoder layer is used to receive the fifth feature information for decoding and output the rewritten text. If the rewritten text is intermediate text, it is recorded as "text_rewrite1", and then "text_rewrite1" is used as the new input and input into the next-level single-decoding network until the finally rewritten professional text "text_rewrite2" is output.
[0072] As another implementation manner in this embodiment, as Figure 3 and Figure 4 shown, if the colloquial dataset includes context information, each generation model in the counterfactual reasoning model is a second part decoding network or a second single-decoding network.
[0073] Specifically, the second-stage decoding network includes a first Embedding layer, a third Embedding layer, an Encoder layer, a second Embedding layer, and a first Decoder layer. Among them, the first Embedding layer performs word embedding and position encoding on the input data and outputs first feature information; the third Embedding layer obtains context information context for word embedding and position encoding and outputs third feature information; the obtained first feature information and third feature information are fused to obtain fused feature information; the Encoder layer receives the fused feature information of the first feature information and the third feature information and outputs fourth feature information; the second Embedding layer performs word embedding and position encoding on the input data and outputs encoded information; the first Decoder layer receives the fourth feature information of the Encoder layer and the encoded information of the second Embedding layer for decoding and outputs the rewritten text. If the rewritten text is an intermediate text, it is recorded as text_rewrite1, and then text_rewrite1 is used as the new input and input into the next-level encoding and decoding network until the finally output professionalized rewritten text text_rewrite2 is obtained.
[0074] In this embodiment, the fusion operation is as follows:
[0075] ;
[0076] Among them, , , , represents the fused feature information, represents the feature information of the colloquial text, represents the feature information of the context, represents the query weight matrix in the attention, represents the key weight matrix in the attention, represents the value weight matrix in the attention matrix, represents the dimension of the query feature, represents the normalized exponential function, and the calculation formula is: .
[0077] In this embodiment, the second single decoding network includes: a fourth embedding layer, a fifth embedding layer, and a second decoding layer; wherein, the fourth embedding layer performs word embedding and position encoding on the input data and outputs fifth feature information; the fifth embedding layer obtains context information "context" for word embedding and position encoding and outputs sixth feature information; after obtaining the fifth feature information and the sixth feature information, the fifth feature information and the sixth feature information are fused to obtain the fused feature information; the second decoding layer receives the fused feature information of the fifth feature information and the sixth feature information for decoding and outputs the rewritten text. If the rewritten text is an intermediate text, it is recorded as text_rewrite1, and then text_rewrite1 is used as a new input and input into the single decoding network at the next level until the finally rewritten professional text text_rewrite2 is output.
[0078] Whether it is a single decoding network or an encoder-decoder network, the process of the fusion operation is the same.
[0079] The encoder-decoder network deeply understands the input sequence through the encoder and extracts rich semantic information, which enables the model to more accurately capture the intent and context of the input when generating the output, making the text generated by the model based on counterfactual assumptions more consistent with the original semantics, and can effectively alleviate the semantic gap problem between colloquial texts and professional texts.
[0080] The single decoding network omits the encoder, and the input sequence and the output sequence are processed in the same module, avoiding the complication of the model structure. During the inference process, only one forward propagation is required, and the inference efficiency is higher. Due to the simplification of the structure, the model can add a multi-order counterfactual reasoning structure, and using multiple orders can enable the model to learn deep semantic information and improve the rewriting effect of the model.
[0081] S103. Input the colloquial text into the trained counterfactual reasoning model to obtain the rewritten professional text.
[0082] As Figure 2 shown, the generation model first performs a rewrite on the colloquial language, and the rewritten text is recorded as text_rewrite1. Then text_rewrite1 is sent into the generation model again as a counterfactual input (note that the parameters of the generation model and are shared and belong to the same model), and the rewritten text is obtained, recorded as text_rewrite2. Based on the two rewritten texts generated and the known professional text (recorded as text_groundtruth), a multi-task loss function is designed accordingly.
[0083] Specifically, the decoder of the generative model is used as the feature extractor of the text. Specifically, the last-token pooling method is adopted. In order to make the rewritten text have a high similarity in semantic features with the professional text, a contrastive learning loss is adopted, as follows:
[0084] ;
[0085] Among them, represents the cosine similarity function, and the formula is ; represents performing a pooling operation on the output layer of the decoder to obtain the text embedding vector. The last-token pooling method is adopted in this method.
[0086] In order to make the rewritten text consistent with the professional text in some aspects (such as semantic content, important information, style, etc.), a counterfactual consistency loss is constructed. Specifically:
[0087] ;
[0088] Among them, represents the cross-entropy loss; represents the rewritten text output by the first generative model; represents the rewritten professional text output by the counterfactual reasoning model; represents the professional text.
[0089] The cross-entropy loss is expressed as:
[0090] ;
[0091] Among them, represents the distribution of the real text, represents the distribution of the predicted text.
[0092] Based on the counterfactual hypothesis, a counterfactual contrast loss is constructed by forcing the similarity between text_rewrite2 and the professional text to be higher than the similarity between text_rewrite1 and the professional text. Specifically:
[0093] Among them, represents the cosine similarity function, and the formula is ; represents performing a pooling operation on the output layer of the decoder to obtain the text embedding vector. The last-token pooling method is adopted in this method. , represents the boundary factor, which is dynamically updated during the training process. The update function is as follows:
[0094]
[0095] Among them, represents the initial , represents the current training round, represents the total number of training rounds.
[0096] In summary, the total loss function is
[0097] ;
[0098] Among them, is the first hyperparameter; is the second hyperparameter; is the third hyperparameter; is the loss function of the counterfactual reasoning model; is the contrastive learning loss; is the counterfactual consistency loss; is the counterfactual contrast loss.
[0099] In this article, the context information is an optional item. Using the context information will have a beneficial impact on the learning of the model, but the cost is an increase in the amount of computation. It can be selected whether to use it according to the actual domain knowledge, which is more flexible. The context information is visible to the generative model and invisible to the loss function, so it will not directly affect the loss function. That is to say, whether the generative model is an encoder-decoder network or a single-decoder network, the total loss function is used as the loss function for training and optimization.
[0100] According to the embodiments of the present invention, based on the counterfactual reasoning hypothesis, multi-level counterfactual modeling is introduced, that is, instead of a single rewrite, a multi-step progressive rewrite method is introduced. For example, in this article, a two-level counterfactual modeling from colloquial language to formal language and then to professional language is used, enabling the model to more clearly learn the contributions of language style, tone, and grammar to professional expression, and reducing the difficulty of the model learning from "colloquial" to "professional".
[0101] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0102] In the technical solution of the present invention, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0103] According to an embodiment of the present invention, the present invention further provides an electronic device.
[0104] Figure 5 FIG. shows a schematic block diagram of an electronic device 500 that can be used to implement embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0105] The electronic device 500 includes a computing unit 501 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0106] A plurality of components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0107] The computing unit 501 may be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above, such as methods S101 to S103. For example, in some embodiments, methods S101 to S103 may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of methods S101 to S103 described above may be executed. Alternatively, in other embodiments, the computing unit 501 may be configured to execute methods S101 to S103 in any other suitable manner (e.g., by means of firmware).
[0108] Various embodiments of the systems and techniques described above in this document may be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-a-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0109] The program code for implementing the methods of the present invention may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes may be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0110] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0111] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0112] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0113] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0114] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.
[0115] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for rewriting colloquial text based on counterfactual reasoning, characterized in that: include: Constructing a spoken language dataset, wherein the spoken language dataset includes professional texts and spoken language texts; Constructing a counterfactual reasoning model, using a spoken language dataset, and training the counterfactual reasoning model according to a loss function of the counterfactual reasoning model to obtain a trained counterfactual reasoning model; Input the colloquial text into the trained counterfactual reasoning model to obtain the rewritten professional text; The counterfactual reasoning model includes at least two generation models arranged in sequence; and each generation model of the counterfactual reasoning model has the same structure; wherein the first generation model is used to rewrite the colloquial text once and output the rewritten text; the remaining generation models in the counterfactual reasoning model are used to rewrite the rewritten text again and output the rewritten text; The loss function of the counterfactual reasoning model is: in, is the first hyperparameter; is the second hyperparameter; is the third hyperparameter; is the loss function of the counterfactual reasoning model; To contrast learning loss; is the counterfactual consistency loss; is the counterfactual comparison loss; The contrastive learning loss is: in, represents the cosine similarity function; Indicates that a pooling operation is performed on the output layer of the decoder to obtain a text embedding vector; represents the rewritten text output by the first generation model; A rewritten specialized text representing the output of the counterfactual reasoning model; Indicates specialized text; The counterfactual consistency loss is: in, represents the cross entropy loss; The counterfactual contrast loss is: in, Represents the boundary factor, which is dynamically updated during the training process and .
2. The method according to claim 1, characterized in that: If the spoken data set does not include context information, each generative model in the counterfactual reasoning model is a first encoding-decoding network or a first single decoding network.
3. The method according to claim 2, characterized in that The first encoding and decoding network includes: The first embedding layer is used to embed words and position encode the input data and output first feature information; The encoding layer is used to receive the first feature information and output the second feature information; The second embedding layer is used to embed words and position encode the input data and output the encoded information; The first decoding layer is used to receive the second feature information of the encoding layer and the encoded information of the second embedding layer, decode them, and output the rewritten text.
4. The method according to claim 2, characterized in that: The first single decoding network comprises: The fourth embedding layer is used to embed words and position encode the input data and output fifth feature information; The second decoding layer is used to receive the fifth feature information for decoding and output the rewritten text.
5. The method according to claim 1, characterized in that If the spoken data set includes context information, each generative model in the counterfactual reasoning model is a second encoding-decoding network or a second single decoding network.
6. The method according to claim 5, characterized in that The second encoding and decoding network includes: The first embedding layer is used to embed words and position encode the input data and output first feature information; The third embedding layer is used to obtain context information for word embedding and position encoding, and output the third feature information; The coding layer is used to receive the feature information obtained by fusing the first feature information and the third feature information, and output the fourth feature information; The second embedding layer is used to embed words and position encode the input data and output the encoded information; The first decoding layer is used to receive the fourth feature information of the encoding layer and the encoded information of the second embedding layer, decode them, and output the rewritten text.
7. The method according to claim 5, characterized in that The second single decoding network comprises: The fourth embedding layer is used to embed words and position encode the input data and output fifth feature information; The fifth embedding layer is used to obtain context information for word embedding and position encoding, and output sixth feature information; The second decoding layer is used to receive the feature information obtained by merging the fifth feature information and the sixth feature information, perform decoding, and output the rewritten text.
8. An electronic device comprising at least one processor; and a memory connected in communication with the at least one processor; characterized in that: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Text style migration model training method and device and text style migration method and device
CN114818728A