A text correction method, device and related equipment

By constructing a knowledge base to compare and correct the initial triplet to generate the target paragraph, the problem of existing technologies being unable to assess the authenticity of text and correct low-quality articles is solved, thereby improving text quality.

CN116070617BActive Publication Date: 2026-01-09IFLYTEK SOUTH CHINA ARTIFICIAL INTELLIGENCE RES INST GUANGZHOU CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211714345.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-29
Publication Date
2026-01-09
Estimated Expiration
2042-12-29

AI Technical Summary

Technical Problem

Existing text quality assessment methods are unable to effectively evaluate the authenticity and reliability of text information, and cannot correct low-quality articles.

Method used

By constructing a knowledge base, the initial triples in the initial paragraph are compared with reference triples to correct erroneous initial triples, generate target triples, and generate target paragraphs to replace the initial paragraphs based on the semantic information of the target triples and adjacent paragraphs.

Benefits of technology

It improved the accuracy and authenticity of the text, thus enhancing its quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116070617B_ABST
    Figure CN116070617B_ABST
Patent Text Reader

Abstract

The application discloses a text correction method and device and related equipment, and the method comprises the following steps: obtaining a basic text and a pre-constructed knowledge base; wherein the basic text contains at least one initial paragraph, and the knowledge base contains a plurality of reference triples related to the basic text; for each initial paragraph, obtaining an initial triple in the current initial paragraph, determining the accuracy of the current initial paragraph based on the initial triple and the reference triple in the knowledge base; in response to the accuracy being less than or equal to a preset threshold, correcting at least part of the initial triple in the current initial paragraph based on the reference triple to obtain a target triple; based on the current initial paragraph, the target triple and the remaining initial paragraphs adjacent to the current initial paragraph, obtaining a target paragraph, and replacing the current initial paragraph with the target paragraph. Through the above method, the application can correct the text with poor quality to improve the accuracy and authenticity of the text.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of natural language processing, in particular to a text correction method and device and related equipment. BACKGROUND

[0002] With the rapid development of science and technology and network technology, more and more people obtain new information through network articles. However, there are also many low-quality articles mixed with false information in the network, which hinders people's rapid acquisition of effective information. Therefore, it is of great significance to evaluate the quality of text content. The existing text quality evaluation method generally only focuses on the syntax level, emotion level or theme level of the text, and cannot evaluate the authenticity and reliability of the text information. Moreover, the low-quality articles cannot be corrected. SUMMARY

[0003] The technical problem solved by the present application is to provide a text correction method, device and related equipment, which can correct the text with poor quality to improve the accuracy and authenticity of the text.

[0004] To solve the above technical problems, one technical solution adopted by the present application is to provide a text correction method, comprising: obtaining a basic text and a pre-constructed knowledge base; wherein the basic text contains at least one initial paragraph, and the knowledge base contains a plurality of reference triples related to the basic text; for each initial paragraph, obtaining an initial triple in the current initial paragraph, determining the accuracy of the current initial paragraph based on the initial triple and the reference triple in the knowledge base; in response to the accuracy being less than or equal to a preset threshold, modifying at least part of the initial triples in the current initial paragraph based on the reference triples to obtain target triples; obtaining a target paragraph based on the current initial paragraph, the target triples and the remaining initial paragraphs adjacent to the current initial paragraph, and replacing the current initial paragraph with the target paragraph.

[0005] To solve the above technical problems, another technical solution adopted by the present application is to provide a text correction device, comprising: a first obtaining module for obtaining a basic text and a pre-constructed knowledge base; wherein the basic text contains at least one initial paragraph, and the knowledge base contains a plurality of reference triples related to the basic text; a second obtaining module for obtaining, for each initial paragraph, an initial triple in the current initial paragraph, determining the accuracy of the current initial paragraph based on the initial triple and the reference triple in the knowledge base; a correction module for, in response to the accuracy being less than or equal to a preset threshold, correcting at least part of the initial triples in the current initial paragraph based on the reference triples to obtain target triples; and a processing module for obtaining a target paragraph based on the current initial paragraph, the target triples, and the remaining initial paragraphs adjacent to the current initial paragraph, and replacing the current initial paragraph with the target paragraph.

[0006] To solve the above technical problems, another technical solution adopted by the present application is to provide an electronic device comprising a memory and a processor coupled to each other, the memory storing program instructions, and the processor being configured to execute the program instructions to implement the text correction method mentioned in the above technical solution.

[0007] To solve the above technical problems, another technical solution adopted by the present application is to provide a computer-readable storage medium storing program instructions executable by a processor, the program instructions being configured to implement the text correction method mentioned in the above technical solution

[0008] The beneficial effects of the present application are: different from the prior art, the text correction method proposed in the present application evaluates the text quality of the current initial paragraph according to the initial triple extracted from the current initial paragraph. For the initial paragraph with low text quality, the reference triples in the knowledge base are used to correct the incorrect initial triples, and the corresponding target paragraph is regenerated according to the semantic information of the triples in the corrected initial paragraph, the initial paragraph, and the part of the paragraph adjacent to the initial paragraph. By replacing the corresponding initial paragraph with the target paragraph, the quality and accuracy of the basic text are improved. BRIEF DESCRIPTION OF DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor. Among them:

[0010] Figure 1is a flowchart of an embodiment of the text correction method of the present application;

[0011] Figure 2 is a flowchart of an embodiment corresponding to step S104;

[0012] Figure 3 is a structural diagram of an embodiment of the text correction model corresponding to step S104;

[0013] Figure 4 is a structural diagram of an embodiment of the text correction device of the present application;

[0014] Figure 5 is a structural diagram of an embodiment of the electronic device of the present application;

[0015] Figure 6 is a structural diagram of an embodiment of the storage device of the present application. DETAILED DESCRIPTION

[0016] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0017] Please refer to Figure 1 , Figure 1 is a flowchart of an embodiment of the text correction method of the present application, which comprises:

[0018] S101: Obtain a basic text and a pre-constructed knowledge base. The basic text contains at least one initial paragraph, and the knowledge base contains a plurality of reference triples related to the basic text.

[0019] In an embodiment, step S101 comprises: obtaining a basic text that needs to be evaluated for text quality, which can be a product marketing copy, a network news, or an article generated by a text generation model, etc. The basic text includes at least one initial paragraph.

[0020] Further, in order to assist in evaluating the text quality of the initial paragraph in the basic text, a knowledge base containing a plurality of reference triples is constructed according to the content information corresponding to the basic text to be evaluated, or the field to which the basic text belongs. For example, in response to the basic text to be evaluated being a news article in the sports field and being closely related to basketball, the knowledge base contains a large number of reference triples related to the basketball field, such as "Yao's height-226cm".

[0021] S102: For each initial paragraph, obtain an initial triple in the current initial paragraph, and determine the accuracy of the current initial paragraph based on the initial triple and a reference triple in the knowledge base.

[0022] In an embodiment, step S102 comprises: for all initial paragraphs in the to-be-evaluated base text, sequentially performing quality evaluation on each initial paragraph to determine whether the current initial paragraph meets the predetermined quality requirement. When performing quality evaluation on the current initial paragraph, all initial triples are first extracted from the current initial paragraph.

[0023] Specifically, the extracted initial triple and the reference triple both include a corresponding subject element, an intermediate element, and an object element, and the intermediate element can be the relationship between the subject element and the object element. The OIE (Open Information Extraction) technology can be used to extract multiple initial triples from the current initial paragraph. For example, for the triple "China - capital - Beijing", the subject element is "China", the object element is "Beijing", and the intermediate element is "capital".

[0024] Alternatively, in other embodiments, an extraction template can also be constructed according to the multiple reference triples in the knowledge base, and the extraction template is used to extract the initial triple from the current initial paragraph. For example, the subject element and the intermediate element of the reference triple are obtained, and in response to the appearance of the subject element and the intermediate element of the reference triple in a certain sentence in the current initial paragraph, the corresponding object element is determined based on the semantic information of the sentence to obtain the corresponding initial triple.

[0025] Further, the initial triple is compared with the reference triple in the knowledge base to determine the accuracy of the current initial paragraph. The higher the accuracy, the higher the quality of the current initial paragraph, and no correction is needed; if the accuracy is low, it is considered that there are many factual errors in the current initial paragraph, and in order to ensure the quality of the base text, the current initial paragraph needs to be corrected.

[0026] In an embodiment, the calculation method of the accuracy of the current initial paragraph can be: obtaining a first number of all initial triples in the current initial paragraph. And, obtaining a second number of initial triples identical to the reference triple in the knowledge base.

[0027] Specifically, in response to the reference triple in the knowledge base being identical to the initial triple, it is considered that the initial triple is a correct initial triple, and the first number of all initial triples in the current initial paragraph and the second number of correct initial triples in the initial paragraph are obtained.

[0028] Further, a ratio of the second quantity and the first quantity is taken as the accuracy of the current initial paragraph. The specific calculation formula is as follows:

[0029]

[0030] wherein, the accuracy of the current initial paragraph is represented by, the second quantity is represented by, the first quantity is represented by.

[0031] S103: In response to the accuracy being less than or equal to a preset threshold, at least part of the initial triples in the current initial paragraph are corrected based on the reference triples to obtain target triples.

[0032] In an embodiment, the step S103 includes: comparing the accuracy of the current initial paragraph with a preset threshold, and in response to the accuracy being less than or equal to the preset threshold, correcting the initial triples with errors in the current initial paragraph. The preset threshold can be estimated by relevant personnel or obtained by multiple experiments.

[0033] Specifically, in response to the initial triples containing a subject element, an intermediate element and an object element, it is determined whether the object element is correct based on the subject element and the intermediate element. If not, it is considered that the corresponding initial triple contains incorrect information, and the candidate reference triples are obtained from the knowledge base based on the subject element and the intermediate element, so as to replace the corresponding initial triple with the candidate reference triples. All the triples in the current initial paragraph after the replacement are taken as the target triples. That is, the target triples contain both the correct initial triples in the current initial paragraph and the corrected initial triples. By replacing the incorrect initial triples with the reference triples, the incorrect information in the current initial paragraph is corrected, thereby improving the accuracy of the information in the current initial paragraph.

[0034] For example, when the initial triples "China - capital - Shanghai" are included in the current initial paragraph, it is determined that the object element "Shanghai" is incorrect based on the semantic information of the subject element "China" and the intermediate element "capital". Then, the reference triples "China - capital - Beijing" are obtained from the knowledge base based on the subject element "China" and the intermediate element "capital", and the corresponding initial triples are replaced with the reference triples.

[0035] Alternatively, in other embodiments, the above-mentioned manner of determining whether the initial triples contain errors can also be determining whether the remaining element is correct based on any two of the subject element, the intermediate element and the object element. If correct, it indicates that the initial triple does not contain errors.

[0036] S104: obtaining a target paragraph based on the current initial paragraph, the target triple, and the remaining initial paragraphs adjacent to the current initial paragraph, and replacing the current initial paragraph with the target paragraph.

[0037] In an embodiment, referring to Figure 2 and Figure 3 , Figure 2 is a flowchart of an embodiment corresponding to step S104, Figure 3 is a structural diagram of an embodiment of a text correction model corresponding to step S104. Step S104 specifically includes:

[0038] S201: obtaining a full-text semantic vector corresponding to the base text based on all initial paragraphs, and obtaining a preceding-text semantic vector corresponding to all initial paragraphs before the current initial paragraph.

[0039] In an embodiment, referring to Figure 3 , step S201 includes: inputting all initial paragraphs into a first semantic extraction network 10 in the text correction model in order of the base text, and the first semantic extraction network 10 performs semantic extraction on the inputted all initial paragraphs to obtain a full-text semantic vector corresponding to the base text, which contains semantic information of the complete base text.

[0040] Specifically, in response to the first semantic extraction network 10 being a Bert model, the specific calculation formula of the full-text semantic vector is as follows:

[0041]

[0042] wherein, denotes the full-text semantic vector, denotes a feature transformation function, denotes all initial paragraphs in the base text. Wherein, in response to the base text containing N initial paragraphs, .

[0043] Optionally, in other embodiments, the above-mentioned first semantic extraction network can also include other commonly used text semantic extraction networks, such as LSTM model, etc.

[0044] Further, referring to Figure 3 , step S201 further includes inputting all initial paragraphs before the current initial paragraph into a second semantic extraction network 20 in the text correction model to obtain a preceding-text semantic vector corresponding to the current initial paragraph. Wherein, the second semantic extraction network 20 is also a Bert model.

[0045] In an embodiment, in response to the current initial paragraph being the i-th paragraph in the base text, the specific calculation formula of the preceding-text semantic vector is as follows:

[0046]

[0047] wherein, denotes the above semantic vector, denotes a feature transformation function, denotes the i-th initial paragraph in the base text.

[0048] Optionally, in other embodiments, step S201 can also obtain a full-text semantic vector corresponding to all initial paragraphs in the base text, and obtain a post-text semantic vector corresponding to all initial paragraphs after the current initial paragraph. Alternatively, a preset number of initial paragraphs before the current initial paragraph and a preset number of initial paragraphs after the current initial paragraph are taken as relevant initial paragraphs, and a relevant semantic vector corresponding to the relevant initial paragraphs is obtained.

[0049] S202: Replacing at least part of the elements in the target triple with the corresponding element category to obtain a paragraph semantic vector of the replaced current initial paragraph.

[0050] In an embodiment, please continue to refer to Figure 3 , step S202 includes: in response to the current initial paragraph containing at least part of the target triple. For each target triple, obtaining the element category to which each entity element in the target triple belongs, and replacing the corresponding subject element or object element in the current initial paragraph with the element category. For example, for the target triple "China - capital - Beijing" in the current initial paragraph, "country" is replaced with "China", and "city" is replaced with "Beijing".

[0051] Optionally, a word classification model can be obtained by training to classify each element in the triple. The specific process can be realized by existing open source algorithm, which will not be described in detail here.

[0052] Further, in response to completing the replacement of all target triples in the current initial paragraph, the replaced current initial paragraph is processed for word segmentation to obtain a plurality of words corresponding to the replaced current initial paragraph, and a word vector corresponding to each word in the current initial paragraph is obtained through a semantic feature extraction technique.

[0053] ​Further, the word vector is sequentially input into the third semantic extraction network 30 in the text correction model according to the position information of the corresponding word in the current initial paragraph, to obtain the paragraph semantic vector corresponding to the current initial paragraph. The third semantic extraction network can be a BiGRU model. The semantic information corresponding to all words in the current initial paragraph is fused through the third semantic extraction network, so that the obtained paragraph semantic vector contains not only the forward semantic information of the current initial paragraph, but also the reverse semantic information of the current initial paragraph, to prevent the obtained paragraph semantic vector from ignoring part of the semantic information.

[0054] Optionally, in other embodiments, the third semantic extraction network 30 can also be a BiLSTM model, etc.

[0055] In an embodiment, in response to the current initial paragraph containing words after replacement, the specific calculation formula of the paragraph semantic vector is as follows:

[0056]

[0057] wherein, represents the paragraph semantic vector obtained after inputting the word vector into the third semantic extraction network, represents the word vector corresponding to the word. Wherein, is less than or equal to .

[0058] In other embodiments, step S202 can also replace the corrected target triple only with the corresponding element category, and obtain the paragraph semantic vector based on the current initial paragraph after replacement. That is, the correct initial triple in the current initial paragraph is not replaced.

[0059] S203: Obtain the first triple vector corresponding to all target triples, and take the mean value of all first triple vectors as the target triple vector.

[0060] In an embodiment, please continue to refer to Figure 3 , step S203 includes: in response to the current initial paragraph containing at least part of the target triple, performing semantic feature extraction on each target triple to obtain the semantic vector corresponding to each target triple. The obtained semantic vectors corresponding to all target triples are sequentially input into the fourth semantic extraction network 40 in the text correction model, to obtain the first triple vector corresponding to each target triple. The fourth semantic extraction network 40 can be a BiGRU model or a BiLSTM model, etc.

[0061] Furthermore, the average of all first triplet vectors corresponding to the current initial paragraph is taken to obtain the target triplet vector corresponding to the current initial paragraph.

[0062] In one implementation, in response to the current initial paragraph containing The specific formula for calculating the target triplet vector is as follows:

[0063]

[0064]

[0065] in, Indicates the first paragraph in the current initial paragraph. The semantic vector corresponding to each target triple. Indicates the first Each target triplet corresponds to the first triplet vector. This represents the target triple vector.

[0066] S204: Obtain the target paragraph based on the full-text semantic vector, the preceding text semantic vector, the paragraph semantic vector, and the target triple vector.

[0067] In one implementation, please refer to... Figure 3 Step S204 includes: inputting the full-text semantic vector, the preceding text semantic vector, the paragraph semantic vector, and the target triplet vector obtained through the above steps into the decoding network 50 in the text correction model to obtain the decoding vector output by each node. The decoding network 50 contains multiple hidden layers, each corresponding to a node.

[0068] Furthermore, multiple candidate words are obtained, and the decoding vectors corresponding to each node are normalized to obtain the confidence scores of multiple candidate words corresponding to each decoding vector. The candidate word with the highest confidence score is taken as the target word of the corresponding node.

[0069] Specifically, a candidate word library containing multiple candidate words is constructed. The full-text semantic vector, the preceding text semantic vector, the paragraph semantic vector, and the target triple vector are input into the decoding network 50 to obtain the target word vector and hidden layer state vector output by each node. Based on each hidden layer state vector, the corresponding decoding vector is obtained. The decoding vector is normalized using the softmax algorithm to obtain the confidence score of each candidate word in the candidate word library corresponding to the decoding vector of each node. The candidate word corresponding to the highest confidence score is taken as the target word for each node. All obtained target words are arranged according to the node's order to obtain the target text.

[0070] In an embodiment, the formula for obtaining the confidence distribution of each candidate word corresponding to the output of different nodes is as follows:

[0071]

[0072]

[0073] wherein, represents the hidden layer state vector at the current moment, represents the target word vector obtained at the moment; represents the matrix parameter, represents the decoding vector at the current moment, represents the decoding vector at the current moment, and the decoding vector corresponds to the confidence distribution of each candidate word in the candidate word library.

[0074] In the embodiment, by decoding the full-text semantic vector, the previous semantic vector, the paragraph semantic vector, and the target triple vector, the obtained decoding vector contains multiple levels of semantic information, and the obtained target paragraph is closely related to the basic text and has good readability and reliability. In addition, in the embodiment, the decoding network 50 is a GRU model. Of course, in other embodiments, the decoding network can also be other semantic decoding models, such as an LSTM model.

[0075] The text correction method proposed in the application evaluates the text quality of the current initial paragraph according to the initial triple extracted from the current initial paragraph. For the initial paragraph with low text quality, the incorrect initial triple is corrected by using the reference triple in the knowledge base, and the corresponding target paragraph is regenerated according to the semantic information of the triple in the corrected initial paragraph, the initial paragraph, and the adjacent part of the paragraph. The quality and accuracy of the basic text are improved by replacing the corresponding initial paragraph with the target paragraph.

[0076] In another embodiment, Figure 1The step S102 of determining the accuracy of the current initial paragraph can also include: in response to obtaining the first number of all initial triples in the current paragraph and the second number of correct initial triples, taking the difference between the first number and the second number as a third number of incorrect initial triples in the current initial paragraph. A number threshold is determined based on the length of the current initial text. If the third number is greater than the number threshold, it is considered that the accuracy of the current initial paragraph is low, and the current initial paragraph needs to be corrected; if the third number is less than the number threshold, it is considered that the accuracy of the current initial paragraph is high, and the current initial paragraph does not need to be corrected. The length of the current initial paragraph is positively correlated with the number of words in the current initial paragraph, and the number threshold is positively correlated with the length of the current initial paragraph.

[0077] In yet another embodiment, in response to performing Figure 1 After the step S104 of replacing the current initial paragraph with the obtained target paragraph, it is detected whether the current initial paragraph is located at the end of the base text, i.e., whether the current initial paragraph is the last paragraph of the base text. If yes, the text correction is stopped and the corrected target text of the base text is obtained. If no, an initial paragraph adjacent to and located after the current initial paragraph is taken as the current initial paragraph, and the steps of obtaining initial triples in the current initial paragraph and determining the accuracy of the current initial paragraph based on the initial triples and the reference triples in the knowledge base are returned to. That is, the first paragraph after the current initial paragraph is taken as a new current initial paragraph, and the related steps in the steps S102 to S104 are re-executed until all paragraphs in the base text are corrected to obtain a target text with high reliability. Figure 1 After the step S104 of replacing the current initial paragraph with the obtained target paragraph, it is detected whether the current initial paragraph is located at the end of the base text, i.e., whether the current initial paragraph is the last paragraph of the base text. If yes, the text correction is stopped and the corrected target text of the base text is obtained. If no, an initial paragraph adjacent to and located after the current initial paragraph is taken as the current initial paragraph, and the steps of obtaining initial triples in the current initial paragraph and determining the accuracy of the current initial paragraph based on the initial triples and the reference triples in the knowledge base are returned to. That is, the first paragraph after the current initial paragraph is taken as a new current initial paragraph, and the related steps in the steps S102 to S104 are re-executed until all paragraphs in the base text are corrected to obtain a target text with high reliability.

[0078] Please refer to Figure 4 , Figure 4 is a structural schematic diagram of an embodiment of the text correction device of the present application. The text correction device comprises a first obtaining module 60, a second obtaining module 70, a correction module 80 and a processing module 90.

[0079] Specifically, the first obtaining module 60 is configured to obtain a base text and a pre-constructed knowledge base; wherein the base text comprises at least one initial paragraph, and the knowledge base comprises a plurality of reference triples related to the base text.

[0080] The second obtaining module 70 is configured to, for each initial paragraph, obtain initial triples in the current initial paragraph, and determine the accuracy of the current initial paragraph based on the initial triples and the reference triples in the knowledge base.

[0081] The step of determining the accuracy of the current initial paragraph based on the initial triple and the reference triple in the knowledge base comprises: obtaining a first quantity of all initial triples in the current initial paragraph; obtaining a second quantity of initial triples identical to the reference triple in the knowledge base; and taking the ratio of the second quantity to the first quantity as the accuracy.

[0082] The correction module 80 is configured to, in response to the accuracy being less than or equal to a preset threshold, correct at least part of the initial triples in the current initial paragraph based on the reference triple to obtain target triples.

[0083] The step of correcting at least part of the initial triples in the current initial paragraph based on the reference triple comprises: in response to the initial triple containing a subject element, an intermediate element and an object element, determining whether the object element is correct based on the subject element and the intermediate element; if not, obtaining an alternative reference triple from the knowledge base based on the subject element and the intermediate element, and replacing the corresponding initial triple with the alternative reference triple.

[0084] The processing module 90 is configured to obtain a target paragraph based on the current initial paragraph, the target triple and the remaining initial paragraphs adjacent to the current initial paragraph, and replace the current initial paragraph with the target paragraph.

[0085] The step of obtaining the target paragraph based on the current initial paragraph, the target triple and the remaining initial paragraphs adjacent to the current initial paragraph comprises: obtaining a text semantic vector corresponding to the base text based on all initial paragraphs, and obtaining a preceding semantic vector corresponding to all initial paragraphs before the current initial paragraph; replacing the target triple with a corresponding element category to obtain a paragraph semantic vector of the replaced current initial paragraph; obtaining a first triple vector corresponding to all target triples, and taking the mean of all first triple vectors as the target triple; and obtaining the target paragraph based on the full-text semantic vector, the preceding semantic vector, the paragraph semantic vector and the target triple vector.

[0086] The step of replacing the target triple with a corresponding element category to obtain a paragraph semantic vector of the replaced current initial paragraph comprises: in response to the current initial paragraph containing multiple target triples, obtaining an element category to which a subject element and an object element in the target triple belong; replacing the corresponding subject element or object element with the element category; and performing semantic feature extraction on the replaced current initial paragraph to obtain a paragraph semantic vector.

[0087] The step of obtaining the target paragraph based on the full-text semantic vector, the preceding semantic vector, the paragraph semantic vector and the target triple vector comprises: inputting the full-text semantic vector, the preceding semantic vector, the paragraph semantic vector and the target triple vector into a decoding network to obtain a decoding vector output by each node; wherein the decoding network comprises a plurality of hidden layers, and each hidden layer corresponds to a node; obtaining a plurality of candidate words, performing normalization processing on the decoding vector to obtain a confidence score of the plurality of candidate words corresponding to the node, and taking a candidate word corresponding to the maximum confidence score as a target word of the corresponding node; and obtaining the target paragraph based on the target word corresponding to each node.

[0088] Please continue to refer to the drawings. The text correction device provided in the application further comprises a detection module 100 coupled with the processing module 90, configured to detect whether the current initial paragraph is located at the end of the basic text; if not, an initial paragraph adjacent to and located after the current initial paragraph is taken as the current initial paragraph, and the step of obtaining an initial triple in the current initial paragraph and determining the accuracy of the initial paragraph based on the initial triple and the reference triple in the knowledge base is returned to.

[0089] Please refer to Figure 5 , Figure 5 is a structural schematic diagram of an embodiment of an electronic device. The electronic device comprises a memory 110 and a processor 120 coupled with each other, the memory 110 stores program instructions, and the processor 120 is configured to execute the program instructions to implement the method in any of the above embodiments. Specifically, the electronic device includes but is not limited to a desktop computer, a notebook computer, a tablet computer, a server, etc., which are not limited herein. In addition, the processor 120 can also be referred to as a CPU (Center Processing Unit). The processor 120 can be an integrated circuit chip with signal processing capability. The processor 120 can also be a general processor, a digital signal processor (Digital Signal Processor, DSP), an application specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general processor can be a microprocessor or the processor can also be any conventional processor. In addition, the processor 120 can be realized by integrated circuit chips together.

[0090] Please refer to Figure 6 , Figure 6An implementation of the storage device of the present application is shown in a structural schematic diagram. The storage device 130 stores program instructions 140 that can be executed by the processor. The program instructions 140 are used to implement the method in any of the embodiments described above.

[0091] In several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only schematic, for example, the division of the modules or units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed units can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms.

[0092] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e. can be located in one place or distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment scheme.

[0093] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0094] If the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method of each embodiment of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0095] The above merely describes the embodiments of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation, or direct or indirect application in other related technical fields, which is made by using the content of the present application specification and drawings, is also included in the patent protection scope of the present application.

Claims

1. A text correction method, characterized in that, include: Obtain a base text and a pre-built knowledge base; wherein the base text contains at least one initial paragraph, and the knowledge base contains a number of reference triples related to the base text; For each initial paragraph, obtain the initial triples in the current initial paragraph, and determine the accuracy of the current initial paragraph based on the initial triples and the reference triples in the knowledge base; In response to the accuracy being less than or equal to a preset threshold, at least a portion of the initial triples in the current initial paragraph are corrected based on the reference triples to obtain the target triples; Based on the current initial paragraph, the target triplet, and the other initial paragraphs adjacent to the current initial paragraph, a target paragraph is obtained, and the current initial paragraph is replaced with the target paragraph; The step of obtaining the target paragraph based on the current initial paragraph, the target triplet, and the other initial paragraphs adjacent to the current initial paragraph includes: obtaining the full-text semantic vector corresponding to the base text based on all the initial paragraphs, and obtaining the context semantic vector corresponding to all the initial paragraphs preceding the current initial paragraph; replacing at least some elements in the target triplet with the corresponding element categories to obtain the paragraph semantic vector of the current initial paragraph after replacement; obtaining the first triplet vector corresponding to all the target triplets, and using the mean of all the first triplet vectors as the target triplet vector; and using a decoding network to decode and obtain the target paragraph based on the full-text semantic vector, the context semantic vector, the paragraph semantic vector, and the target triplet vector.

2. The method according to claim 1, characterized in that, The step of determining the accuracy of the current initial paragraph based on the initial triplet and the reference triplet in the knowledge base includes: Obtain the first count of all the initial triplets in the current initial paragraph; Obtain a second number of the initial triplets that are identical to the reference triplets in the knowledge base; The ratio of the second quantity to the first quantity is taken as the accuracy rate.

3. The method according to claim 1, characterized in that, The step of modifying at least a portion of the initial triples in the current initial paragraph based on the reference triples includes: In response to the initial triple containing a subject element, an intermediate element, and an object element, the correctness of the object element is determined based on the subject element and the intermediate element. If not, then alternative reference triples are obtained from the knowledge base based on the main element and the intermediate element, and the alternative reference triples are used to replace the corresponding initial triples.

4. The method according to claim 1, characterized in that, The step of replacing the target triplet with the corresponding element category to obtain the paragraph semantic vector of the current initial paragraph includes: In response to the fact that the current initial paragraph contains multiple target triples, the element categories to which the subject element and object element in the target triples belong are obtained; Replace the corresponding subject element or object element with the element category; Semantic features are extracted from the replaced initial paragraph to obtain the paragraph semantic vector.

5. The method according to claim 1, characterized in that, The steps for obtaining the target paragraph include: The full-text semantic vector, the preceding text semantic vector, the paragraph semantic vector, and the target triple vector are input into the decoding network to obtain the decoding vectors output by each node; wherein, the decoding network contains multiple hidden layers, and each hidden layer corresponds to one node; Multiple candidate words are obtained, the decoding vector is normalized, and the confidence scores of the multiple candidate words corresponding to the node are obtained. The candidate word corresponding to the candidate word with the largest confidence score is taken as the target word of the corresponding node. The target paragraph is obtained based on the target word corresponding to each of the nodes.

6. The method according to claim 1, characterized in that, After the step of replacing the current initial paragraph with the target paragraph, the following is included: Detect whether the current initial paragraph is at the end of the base text; If not, the initial paragraph that is adjacent to and follows the current initial paragraph is taken as the current initial paragraph, and the process returns to the step of obtaining the initial triples in the current initial paragraph and determining the accuracy of the current initial paragraph based on the initial triples and the reference triples in the knowledge base.

7. A text correction device, characterized in that, include: The first acquisition module is used to acquire a base text and a pre-built knowledge base; wherein the base text contains at least one initial paragraph, and the knowledge base contains a number of reference triples related to the base text; The second obtaining module is used to obtain the initial triples in the current initial paragraph for each initial paragraph, and to determine the accuracy of the current initial paragraph based on the initial triples and the reference triples in the knowledge base; The correction module is configured to, in response to the accuracy being less than or equal to a preset threshold, correct at least a portion of the initial triplets in the current initial paragraph based on the reference triplets to obtain the target triplets; The processing module is used to obtain a target paragraph based on the current initial paragraph, the target triplet, and the other initial paragraphs adjacent to the current initial paragraph, and replace the current initial paragraph with the target paragraph; The step of obtaining the target paragraph based on the current initial paragraph, the target triplet, and the other initial paragraphs adjacent to the current initial paragraph includes: obtaining the full-text semantic vector corresponding to the base text based on all the initial paragraphs, and obtaining the context semantic vector corresponding to all the initial paragraphs preceding the current initial paragraph; replacing at least some elements in the target triplet with the corresponding element categories to obtain the paragraph semantic vector of the current initial paragraph after replacement; obtaining the first triplet vector corresponding to all the target triplets, and using the mean of all the first triplet vectors as the target triplet vector; and using a decoding network to decode and obtain the target paragraph based on the full-text semantic vector, the context semantic vector, the paragraph semantic vector, and the target triplet vector.

8. An electronic device, characterized in that, The method includes a memory and a processor coupled to each other, wherein the memory stores program instructions and the processor executes the program instructions to implement the text correction method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The device stores program instructions that can be executed by a processor, the program instructions being used to implement the text correction method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Information error correction method and device based on knowledge graph

    CN113971217A