Text rectification method and system

By constructing a multimodal text deviation correction model, combining visual data and time difference learning algorithms, the problems of inaccurate model training and insufficient adaptability in the existing technology are solved, and efficient and accurate text deviation correction are achieved.

CN120045709BActive Publication Date: 2025-07-08兵器装备集团财务有限责任公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510143618.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-07-08
Estimated Expiration
2045-02-10

AI Technical Summary

Technical Problem

现有的文本纠偏方法在处理多模态数据时,模型训练不够精准,无法根据不同文本类型自适应选择最佳纠偏模型,导致纠偏准确性和效率低下。

Method used

By collecting text and visual data from multiple data sources, a text deviation correction model is built, and a multimodal coding layer, multimodal fusion layer, analysis layer and optimization layer are used to select the optimal deviation correction strategy in combination with the time difference learning algorithm, and adjust the model parameters through error feedback to ensure efficient deviation correction for different types of text.

Benefits of technology

The text deviation correction model is achieved with high accuracy and adaptability, which can improve deviation correction efficiency in different fields and scenarios, ensuring the accuracy of key information and the reliability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045709B_ABST
    Figure CN120045709B_ABST
Patent Text Reader

Abstract

The present invention discloses a text rectification method and system. The method includes: constructing a text rectification model; training different text rectification models through different data sets, comparing the rectification results output by the text rectification models with the actual rectification data to obtain the error values of the rectification results, determining whether the error values exceed the preset error threshold, if so, determining the parameters to be adjusted based on the error values, and adjusting the parameters to be adjusted until the error values of the text rectification model do not exceed the preset error threshold to obtain different optimal text rectification models; calculating the text rectification model with the highest data type matching value for the text to be rectified, determining the text rectification model with the highest matching value as the best text rectification model, and rectifying the text to be rectified based on the best text rectification model, which helps to solve the problems in the prior art that multi-modal data cannot be effectively utilized, model training lacks pertinence, and the best rectification model cannot be adaptively selected according to the text type.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a text rectification method and system. Background Art

[0002] With the rapid development of digital information, text data has emerged in large quantities in various fields. However, these text data often have various biases, including spelling mistakes, grammar errors, and semantic inaccuracies. In many scenarios, text data is also associated with visual data, such as pictures with image annotations, videos with video captions, etc. Traditional text rectification methods often only focus on the text itself and ignore the information contained in the associated visual data. Multimodal data (i.e., data containing multiple types such as text and vision) provides the possibility for more accurate text rectification. Utilizing the visual information in multimodal data to assist text rectification can improve the accuracy and efficiency of rectification.

[0003] However, when existing methods process multimodal data for text rectification, there are often problems such as inaccurate model training and inability to adaptively select the best model according to different text types.

[0004] Therefore, there is an urgent need for a method that can comprehensively utilize multimodal data, accurately train, and select the best rectification model. Summary of the Invention

[0005] In view of this, the present invention proposes a text rectification method and system, which can solve the problems that existing text rectification methods are difficult to effectively utilize multimodal data, lack pertinence in model training, and cannot adaptively select the best rectification model according to text types, and achieve more accurate text rectification.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A text rectification method, comprising:

[0008] Collecting different categories of data from multiple data sources, the data including text data and visual data associated with the text data;

[0009] Obtaining the category identifier corresponding to each data, and dividing the data into different data sets based on the category identifier;

[0010] Constructing a text rectification model, the text rectification model including an input layer, a multimodal encoding layer, a multimodal fusion layer, an analysis layer, a rectification strategy generation layer, an optimization layer, and an output layer;

[0011] Receive text data and visual data associated with the text data through an input layer, encode the text data and the visual data separately through a multi-modal encoding layer, mark key text segments, convert the text data into a text feature vector and a text key segment vector, convert the visual data into a visual feature vector, fuse the visual feature vector, the text feature vector and the text key segment vector through a multi-modal fusion layer to obtain a fused feature vector, locate the text error type and text error position based on the fused feature vector through an analysis layer, generate a candidate set of correction strategies based on the text error type and text error position through a correction strategy generation layer, calculate the Q value corresponding to each correction strategy in the candidate set of correction strategies through an optimization layer, update the Q value based on the immediate reward and future rewards using the temporal difference learning algorithm, select the optimal correction strategy from the candidate set of correction strategies based on the updated Q value, execute the optimal correction strategy, and output a correction result through the output layer;

[0012] Train different text correction models with different data sets, compare the correction results output by the text correction models with actual correction data to obtain an error value of the correction result, determine whether the error value exceeds a preset error threshold, if it exceeds, determine parameters to be adjusted based on the error value, and adjust the parameters to be adjusted until the error value of the text correction model does not exceed the preset error threshold to obtain different optimal text correction models;

[0013] Obtain the text to be corrected, determine the data type of the text to be corrected, calculate the text correction model with the highest matching value for the data type of the text to be corrected, determine the text correction model with the highest matching value as the best text correction model, and correct the text to be corrected based on the best text correction model.

[0014] Based on the above technical solutions, the present invention can also be improved as follows:

[0015] Optionally, the marking of the key text segments includes:

[0016] Calculate the comprehensive weight of the key segment of the word through formula (1);

[0017] w keul =α×w pl +(1 - α)×w t_fieldl Formula (1);

[0018] In the formula, w keul is the comprehensive weight of the key segment of word l, α is the weight coefficient, w pl is the part-of-speech tagging weight of word l, w t_fieldi is the term frequency-inverse document frequency weight of word l;

[0019] Judge wkeyl Whether it is greater than the critical segment threshold. If so, the word l and the words adjacent to the word l with a certain relevance are jointly determined as a critical segment.

[0020] Optionally, the converting the visual data into a visual feature vector, and fusing the visual feature vector, the text feature vector, and the text critical segment vector through a multimodal fusion layer to obtain a fused feature vector includes:

[0021] Calculating the fused feature vector through formula (2);

[0022]

[0023] In the formula, is the fused feature vector, tanh is the hyperbolic tangent function, W is the weight matrix corresponding to the visual feature vector, is the visual feature vector, is the bias vector, ⊙ is the element-wise multiplication, is the text feature vector, V is the weight matrix corresponding to the critical segment feature vector, and H is the critical segment feature vector.

[0024] Optionally, the analyzing layer locates the text error type and the text error position based on the fused feature vector, including:

[0025] Calculating the text error type through formula (3);

[0026]

[0027] In the formula, Error Type is the text error type, is to find the k that maximizes the expression value among all possible k, k is the index of different error types, and K is the error type currently being calculated, is the weight vector of the error type, is the fused feature vector, b k is the bias term of the error type;

[0028] Calculating the text error position through formula (4);

[0029]

[0030] In the formula, Error Position is the text error position, is to find the s that maximizes the expression value among all possible s, s is the index of different error positions, is the weight vector of the error position, b s is the bias term of the error position, is the fused feature vector.

[0031] Optionally, comparing the rectification result output by the text rectification model with the actual rectification data to obtain an error value of the rectification result, including:

[0032] Calculating the error value through formula (5);

[0033]

[0034] In the formula, Error is the error value, α is the weight coefficient, p is the number of samples, is the rectification result of the text rectification model for the i-th sample, y i is the actual rectification data of the i-th sample.

[0035] Optionally, comparing the rectification result output by the text rectification model with the actual rectification data to obtain an error value of the rectification result, determining whether the error value exceeds a preset error threshold, and if it exceeds, determining the parameter to be adjusted based on the error value and adjusting the parameter to be adjusted until the error value of the text rectification model does not exceed the preset error threshold to obtain different optimal text rectification models, including:

[0036] Dividing the parameters of the text rectification model into a parameter group including coding layer parameters, fusion layer parameters, analysis layer decision threshold parameters, rectification strategy generation layer rule parameters, and optimization layer learning rate parameters;

[0037] Conducting a small-range perturbation experiment on each parameter group one by one, determining the sensitive parameter sensitive to the error value based on the error value between the rectification result output by the text rectification model and the actual rectification data after each perturbation, and determining the sensitive parameter as the parameter to be adjusted;

[0038] Calculating the new parameter value corresponding to the parameter to be adjusted, adjusting the parameter to be adjusted based on the new parameter value, obtaining the new rectification result output by the adjusted text rectification model, and calculating the new error value based on the rectification result;

[0039] Comparing the new error value with the preset error threshold, and if the new error value does not exceed the preset error threshold after continuous multiple iterations, terminating the parameter adjustment to obtain the optimal text rectification model.

[0040] Optionally, the calculating the new parameter value corresponding to the parameter to be adjusted includes:

[0041] Calculating the new parameter value through formula (6);

[0042]

[0043] In the formula, θ new is the new parameter value, θ old is the old parameter value, η is the learning rate, p is the number of samples, For the rectification result The partial derivative of the parameter θ is the rectification result of the model for the i-th sample, y i is the actual rectification data of the i-th sample.

[0044] A text rectification system, comprising:

[0045] A data acquisition module, configured to acquire different types of data from multiple data sources, where the data includes text data and visual data associated with the text data;

[0046] A data partitioning module, configured to obtain the class identifier corresponding to each data, and partition the data into different data sets based on the class identifier;

[0047] A model construction module, configured to construct a text rectification model, where the text rectification model includes an input layer, a multi-modal encoding layer, a multi-modal fusion layer, an analysis layer, a rectification strategy generation layer, an optimization layer, and an output layer; receive text data and visual data associated with the text data through the input layer, encode the text data and the visual data respectively through the multi-modal encoding layer, mark the key text segments, convert the text data into a text feature vector and a text key segment vector, convert the visual data into a visual feature vector, fuse the visual feature vector, the text feature vector, and the text key segment vector through the multi-modal fusion layer to obtain a fused feature vector, locate the text error type and text error location based on the fused feature vector through the analysis layer, generate a candidate set of rectification strategies based on the text error type and text error location through the rectification strategy generation layer, calculate the Q value corresponding to each rectification strategy in the candidate set of rectification strategies through the optimization layer, update the Q value based on the immediate reward and future rewards based on the temporal difference learning algorithm, select the optimal rectification strategy from the candidate set of rectification strategies based on the updated Q value, execute the optimal rectification strategy, and output the rectification result through the output layer;

[0048] A model optimization module, configured to train different text rectification models through different data sets, compare the rectification result output by the text rectification model with the actual rectification data to obtain an error value of the rectification result, determine whether the error value exceeds a preset error threshold, if it exceeds, determine the parameter to be adjusted based on the error value, and adjust the parameter to be adjusted until the error value of the text rectification model does not exceed the preset error threshold, so as to obtain different optimal text rectification models;

[0049] A text rectification execution module is used to obtain the text to be rectified, determine the data type of the text to be rectified, calculate the text rectification model with the highest matching value for the data type of the text to be rectified, determine the text rectification model with the highest matching value as the best text rectification model, and rectify the text to be rectified based on the best text rectification model.

[0050] An electronic device includes a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the computer program, the steps of the method are implemented.

[0051] A non-transitory computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method are implemented.

[0052] The present invention has the following advantages:

[0053] In the text rectification method of the present invention, text data and visual data are collected from multiple data sources, greatly expanding the richness and diversity of the data. The data set is subdivided based on the category identifier corresponding to each data, and different text rectification models are trained through different data sets, enabling each text rectification model to adapt to specific types of data, improving the rectification accuracy of the text rectification model for specific types of data, not only avoiding interference between different types of data, but also avoiding the limitations of a single model when processing diverse data.

[0054] In the text rectification method of the present invention, the key content of the text is focused by marking key segments. For example, in a contract text, key information such as amounts, dates, and liability clauses is highlighted. When rectifying, the text rectification model will give priority to these key points that cannot afford to make mistakes, avoiding missing major errors due to the interference of secondary modifiers, greatly improving the pertinence of rectification and the accuracy of key information. The text data is further refined into text feature vectors and text key segment vectors, enabling the text rectification model to understand the text more stereoscopically and deeply. The text feature vectors can capture the overall grammar and semantic style of the text, while the text key segment vectors focus on key details. The combination of the two enables the text rectification model to not only have an overall view of the text and grasp the general meaning of the text, but also be able to insight into the details and accurately locate key errors.

[0055] In the text rectification method of the present invention, the text data and visual data are respectively encoded into text feature vectors and visual feature vectors, which can fully explore and utilize the information in multi-modal data. Different types of feature vectors are fused to obtain fused feature vectors, enabling the text rectification model to comprehensively consider text and visual information and improve the ability to locate text errors; based on the fused feature vectors, the types and positions of text errors are located to achieve accurate error analysis.

[0056] The text correction method in the present invention can select the optimal solution from multiple candidate strategies, improve the effectiveness of correction, and the output layer outputs the correction result to achieve the final presentation of the entire correction process.

[0057] The text correction method in the present invention obtains an error value by comparing the correction result with the actual correction data, and adjusts the model parameters according to the error value. This adjustment mechanism based on error feedback can continuously optimize the model, making it gradually converge to the optimal state to ensure the accuracy and reliability of the model.

[0058] The text correction method in the present invention selects the text correction model with the highest matching value according to the data type of the text to be corrected, which can ensure the most suitable correction method for different types of texts. Greatly improve the accuracy and efficiency of correction, enabling the model to better adapt to the text correction requirements in different fields and scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] For purposes of illustration and not limitation, the present invention will now be described in connection with the embodiments and drawings of the present invention, wherein:

[0060] Figure 1 is a schematic flowchart of the text correction method in the embodiment of the present invention;

[0061] Figure 2 is a schematic diagram of the main components of the text correction system in the embodiment of the present invention;

[0062] Figure 3 is a schematic diagram of the physical structure of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0063] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0064] It should be noted that in the description of the present invention and the above-mentioned drawings, terms such as "first" and "second" are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so as to implement the embodiments of the present invention described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0065] It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other. The embodiments of the present invention will be described in detail below with reference to the drawings.

[0066] Figure 1 It is a schematic flowchart of the text correction method in the embodiments of the present invention, as Figure 1 shown, the text correction method provided by the embodiments of the present invention includes the following steps S101 to S105.

[0067] S101, collect different types of data from multiple data sources, where the data includes text data and visual data associated with the text data.

[0068] Specifically, the data sources widely cover multiple fields such as Internet web pages, academic databases, social media platforms, e-books, news information websites, etc. A large amount of user comments and blog articles are crawled from Internet web pages. These texts have diverse styles and wide-ranging topics, covering many aspects such as daily trivia, discussions on the frontiers of technology, and sharing of entertainment news. At the same time, the associated web page pictures are crawled, such as product display pictures, landscape photos, and portraits, to endow the text with an intuitive visual context; by delving into academic databases, professional papers and research reports are extracted. Their rigorous terms and complex sentence structures are paired with experimental charts and data visualization graphics to help the model understand the profound academic context. On social media platforms, short copywriting with strong timeliness, caption for emoticons, as well as life photos and short video clips uploaded by users are collected to capture the current popular language trends and emotional colors; e-books provide rich literary genre materials, the delicate descriptions in novels and the rhythmic expressions in poems, combined with book covers and illustrations to build an artistic atmosphere; the current affairs reports and special news on news information websites, paired with on-site pictures and screenshots of news videos, make the text closely connected with reality. Such diverse data collection lays a solid foundation for the subsequent accurate and powerful text correction model training.

[0069] S102, obtain the category identifier corresponding to each data, and divide the data into different data sets based on the category identifier.

[0070] Specifically, for each piece of collected data, its category label is determined through fine-grained text analysis and visual feature recognition techniques. For text content, natural language processing algorithms are used to identify keywords, sentence structures, grammatical features, semantic categories, etc. For example, text containing a large number of professional terms and with a standard sentence pattern is determined to be academic, while text filled with internet buzzwords and emojis is classified as social media; from the perspective of visual data, image recognition technology is used to judge the image type. For example, product images correspond to text for commercial product introductions, landscape images are associated with travel and geography-related copywriting, and portrait photos are matched with biographies and social dynamics-related text. Once the category label is determined, the data is systematically divided into different sets such as academic literature datasets, social media datasets, literary work datasets, commercial copywriting datasets, news report datasets, etc., ensuring that data of the same type is gathered together, facilitating targeted training of the model, avoiding interference from data of different styles, and enhancing the model's ability to identify and correct errors in texts in various fields.

[0071] S103, construct a text correction model, which includes an input layer, a multi-modal encoding layer, a multi-modal fusion layer, an analysis layer, a correction strategy generation layer, an optimization layer, and an output layer.

[0072] Specifically, the input layer receives text data and visual data associated with the text data. The multi-modal encoding layer encodes the text data and visual data separately, marks the key segments of the text, converts the text data into text feature vectors and text key segment vectors, and converts the visual data into visual feature vectors. The multi-modal fusion layer fuses the visual feature vectors, text feature vectors, and text key segment vectors to obtain a fused feature vector. The analysis layer locates the text error type and text error location based on the fused feature vector. The correction strategy generation layer generates a candidate set of correction strategies based on the text error type and text error location. The optimization layer calculates the Q values corresponding to each correction strategy in the candidate set of correction strategies, updates the Q values based on the temporal difference learning algorithm using immediate rewards and future rewards, selects the optimal correction strategy from the candidate set of correction strategies based on the updated Q values, executes the optimal correction strategy, and the output layer outputs the correction result.

[0073] The marking of the key segments of the text includes:

[0074] Calculate the comprehensive weight of the key segments of the word through formula (1);

[0075] w keyl =α×w pl +(1-α)×w t_fieldl Formula (1);

[0076] In the formula, w keylThe comprehensive weight of the key segment for word l, α is the weight coefficient, and its value range is between 0 and 1, which determines the pl and w t_fieldi 's relative importance in the comprehensive weight, ω pl is the part-of-speech tagging weight of word l, which reflects the influence of the word's part of speech on the key segment weight, ω t_fieldi is the term frequency-inverse document frequency weight of word l, which reflects the importance of the word in the document and is usually used in information retrieval and text mining;

[0077] When α = 1, the formula becomes w keyl = w pl , meaning that the comprehensive weight of the key segment is completely determined by the part-of-speech tagging weight; when α = 0, the formula becomes w keyl = w t_fieldl , which means that the comprehensive weight of the key segment is completely determined by the term frequency-inverse document frequency weight.

[0078] For the case of 0 < α < 1, the comprehensive weight of the key segment is a weighted average of the part-of-speech tagging weight and the term frequency-inverse document frequency weight. Specifically:

[0079] The α×w pl part represents the contribution of the part-of-speech tagging weight to the comprehensive weight, and the (1 - α)×w t_fieldl part represents the contribution of the term frequency-inverse document frequency weight to the comprehensive weight.

[0080] Judge whether w keyl is greater than the key segment threshold. If so, jointly determine word l and the words adjacent to word l with a certain relevance as a key segment.

[0081] The conversion of visual data into visual feature vectors, and the fusion of visual feature vectors, text feature vectors, and text key segment vectors through a multi-modal fusion layer to obtain a fusion feature vector, including:

[0082] Calculate the fusion feature vector through formula (2);

[0083]

[0084] In the formula, is the fusion feature vector, tanh is the hyperbolic tangent function, W is the weight matrix corresponding to the visual feature vector, is the visual feature vector, is the bias vector, ⊙ is element-wise multiplication, is the text feature vector, V is the weight matrix corresponding to the key segment feature vector, and H is the key segment feature vector.

[0085] is for the visual feature vector Perform a linear transformation (using the weight matrix W), and then add the bias vector

[0086] Apply the result to the hyperbolic tangent function tanh, which compresses the input value into the interval (-1, 1) and serves as a non-linear activation

[0087] Then multiply element-wise with and VH (denoted by ⊙) to finally obtain the fused feature vector. The fused feature vector combines visual features, text features, and key segment features, and has been comprehensively processed through the weight matrix and non-linear activation function

[0088] The analysis layer locates the text error type and text error position based on the fused feature vector, including:

[0089] Calculate the text error type through formula (3);

[0090]

[0091] where Error Type is the text error type, is to find the k that maximizes the expression value among all possible k. k is the index of different error types, and K is the error type currently being calculated, is the weight vector of the error type, is the fused feature vector, and b k is the bias term of the error type;

[0092] is the sum of similar exponential terms for all k error types. The entire formula calculates a score for each error type k, which is obtained by dividing the numerator by the denominator. Then, through operation, find the k value that maximizes this score. This k value is the final text error type. Determine the text error type by calculating the scores of different error types and selecting the type with the highest score

[0093] Calculate the text error position through formula (4);

[0094]

[0095] where Error Position is the text error position, is to find the s that maximizes the expression value among all possible s. s is the index of different error positions, is the weight vector of the error position, and b sis the bias term for the error position, is the fused feature vector.

[0096] For each possible error position index s, calculate the expression value.

[0097] Here is the weight vector and the dot product of the fused feature vector , then add the bias term b s .

[0098] By comparing the values corresponding to all s, find the s value that maximizes the value of this expression.

[0099] This maximum s value is the text error position Error Position calculated by the formula.

[0100] An example of calculating the Q value corresponding to each correction strategy in the correction strategy candidate set through the optimization layer, updating the Q value based on immediate reward and future reward based on the temporal difference learning algorithm, and selecting the optimal correction strategy from the correction strategy candidate set based on the updated Q value is as follows:

[0101] Suppose we have a text to be corrected: "I are going to the park tomorrow". After being judged by the analysis layer, its error type is a grammar error (subject-verb disagreement), and the error position is at "are".

[0102] Based on this error type and position, generate a correction strategy candidate set:

[0103] Strategy 1: Replace "are" with "am";

[0104] Strategy 2: Insert "will" before "are" to become "I will are going to the parktomorrow";

[0105] Strategy 3: Delete "are" and change it to "I going to the park tomorrow".

[0106] Optimization layer: Calculate the Q value corresponding to each strategy. Initially, these Q values can be initialized based on some preset rules or randomly.

[0107] Suppose currently: The Q value of Strategy 1 is 0.4, the Q value of Strategy 2 is 0.2, and the Q value of Strategy 3 is 0.3.

[0108] Update the Q-value based on the temporal difference learning algorithm. Assume that after implementing Policy 1, the immediate reward is obtained: - The immediate reward is measured based on the similarity between the corrected text and the standard text ("I am going to the park tomorrow."). Here, assume the similarity is 0.8 (calculated by a certain text similarity algorithm, such as edit distance, etc.). The prediction of the future reward can be based on the model's past experience. For example, after correcting similar grammar errors in the past, the subsequent improvement in text coherence and readability and other comprehensive benefits. Assume the predicted future reward is 0.1.

[0109] According to the temporal difference learning algorithm, the Q-value update formula is approximately

[0110]

[0111] Substitute the values to calculate the Q-value after updating Policy 1:

[0112] The Q-value of Policy 1 becomes 0.476, the Q-value of Policy 2 remains 0.2 (not executed, not updated yet), and the Q-value of Policy 3 remains 0.3 (not executed, not updated yet).

[0113] Obviously, at this time, the Q-value of Policy 1 is the highest. Select Policy 1 as the optimal correction policy, and execute the optimal correction policy and output the result: Replace "are" with "am" to get the corrected text: "I am going to the parktomorrow.", and output this correction result through the output layer.

[0114] Compare the correction result output by the text correction model with the actual correction data to obtain the error value of the correction result, including:

[0115] Calculate the error value through formula (5);

[0116]

[0117] In the formula, Error is the error value, α is the weight coefficient, p is the number of samples, is the correction result of the text correction model for the i-th sample, y i is the actual correction data of the i-th sample.

[0118] What is calculated is the mean square error. For each sample p, first calculate the difference between the actual correction data y i and the model correction result , and then square the difference. Sum the squared differences of all samples, divide by the number of samples to get the average square error, and finally multiply by the weight coefficient;

[0119] The average absolute error is calculated. For each sample, the absolute value of the difference between the actual rectification data and the rectification result of the model is calculated first. The sum of the absolute differences of all samples is calculated, then divided by the number of samples to obtain the average absolute error, and finally multiplied by the weight coefficient. Finally, the two parts are added together to obtain the total error value.

[0120] S104. Different text rectification models are trained with different data sets. The rectification result output by the text rectification model is compared with the actual rectification data to obtain the error value of the rectification result. It is judged whether the error value exceeds the preset error threshold. If it exceeds, the parameter to be adjusted is determined based on the error value, and the parameter to be adjusted is adjusted until the error value of the text rectification model does not exceed the preset error threshold, so as to obtain different optimal text rectification models.

[0121] Specifically, the parameters of the text rectification model are divided into parameter groups including encoding layer parameters, fusion layer parameters, analysis layer decision threshold parameters, rectification strategy generation layer rule parameters, and optimization layer learning rate parameters;

[0122] The encoding layer parameters include weight parameters and bias parameters;

[0123] Weight parameters: The weights for converting text data into text feature vectors and text key segment vectors and the weights for converting visual data into visual feature vectors;

[0124] Bias parameters: The bias terms in the text data encoding process and the bias terms in the visual data encoding process.

[0125] The fusion layer parameters include fusion weight parameters and fusion bias parameters;

[0126] Fusion weight parameters: The weights for fusing visual feature vectors, text feature vectors, and text key segment vectors.

[0127] Fusion bias parameters: The bias terms in the fusion process.

[0128] The analysis layer decision threshold parameters include error type judgment threshold and error position judgment threshold;

[0129] Error type judgment threshold: The threshold for judging the text error type based on the fused feature vector; Error position judgment threshold: The threshold for judging the text error position based on the fused feature vector.

[0130] The rectification strategy generation layer rule parameters include strategy generation weights and strategy generation biases;

[0131] Strategy generation weights: The weight parameters used when generating the rectification strategy candidate set.

[0132] Strategy generation biases: The bias parameters used when generating the rectification strategy candidate set.

[0133] The optimized layer learning rate parameters include the learning rate parameter and the parameters related to the temporal difference learning algorithm;

[0134] Learning rate parameter: The learning rate used by the optimized layer when calculating the Q-values corresponding to each rectification strategy in the rectification strategy candidate set.

[0135] The parameters related to the temporal difference learning algorithm include the immediate reward and the parameters involved in updating the Q-value for future rewards.

[0136] Perform small-range perturbation experiments on each parameter group one by one. Based on the error value between the rectification result output by the text rectification model after each perturbation and the actual rectification data, determine the sensitive parameters sensitive to the error value, and determine the sensitive parameters as the parameters to be adjusted;

[0137] Calculate the new parameter values corresponding to the parameters to be adjusted. Based on the new parameter values, adjust the parameters to be adjusted, obtain the new rectification result output by the adjusted text rectification model, and calculate the new error value based on the rectification result;

[0138] Compare the new error value with the preset error threshold. If the new error value does not exceed the preset error threshold after continuous iterations for multiple times, terminate the parameter adjustment to obtain the optimal text rectification model.

[0139] The calculation of the new parameter values corresponding to the parameters to be adjusted includes:

[0140] Calculate the new parameter values through formula (6);

[0141]

[0142] In the formula, θ new is the new parameter value, θ old is the old parameter value, η is the learning rate, p is the number of samples, is the rectification result the partial derivative of the parameter θ, is the rectification result of the model for the i-th sample, y i is the actual rectification data of the i-th sample.

[0143] An example:

[0144] Suppose there is a text rectification model. The current preset error threshold is 0.05, and the model is processing a small dataset containing 100 text samples.

[0145] Parameter grouping: Encoding layer parameters: Suppose the text encoding weight matrix W text The initial value is a randomly generated 50×100 matrix, and the bias vector b text The initial value is a zero vector of length 50.

[0146] Visual coding weight matrix W υision The initial value is a randomly generated 30×80 matrix, and the bias vector b υision The initial value is a zero vector with a length of 30.

[0147] Fusion weight matrix W fusion The initial value is a randomly generated 80×130 matrix, and the bias vector b fusion The initial value is a zero vector with a length of 80.

[0148] Error type judgment threshold T type The initial value is 0.7;

[0149] Error position judgment threshold T position The initial value is 0.6.

[0150] Policy generation weight matrix W strategy The initial value is a randomly generated 20×80 matrix, and the bias vector b strategy The initial value is a zero vector with a length of 20.

[0151] The initial value of the learning rate η is 0.01.

[0152] Text encoding weight matrix W text Perform a small range of perturbations, for example, increase or decrease each element by a small random value (ranging from -0.05 to 0.05).

[0153] Calculate the error value between the rectification result of the perturbed model and the actual rectification data. Assume that the error value increases from 0.06 to 0.08 after the first perturbation, indicating that this parameter is relatively sensitive to the error.

[0154] For the visual coding weight matrix W υision Perform a similar operation to evaluate its impact on the error. For the fusion weight matrix W fusion Perform a perturbation and calculate the error value. Assume that the error value increases from 0.06 to 0.09 after the perturbation, and determine it as a sensitive parameter.

[0155] Change the error type judgment threshold T type For example, increase or decrease it by 0.1. Assume that the error value decreases from 0.06 to 0.04 after the increase, indicating that this parameter has a significant impact on the error.

[0156] For the policy generation weight matrix W strategy Perform a perturbation and observe the change in the error. Assume that the error value increases from 0.06 to 0.07 after the perturbation, and judge the degree of its impact on the error.

[0157] Change the learning rate η, for example, increase or decrease it by 0.005. Assume that after the increase, the error value increases from 0.06 to 0.07, and evaluate its impact on the error.

[0158] Through the above perturbation experiments, it is found that the fusion weight matrix W fusion and the error type judgment threshold T type have a greater impact on the error value, and they are determined as the parameters to be adjusted.

[0159] Use a certain optimization algorithm (such as gradient descent) to calculate the new parameter values of the fusion weight matrix W fusion Assume that after one iteration of calculation, the new value of W fusion is updated.

[0160] According to the performance of the model and the error analysis, adjust the error type judgment threshold T type from 0.7 to 0.65.

[0161] Use the adjusted parameters to re-run the text correction model, and correct 100 text samples. Calculate the error value between the new correction result and the actual correction data. Assume that the new error value is 0.04. Since the new error value 0.04 is less than the preset error threshold 0.05, terminate the parameter adjustment. The model obtained at this time is the optimized text correction model (under the current data set and evaluation criteria) after adjustment.

[0162] S105. Obtain the text to be corrected, determine the data type of the text to be corrected, calculate the text correction model with the highest matching value for the data type of the text to be corrected, determine the text correction model with the highest matching value as the best text correction model, and correct the text to be corrected based on the best text correction model.

[0163] Specifically, extract the features of the text to be corrected, obtain the feature vector of the text to be corrected, regard the feature vector of the text to be corrected and the feature vectors of each text correction model as points in space, calculate the Euclidean distance between them, the smaller the distance, the higher the matching value. Calculate the cosine similarity between the feature vector of the text to be corrected and the feature vector of the model, the higher the similarity, the higher the matching value.

[0164] Figure 2 It is a schematic diagram of the main components of the text correction system in the embodiment of the present invention. As Figure 2 shown, the text correction system 1 provided by the embodiment of the present invention includes a data acquisition module 10, a data division module 20, a model construction module 30, a model optimization module 40, and a text correction execution module 50.

[0165] The data acquisition module 10 is used to collect different types of data from multiple data sources, and the data includes text data and visual data associated with the text data;

[0166] The data partitioning module 20 is used to obtain the category identifier corresponding to each piece of data, and partition the data into different data sets based on the category identifier;

[0167] The model construction module 30 is used to construct a text correction model, which includes an input layer, a multi-modal encoding layer, a multi-modal fusion layer, an analysis layer, a correction strategy generation layer, an optimization layer, and an output layer; receive text data and visual data associated with the text data through the input layer, encode the text data and the visual data respectively through the multi-modal encoding layer, mark the key text segments, convert the text data into a text feature vector and a text key segment vector, convert the visual data into a visual feature vector, fuse the visual feature vector, the text feature vector, and the text key segment vector through the multi-modal fusion layer to obtain a fused feature vector, locate the text error type and text error location based on the fused feature vector through the analysis layer, generate a candidate set of correction strategies based on the text error type and text error location through the correction strategy generation layer, calculate the Q value corresponding to each correction strategy in the candidate set of correction strategies through the optimization layer, update the Q value based on the immediate reward and future rewards based on the temporal difference learning algorithm, select the optimal correction strategy from the candidate set of correction strategies based on the updated Q value, execute the optimal correction strategy, and output the correction result through the output layer;

[0168] The model optimization module 40 is used to train different text correction models through different data sets, compare the correction results output by the text correction models with the actual correction data to obtain the error value of the correction results, determine whether the error value exceeds a preset error threshold, if it exceeds, determine the parameters to be adjusted based on the error value, and adjust the parameters to be adjusted until the error value of the text correction model does not exceed the preset error threshold, so as to obtain different optimal text correction models;

[0169] The text correction execution module 50 is used to obtain the text to be corrected, determine the data type of the text to be corrected, calculate the text correction model with the highest matching value for the data type of the text to be corrected, determine the text correction model with the highest matching value as the best text correction model, and correct the text to be corrected based on the best text correction model.

[0170] Figure 3 It is a schematic diagram of the physical structure of the electronic device provided by the embodiment of the present invention. As Figure 3 shown, the electronic device 60 includes: a processor 601 (processor), a memory 602 (memory), and a bus 603;

[0171] Among them, the processor 601 and the memory 602 complete communication with each other through the bus 603;

[0172] The processor 601 is configured to call program instructions in the memory 602 to execute the methods provided in the foregoing method embodiments, so as to execute the methods provided in the embodiments of the present invention.

[0173] This embodiment provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the methods provided in the embodiments of the present invention.

[0174] Those of ordinary skill in the art can understand that all or part of the steps of implementing the foregoing method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the foregoing method embodiments; and the foregoing storage medium includes various storage media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0175] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A text rectification method, characterized in that, Including: Collecting different types of data from multiple data sources, where the data includes text data and visual data associated with the text data; Obtaining the category identifier corresponding to each data, and dividing the data into different data sets based on the category identifier; Constructing a text correction model, where the text correction model includes an input layer, a multi-modal encoding layer, a multi-modal fusion layer, an analysis layer, a correction strategy generation layer, an optimization layer, and an output layer; Receiving the text data and the visual data associated with the text data through the input layer, encoding the text data and the visual data respectively through the multi-modal encoding layer, marking the key text segments, converting the text data into a text feature vector and a text key segment vector, converting the visual data into a visual feature vector, fusing the visual feature vector, the text feature vector, and the text key segment vector through the multi-modal fusion layer to obtain a fused feature vector, locating the text error type and the text error position based on the fused feature vector through the analysis layer, generating a candidate set of correction strategies based on the text error type and the text error position through the correction strategy generation layer, calculating the Q value corresponding to each correction strategy in the candidate set of correction strategies through the optimization layer, updating the Q value based on the temporal difference learning algorithm and the immediate reward and the future reward, selecting the optimal correction strategy from the candidate set of correction strategies based on the updated Q value, executing the optimal correction strategy, and outputting the correction result through the output layer; Training different text correction models with different data sets, comparing the correction result output by the text correction model with the actual correction data to obtain the error value of the correction result, determining whether the error value exceeds the preset error threshold, if it exceeds, determining the parameter to be adjusted based on the error value, and adjusting the parameter to be adjusted until the error value of the text correction model does not exceed the preset error threshold to obtain different optimal text correction models; Obtaining the text to be corrected, determining the data type of the text to be corrected, calculating the text correction model with the highest matching value for the data type of the text to be corrected, determining the text correction model with the highest matching value as the best text correction model, and correcting the text to be corrected based on the best text correction model.

2. The text rectification method according to claim 1, wherein The marking of the key text segments includes: Calculating the comprehensive weight of the key segments of the word through formula (1); Formula (1); In the formula, is the comprehensive weight of the key segment of the word , is the weight coefficient, is the part-of-speech tagging weight of the word , is the word 's term frequency-inverse document frequency weight; Determine whether it is greater than the key segment threshold. If so, determine the word and the word with a certain correlation adjacent to the word together as a key segment.

3. The text rectification method according to claim 1, characterized in that, The conversion of the visual data into a visual feature vector and the fusion of the visual feature vector, the text feature vector, and the text key segment vector through the multi-modal fusion layer to obtain a fused feature vector includes: Calculating the fused feature vector through formula (2); Formula (2); In the formula, is the fused feature vector, is the hyperbolic tangent function, is the weight matrix corresponding to the visual feature vector, is the visual feature vector, is the bias vector, is the element-wise multiplication, is the text feature vector, is the weight matrix corresponding to the key segment feature vector, is the key segment feature vector.

4. The text rectification method according to claim 1, characterized in that The positioning of the text error type and the text error position based on the fused feature vector through the analysis layer includes: Calculating the text error type through formula (3); Formula (3); Wherein, is the text error type, is to find the that maximizes the expression value among all possible , is the index of different error types, is the error type currently being calculated, is the weight vector of the error type, is the fused feature vector, is the bias term of the error type; Calculating the text error position through formula (4); Formula (4); In the formula, is the text error position, is to find the one that maximizes the expression value among all possible , , is the index of different error positions, is the weight vector of the error position, is the bias term of the error position, is the fused feature vector.

5. The text rectification method according to claim 1, characterized in that The comparison of the correction result output by the text correction model with the actual correction data to obtain the error value of the correction result includes: Calculating the error value through formula (5); Formula (5); Wherein, is the error value, is the weight coefficient, is the number of samples, is the correction result of the text correction model for the th sample, is the actual correction data of the th sample.

6. The text rectification method according to claim 1, characterized in that Compare the rectification result output by the text rectification model with the actual rectification data to obtain the error value of the rectification result, and determine whether the error value exceeds the preset error threshold. If it exceeds, determine the parameter to be adjusted based on the error value, and adjust the parameter to be adjusted until the error value of the text rectification model does not exceed the preset error threshold, so as to obtain different optimal text rectification models, including: Divide the parameters of the text rectification model into a parameter group including encoding layer parameters, fusion layer parameters, analysis layer decision threshold parameters, rectification strategy generation layer rule parameters, and optimization layer learning rate parameters; Conduct a small-range perturbation experiment on each parameter group one by one. Based on the error value between the rectification result output by the text rectification model after each perturbation and the actual rectification data, determine the sensitive parameters sensitive to the error value, and determine the sensitive parameters as the parameters to be adjusted; Calculate the new parameter value corresponding to the parameter to be adjusted, adjust the parameter to be adjusted based on the new parameter value, obtain the new rectification result output by the adjusted text rectification model, and calculate the new error value based on the rectification result; Compare the new error value with the preset error threshold. If the new error value does not exceed the preset error threshold after continuous multiple iterations, terminate the parameter adjustment to obtain the optimal text rectification model.

7. The text rectification method according to claim 6, characterized in that, The calculating the new parameter value corresponding to the parameter to be adjusted includes: Calculating the new parameter value through formula (6); Formula (6); Wherein, is the new parameter value, is the old parameter value, is the learning rate, is the number of samples, is the rectification result is the partial derivative of the parameter , is the rectification result of the model for the -th sample, is the -th actual rectification data of the sample.

8. A text rectification system, characterized in that, Including: A data acquisition module for collecting different types of data from multiple data sources, where the data includes text data and visual data associated with the text data; A data division module for obtaining the category identifier corresponding to each data, and dividing the data into different data sets based on the category identifier; A model construction module for constructing a text rectification model, where the text rectification model includes an input layer, a multi-modal encoding layer, a multi-modal fusion layer, an analysis layer, a rectification strategy generation layer, an optimization layer, and an output layer; receive text data and visual data associated with the text data through the input layer, encode the text data and the visual data respectively through the multi-modal encoding layer, mark the key text segments, convert the text data into a text feature vector and a text key segment vector, convert the visual data into a visual feature vector, fuse the visual feature vector, the text feature vector, and the text key segment vector through the multi-modal fusion layer to obtain a fused feature vector, locate the text error type and text error position based on the fused feature vector through the analysis layer, generate a candidate set of rectification strategies based on the text error type and text error position through the rectification strategy generation layer, calculate the Q value corresponding to each rectification strategy in the candidate set of rectification strategies through the optimization layer, update the Q value based on the temporal difference learning algorithm and immediate reward and future reward, select the optimal rectification strategy from the candidate set of rectification strategies based on the updated Q value, execute the optimal rectification strategy, and output the rectification result through the output layer; A model optimization module, which is used to train different text correction models through different data sets, compare the correction results output by the text correction models with the actual correction data to obtain the error value of the correction results, determine whether the error value exceeds a preset error threshold, and if it exceeds, determine the parameters to be adjusted based on the error value, and adjust the parameters to be adjusted until the error value of the text correction model does not exceed the preset error threshold, so as to obtain different optimal text correction models; A text correction execution module, which is used to obtain the text to be corrected, determine the data type of the text to be corrected, calculate the text correction model with the highest matching value for the data type of the text to be corrected, determine the text correction model with the highest matching value as the best text correction model, and correct the text to be corrected based on the best text correction model.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A non-transitory computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Text error correction method and system, model training method, medium and equipment

    CN116681070A

  • Multi-modal text error correction method and system based on multiple inputs

    CN118036594A