A precise verification method and system for the text information of a packaging box based on OCR error processing
By introducing error processing-based packaging box text information verification method in OCR technology, using technical means such as text rearrangement, similarity matching and difference ratio, the problems of high error rate and low automation in the existing technology are solved, and accurate verification and efficient processing of packaging box text information is achieved.
Patent Information
- Application Number
- CN202410339260.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-03-25
AI Technical Summary
The existing OCR technology has problems such as high error rate, difficulty in processing special characters and formats, and low degree of automation in text processing. Especially in the case of packaging box text information detection scenarios, when the same content is different in layout, it is difficult for the existing technology to achieve accurate verification.
A precise verification method for packaging box text information based on OCR error processing is adopted. By importing the original text information file and reading the ignorance rules, OCR recognition, text rearrangement, Cartesian combination, similarity matching and best sentence matching are carried out, and the difference comparison and ignorance rules are combined to achieve effective processing of OCR errors.
It improves the accuracy of information verification, eliminates interference from OCR errors, ensures the accuracy and credibility of the final result, is applicable to different typesettings of the same content, and adapts to different typesettings of text information verification of different products and typesettings through flexible rules.
Smart Images

Figure CN118212641B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text information verification, and specifically relates to a precise verification method and system for packaging box text information based on OCR error processing. Background Art
[0002] In today's digital age, a large amount of text information needs to be processed, compared, and verified; optical character recognition technology is widely used in the field of text processing due to its high efficiency and convenience.
[0003] Currently, in the application of OCR technology, there are often some challenges and limitations, specifically manifested in high error rates, difficulties in processing special characters and formats, and low automation. Existing technologies have significant deficiencies in terms of the accuracy, fault tolerance, processing efficiency, and automation of OCR text processing. For the product packaging box information detection scenario, product description information usually needs to be typeset by designers, and the typesetting of each product is different. In the case of the same product with different packaging boxes, the same product information typesetting is adopted, that is, the content is the same but the typesetting is different; therefore, it is necessary to design a precise verification method and system for packaging box text information based on OCR error processing. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies of the prior art, and to better and effectively solve the problems that there are often some challenges and limitations in the application of OCR technology, specifically manifested in high error rates, difficulties in processing special characters and formats, and low automation, and existing technologies have significant deficiencies in terms of the accuracy, fault tolerance, processing efficiency, and automation of OCR text processing. A precise verification method and system for packaging box text information based on OCR error processing are provided, which realizes differential comparison of OCR errors and effective processing in combination with ignoring rules, improves the accuracy of information verification, excludes the interference of OCR errors, ensures the accuracy and reliability of the final result, and is not only applicable to the situation of the same content with different typesetting, but also can adapt to the text information verification of different products with different typesetting through flexible rules.
[0005] In order to achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0006] A precise verification method for packaging box text information based on OCR error processing includes the following steps:
[0007] Step A, import the original text information file and read the ignoring rules to obtain text fragments of the original text information file;
[0008] Step B, perform OCR recognition on the image area containing the original text information file to obtain the recognized text content;
[0009] Step C, rearrange the recognized text content to obtain the rearranged text content;
[0010] Step D, perform a Cartesian combination of the rearranged text content and the text fragments of the original text document of the text information, and use similarity matching and best sentence matching to obtain the matching text;
[0011] Step E, perform a difference comparison on the matching text and use the ignore rule to retrieve the comparison result to obtain the difference after retrieval;
[0012] Step F, mark and display the difference after retrieval to complete the verification operation of the text information of the packaging box.
[0013] The above-mentioned precise verification method for the text information of the packaging box based on OCR error processing, step A, import the original text document of the text information and read the ignore rule to obtain the text fragments of the original text document of the text information. The specific steps are as follows.
[0014] Step A1, open the original text document of the text information in read-only mode and set the encoding of the original text document of the text information, and then read the content in the original text document of the text information line by line;
[0015] Step A2, remove the blank characters at the head and tail of the string for each line in the original text document of the text information, and split each line into two parts with a colon as the delimiter, and then represent different characters and the corresponding ignore character list respectively;
[0016] Step A3, split the string into a character list with a comma as the delimiter, and check whether each character in the split character list is a null string. If the check result is yes, replace the null string with an empty string. If the check result is no, keep it unchanged;
[0017] Step A4, assign the processed character list to the corresponding key in the first dictionary.
[0018] The above-mentioned precise verification method for the text information of the packaging box based on OCR error processing, step B, perform OCR recognition on the image area containing the original text document of the text information to obtain the recognized text content. The specific steps are as follows.
[0019] Step B1, set the pixel value in the image area R i to be represented as p(x, y), where (x, y) is the coordinate of the pixel, and use the convolution expression to perform feature extraction on the image area R i and generate a feature vector F i As shown in formula (1).
[0020]
[0021] Among them, k(m,n) is the weight of the convolution kernel, and M and N are respectively half of the height and width of the convolution kernel;
[0022] Step B2, convert the feature vector F i into the text sequence T i , as shown in formula (2),
[0023] T i =σ(∑ j W j F i (j)+b)(2)
[0024] Among them, W j is the model weight, b is the bias term, j is the index of the feature vector F i , and σ is the softmax activation function. The softmax activation function is used to convert the result of weighted summation into a probability distribution and represent the possibility of each character.
[0025] The aforementioned accurate verification method for the text information of the packaging box based on OCR error processing, step C, rearrange the recognized text content to obtain the rearranged text content. The specific steps are as follows.
[0026] Step C1, initialize an empty list to store the matching results, then traverse each text paragraph recognized by OCR. Then, for each text paragraph, call a function to compare it with the paragraphs in the reference text paragraph set and find the paragraph with the highest similarity to it and its similarity. If the similarity is not 0, add this OCR text paragraph, the best matching paragraph, and the similarity between the OCR text paragraph and the best matching paragraph as an element to the empty list. The reference text paragraph set is the set of text fragments in the original information file.
[0027] Step C2, use list comprehension to filter out the matching items with similarity greater than or equal to the threshold from the empty list and store them in the first list. Then, filter out the matching items with low similarity based on the similarity threshold and only keep the matching results with high similarity.
[0028] Step C3, find the optimal similarity text pair matching function. The specific steps are as follows.
[0029] Step C31, use a function to check whether the second list containing multiple reference text paragraphs is not empty. If the second list is not empty, the matching function enters the comparison process.
[0030] Step C32, initialize the best matching paragraph as the first element in the reference text list and use it as the initial value of the current best matching item. Then, calculate the cosine similarity between the OCR text paragraph and the initial value and store the calculation result in the third list.
[0031] Step C33: Calculate the similarity of other text paragraphs in the second list using cosine similarity. Specifically, a matching function is used to receive two string parameters, text1 and text2, where the two string parameters text1 and text2 represent two paragraphs of text for comparing similarity. The specific steps are as follows:
[0032] Step C331: Convert the two paragraphs of text into term frequency vectors. Specifically, analyze all unique words in the text, and then generate a sparse matrix according to the occurrence frequency of each word in each paragraph of text. The sparse matrix X is shown in formula (3).
[0033]
[0034] Among them, each row represents the vectorized representation of a document, and each column corresponds to a word in the vocabulary V.
[0035] Step C332: Calculate the cosine similarity. Specifically, use the similarity function to calculate the cosine similarity between two term frequency vectors from the sparse matrix to obtain a similarity matrix, and then the similarity function returns the similarity matrix. The similarity matrix contains the similarity values between sparse matrices. The calculation of the cosine similarity is as described in formula (4), and the similarity matrix is shown in formula (5).
[0036]
[0037] Among them, cosine_similarity is the similarity function, ‖X‖ and ‖Y‖ are the Euclidean norms of X and Y respectively, n is the number of features, and K ij is the cosine similarity between the i-th sample in the sparse matrix X and the j-th sample in Y.
[0038] Step C333: Construct a vocabulary using the similarity matrix and convert the text into term frequency vectors, as shown in formula (6).
[0039] D = {d 1 , d 2}(6)
[0040] Among them, d 1 and d 2 are the string parameters text1 and text2 respectively, and D is the term frequency vector.
[0041] Step C334: Remove duplicates from the words in all documents and form a vocabulary V = {v 1 , v 2 , …, v n}, where n is the total number of words in the vocabulary. Then, for each document d iConstruct a vector x with the same length as the vocabulary V i ={x i1 , x i2 , …, x in}, where x ij represents the number of occurrences of the vocabulary v j in the document d i .
[0042] For the aforementioned precise verification method of the packaging box text information based on OCR error processing, step E: perform a differential comparison on the matching text and use the ignore rule to retrieve the comparison result to obtain the retrieved difference. The specific steps are as follows
[0043] Step E1: Use two-layer nested traversal to traverse all strings in the comparison result string. Among them, the outer loop traverses each starting index, and the inner loop traverses each ending index starting from the starting index + 1
[0044] Step E2: In each inner loop, slice the comparison result string according to the starting index and the ending index, and obtain the current substring
[0045] Step E3: Check whether the current substring exists in the second dictionary. The specific steps are as follows
[0046] Step E31: If the length of the current substring is 1 and it does not exist in the empty dictionary, then use the current substring as the key and the value as the fourth list, where the fourth list contains the initial index list and the corresponding second dictionary value
[0047] Step E32: If the length of the current substring is 1 and it already exists in the empty dictionary, then append the starting index to the index list of the existing key
[0048] Step E33: If the length of the current substring is greater than 1 and it does not exist in the empty dictionary, then use the current substring as the fifth list, where the fifth list contains tuples and the corresponding second dictionary value
[0049] Step E34: If the length of the current substring is greater than 1 and it already exists in the empty dictionary, then append the tuple to the index list of the existing key
[0050] For the aforementioned precise verification method of the packaging box text information based on OCR error processing, step F: mark and display the retrieved difference to complete the verification operation of the packaging box text information. Specifically, retrieve the corresponding sentence in the information original file at the error position. If such an error exists, cancel the highlighting here
[0051] An accurate verification system for the text information of the packaging box based on OCR error processing, including an import and reading module, an image recognition module, a text rearrangement module, a text matching module, a difference comparison module, and a difference display module. The import and reading module is used to import the original text information file and read the ignore rules to obtain the text fragments of the original text information file. The image recognition module is used to perform OCR recognition on the image area containing the original text information file to obtain the recognized text content. The text rearrangement module is used to rearrange the recognized text content to obtain the rearranged text content. The text matching module is used to perform Cartesian combination on the rearranged text content and the text fragments of the original text information file and use similarity matching and best sentence matching to obtain the matching text. The difference comparison module is used to perform difference comparison on the matching text and use the ignore rules to retrieve the comparison results to obtain the retrieved differences. The difference display module is used to mark and display the retrieved differences to complete the verification operation of the text information of the packaging box.
[0052] The beneficial effects of the present invention are as follows:
[0053] (1). The accuracy of information verification is improved. The present invention effectively processes the OCR errors through difference comparison and combines the ignore rules, improving the accuracy of information verification, excluding the interference of OCR errors, and ensuring that the final result is more accurate and reliable.
[0054] (2). The need for manual intervention is reduced. The present invention adopts an automated processing flow, reducing the dependence on manual intervention, improving the efficiency of information processing, reducing the workload of operators, and making information verification more convenient.
[0055] (3). The OCR error tolerance is improved. The present invention introduces flexible ignore rules, which can ignore specific OCR misrecognition situations, thereby improving the system's tolerance to errors and enabling the system to better adapt to diverse text inputs.
[0056] (4). It is applicable to various typesetting differences. The present invention is not only applicable to the situation of the same content with different typesettings, but also can adapt to the text information verification of different products and different typesettings through flexible rules, with strong versatility.
[0057] (5). Quickly locate and mark differences. The present invention quickly locates text differences through similarity matching and difference comparison, and enables operators to quickly pay attention to and handle specific problems through difference marking and display, improving work efficiency.
[0058] (6). An ever-optimizing fault tolerance mechanism. The present invention enables the ignore rules of the system to be gradually optimized as misrecognition is continuously discovered, thereby continuously improving the fault tolerance mechanism and further enhancing the robustness of the system. Brief Description of the Drawings
[0059] Figure 1 is a flowchart of a method and system for accurately verifying the text information of a packaging box based on OCR error processing according to the present invention. Detailed Embodiment
[0060] The present invention will be further described below in conjunction with the accompanying drawings of the specification.
[0061] As Figure 1 shown, a method for accurately verifying the text information of a packaging box based on OCR error processing according to the present invention includes the following steps
[0062] Step A: Import the original text information file and read the ignore rules to obtain the text fragments of the original text information file. The specific steps are as follows
[0063] Step A1: Open the original text information file in read-only mode and set the encoding of the original text information file, and then read the content in the original text information file line by line;
[0064] Step A2: Remove the blank characters at the beginning and end of the string for each line in the original text information file, and split each line into two parts with a colon as the delimiter, and then represent different characters and the corresponding ignore character list respectively;
[0065] Step A3: Split the string into a character list with a comma as the delimiter, and check whether each character in the split character list is a null string. If the check result is yes, replace the null string with an empty string. If the check result is no, keep it unchanged;
[0066] Step A4: Assign the processed character list to the corresponding key in the first dictionary.
[0067] Step B: Perform OCR recognition on the image area containing the original text information file to obtain the recognized text content. The specific steps are as follows
[0068] Step B1: Set the pixel value in the image area R i as p(x, y), where (x, y) is the coordinate of the pixel, and use the convolution expression to perform feature extraction on the image area R i and generate a feature vector F i As shown in formula (1),
[0069]
[0070] where k(m, n) is the weight of the convolution kernel, and M and N are respectively half of the height and width of the convolution kernel;
[0071] Step B2: The feature vector Fi Convert to text sequence T i , as shown in formula (2),
[0072] T i = σ(∑ j W j F i (j)+b)(2)
[0073] where W j is the model weight, b is the bias term, j is the index of the feature vector F i , and σ is the softmax activation function, and the softmax activation function is used to convert the result of weighted summation into a probability distribution and represent the possibility of each character.
[0074] Step C, rearrange the recognized text content to obtain the rearranged text content. The specific steps are as follows.
[0075] Step C1, initialize an empty list to store the matching results, then traverse each text paragraph recognized by OCR, and then for each text paragraph, call a function to compare it with the paragraphs in the reference text paragraph set and find the paragraph with the highest similarity and its similarity. If the similarity is not 0, add this OCR text paragraph, the best matching paragraph, and the similarity between the OCR text paragraph and the best matching paragraph as an element to the empty list, where the reference text paragraph set is the set of text fragments in the information original file;
[0076] Step C2, use list comprehension to filter out the matching items with similarity greater than or equal to the threshold from the empty list and store them in the first list, and then filter out the matching items with low similarity based on the similarity threshold and only keep the matching results with high similarity;
[0077] Step C3, find the optimal similarity text pair matching function. The specific steps are as follows.
[0078] Step C31, use a function to check whether the second list containing multiple reference text paragraphs is not empty. If the second list is not empty, the matching function enters the comparison process;
[0079] Step C32, initialize the best matching paragraph as the first element in the reference text list and use it as the initial value of the current best matching item, then calculate the cosine similarity between the OCR text paragraph and the initial value and store the calculation result in the third list;
[0080] Step C33. Calculate the similarity of other text paragraphs in the second list using cosine similarity. Specifically, a matching function is used to receive two string parameters text1 and text2, and the two string parameters text1 and text2 represent two paragraphs of text for comparing similarity. The specific steps are as follows:
[0081] Step C331. Convert the two paragraphs of text into term frequency vectors. Specifically, analyze all unique words in the text, and then generate a sparse matrix according to the occurrence frequency of each word in each paragraph of text. The sparse matrix X is shown in formula (3).
[0082]
[0083] Among them, each row represents the vectorized representation of a document, and each column corresponds to a word in the vocabulary V;
[0084] Step C332. Calculate the cosine similarity. Specifically, use the similarity function to calculate the cosine similarity between two term frequency vectors from the sparse matrix to obtain a similarity matrix, and then the similarity function returns the similarity matrix. The similarity matrix contains the similarity values between the sparse matrices. The calculation of the cosine similarity is as described in formula (4), and the similarity matrix is shown in formula (5).
[0085]
[0086] Among them, cosine_similarity is the similarity function, ‖X‖ and ‖Y‖ are the Euclidean norms of X and Y respectively, n is the number of features, and K ij is the cosine similarity between the i-th sample in the sparse matrix X and the j-th sample in Y;
[0087] Step C333. Construct a vocabulary using the similarity matrix and convert the text into term frequency vectors, as shown in formula (6).
[0088] D = {d 1 , d 2}(6)
[0089] Among them, d 1 and d 2 are the string parameters text1 and text2 respectively, and D is the term frequency vector;
[0090] Step C334. Remove duplicates from the words in all documents and form a vocabulary V = {v 1 , v 2 , …, v n}, where n is the total number of words in the vocabulary. Then, for each document d i construct a vector x with the same length as the vocabulary Vi = {x i1 , x i2 , …, x in}, where x ij represents the number of occurrences of the vocabulary v j in the document d i .
[0091] Step D: Perform a Cartesian combination of the rearranged text content and the text fragments of the original text information file, and use similarity matching and best sentence matching to obtain the matching text;
[0092] Step E: Compare the differences of the matching text and use the ignore rule to retrieve the comparison result to obtain the retrieved difference. The specific steps are as follows:
[0093] Step E1: Use two-layer nested traversal to iterate through all strings in the comparison result string. The outer loop iterates through each starting index, and the inner loop iterates through each ending index starting from the index immediately after the starting index;
[0094] Step E2: In each inner loop, slice the comparison result string according to the starting index and the ending index, and obtain the current substring;
[0095] Step E3: Check whether the current substring exists in the second dictionary. The specific steps are as follows:
[0096] Step E31: If the length of the current substring is 1 and it does not exist in the empty dictionary, use the current substring as the key and the value as the fourth list, where the fourth list contains the initial index list and the corresponding second dictionary value;
[0097] Step E32: If the length of the current substring is 1 and it already exists in the empty dictionary, append the starting index to the index list of the existing key;
[0098] Step E33: If the length of the current substring is greater than 1 and it does not exist in the empty dictionary, use the current substring as the fifth list, where the fifth list contains the tuple and the corresponding second dictionary value;
[0099] Step E34: If the length of the current substring is greater than 1 and it already exists in the empty dictionary, append the tuple to the index list of the existing key.
[0100] Step F: Mark and display the retrieved difference to complete the verification operation of the packaging box text information. Specifically, retrieve the corresponding sentence in the information original file at the incorrect position. If such an error exists, cancel the highlighting here.
[0101] A precise verification system for the text information of a packaging box based on OCR error processing, comprising an import and reading module, an image recognition module, a text rearrangement module, a text matching module, a difference comparison module, and a difference display module. The import and reading module is used to import the original text information file and read the ignoring rules to obtain the text fragments of the original text information file. The image recognition module is used to perform OCR recognition on the image area containing the original text information file to obtain the recognized text content. The text rearrangement module is used to rearrange the recognized text content to obtain the rearranged text content. The text matching module is used to perform Cartesian combination on the rearranged text content and the text fragments of the original text information file and use similarity matching and best sentence matching to obtain the matching text. The difference comparison module is used to perform difference comparison on the matching text and use the ignoring rules to retrieve the comparison result to obtain the retrieved difference. The difference display module is used to mark and display the retrieved difference to complete the verification operation of the text information of the packaging box.
[0102] In summary, for the precise verification method and system for the text information of a packaging box based on OCR error processing of the present invention, first, the original text information file is imported and the ignoring rules are read to obtain the text fragments of the original text information file. Then, OCR recognition is performed on the image area containing the original text information file to obtain the recognized text content. Next, the recognized text content is rearranged to obtain the rearranged text content. Then, Cartesian combination is performed on the rearranged text content and the text fragments of the original text information file, and similarity matching and best sentence matching are used to obtain the matching text. Subsequently, difference comparison is performed on the matching text, and the ignoring rules are used to retrieve the comparison result to obtain the retrieved difference. Then, the retrieved difference is marked and displayed to complete the verification operation of the text information of the packaging box. The present invention effectively processes the OCR errors through difference comparison and in combination with the ignoring rules, improves the accuracy of information verification, eliminates the interference of OCR errors, and ensures that the final result is more accurate and reliable. The present invention adopts an automated processing flow, reduces the dependence on manual intervention, improves the efficiency of information processing, reduces the workload of operators, and makes information verification more convenient. The present invention introduces flexible ignoring rules, which can ignore specific OCR misrecognition situations, thereby improving the system's tolerance to errors and enabling the system to better adapt to diverse text inputs. The present invention is not only applicable to the situation of the same content with different layouts, but also can adapt to the text information verification of different products and different layouts through flexible rules, and has strong versatility. The present invention quickly locates text differences through similarity matching and difference comparison, and enables operators to quickly focus on and handle specific problems through difference marking and display, improving work efficiency. With the continuous discovery of misrecognition, the ignoring rules of the system of the present invention will be gradually optimized, thereby continuously improving the fault tolerance mechanism and further enhancing the robustness of the system.
[0103] The basic principles, main features and advantages of the present invention have been shown and described above. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for accurately verifying text information on a packaging box based on OCR error processing, characterized in that: The following steps are included: Step A, importing the original text information file and reading the ignore rule to obtain the text fragments of the original text information file; Step B, performing OCR recognition on the image area of the original document containing text information to obtain the recognized text content; Step C, rearranging the recognized text content to obtain rearranged text content; Step D, Cartesian combination of the rearranged text content and the original text fragments of the text information and similarity matching and best sentence matching to obtain matching text; Step E, performing a difference comparison on the matching text and searching the comparison result using the ignore rule to obtain the difference after searching; Step F, marking and displaying the differences after retrieval, completing the packaging box text information verification operation; The specific steps of step A are as follows: Step A1, opening the original text information file in read-only mode and setting the encoding of the original text information file, and then reading the content of the original text information file line by line; Step A2, for each line in the original text information file, remove the blank characters at the beginning and end of the character string, and divide each line into two parts using a colon as a separator, which respectively represent different characters and the corresponding ignored character list; Step A3, split the string into a character list using a comma as a separator, and check whether each character in the split character list is a null string. If the check result is yes, replace the null string with an empty string; if the check result is no, keep it unchanged; Step A4, assigning the processed character list to the corresponding key in the first dictionary; The specific steps of step E are as follows: Step E1, using two layers of nested traversal to traverse all the strings in the comparison result string, wherein the outer layer loop traverses each starting index, and the inner layer loop traverses each ending index starting from the starting index plus 1; Step E2, in each inner loop, slice and compare the result string according to the start index and the end index, and obtain the current substring; Step E3, check whether the current substring exists in the second dictionary, the specific steps are as follows: Step E31, if the length of the current substring is 1 and does not exist in the empty dictionary, then use the current substring as the key and the value as the fourth list, where the fourth list contains the initial index list and the corresponding second dictionary value; Step E32, if the length of the current substring is 1 and already exists in the empty dictionary, then append the starting index to the index list of the existing key; Step E33, if the length of the current substring is greater than 1 and does not exist in the empty dictionary, the current substring is used as a fifth list, wherein the fifth list contains tuples and corresponding second dictionary values; Step E34, if the length of the current substring is greater than 1 and already exists in the empty dictionary, then append the tuple to the index list of the existing key.
2. According to claim 1, a method for accurately verifying package box text information based on OCR error processing is characterized in that: Step B, performing OCR recognition on the image area of the original document containing text information to obtain the recognized text content. The specific steps are as follows: Step B1, setting the image area R i The pixel value in is expressed as p(x,y), where (x,y) is the coordinate of the pixel, and the convolution expression is used to perform the image region R i Perform feature extraction and generate a feature vector F i As shown in formula (1), Among them, k(g,h) is the weight of the convolution kernel, M and N are half of the height and width of the convolution kernel respectively; Step B2: transform the feature vector F i Convert to text sequence T i , as shown in formula (2), T i =σ(∑ j W j F i (j)+b) (2) Among them, W j is the model weight, b is the bias term, and j is the feature vector F i The index of , σ is the softmax activation function, and the softmax activation function is used to convert the result of the weighted summation into a probability distribution and represent the possibility of each character.
3. The method for accurately verifying package box text information based on OCR error processing according to claim 2, characterized in that: Step C, rearrange the recognized text content to obtain rearranged text content, the specific steps are as follows: Step C1, initialize an empty list and use it to store the matching results, then traverse each text paragraph recognized by OCR, then call a function for each text paragraph to compare it with the paragraphs in the reference text paragraph set and find the paragraph with the highest similarity and its similarity, if the similarity is not 0, add this OCR text paragraph, the best matching paragraph and the similarity between the OCR text paragraph and the best matching paragraph as an element to the empty list, where the reference text paragraph set is the text fragment set in the original information file; Step C2, using list derivation to filter out matching items with a similarity greater than or equal to a threshold from an empty list and store them in a first list, then filtering out matching items with low similarity based on the similarity threshold and retaining only matching results with high similarity; Step C3, find the optimal similarity text pair matching function, the specific steps are as follows: Step C31, using a function to check whether the second list containing multiple reference text paragraphs is not empty, if the second list is not empty, the matching function enters the comparison process; Step C32, initializing the best matching paragraph as the first element in the reference text list and as the initial value of the current best matching item, then calculating the cosine similarity between the OCR text paragraph and the initial value and storing the calculation result in a third list; Step C33, using cosine similarity to calculate similarity of other text paragraphs in the second list, wherein specifically a matching function is used to receive two string parameters text1 and text2, and the two string parameters text1 and text2 respectively represent two texts to be compared for similarity, and the specific steps are as follows: Step C331, converting the two texts into word frequency vectors, specifically, analyzing all unique words in the texts, and then generating a sparse matrix according to the frequency of occurrence of each word in each text, where the sparse matrix X is shown in formula (3), Among them, each row represents the vectorized representation of a document, and each column corresponds to a word; Step C332, calculate the cosine similarity, specifically, use the similarity function to calculate the cosine similarity between the two word frequency vectors from the sparse matrix to obtain a similarity matrix, and then return the similarity matrix from the similarity function, wherein the similarity matrix includes the similarity values between the sparse matrices, calculate the cosine similarity as described in formula (4), and the similarity matrix is shown in formula (5), Among them, cosine_similarity is the similarity function, ‖X‖ and ‖Y‖ are the Euclidean norms of X and Y respectively, k is the number of features, K ij is the cosine similarity between the i-th sample in the sparse matrix X and the j-th sample in Y; Step C333, using the similarity matrix to build a vocabulary and convert the text into a word frequency vector, as shown in formula (6), D={d1,d2}(6) Among them, d1 and d2 are string parameters text1 and text2 respectively, and D is the word frequency vector; Step C334: remove duplicate words from all documents and form a vocabulary V = {v1, v2, ..., v n }, where n is the total number of words in the vocabulary, and then for each document d i Construct a vector x of the same length as the vocabulary V i ={x i1 ,x i2 ,…,x in }, where x ij Represents vocabulary v j In the document d i The number of occurrences in .
4. The method for accurately verifying package box text information based on OCR error processing according to claim 3, characterized in that: Step F, marking and displaying the differences after retrieval, completing the text information verification operation of the packaging box, wherein specifically, the corresponding sentence in the original information file is retrieved at the wrong position, and if such an error exists, the highlight here is cancelled.
5. A system for accurately verifying package box text information based on OCR error processing, the system for accurately verifying package box text information is based on the method for accurately verifying package box text information according to any one of claims 1 to 4, characterized in that: It includes an import and read module, an image recognition module, a text rearrangement module, a text matching module, a difference comparison module and a difference display module, wherein the import and read module is used to import the original text information file and read the ignore rules to obtain the text fragments of the original text information file; The image recognition module is used to perform OCR recognition on the image area of the original document containing text information to obtain the recognized text content; The text rearrangement module is used to rearrange the identified text content to obtain rearranged text content; The text matching module is used to perform Cartesian combination of the rearranged text content and the original text fragments of the text information, and use similarity matching and best sentence matching to obtain matching text; The difference comparison module is used to compare the matching texts and retrieve the comparison results using the ignore rule to obtain the post-retrieval differences; The difference display module is used to mark and display the differences after retrieval, completing the packaging box text information verification operation.
Citation Information
Patent Citations
Consistency auditing method for different source files
CN109190092A
Method and device for calculating text similarity based on text labels
CN116186582A