Archive box information intelligent extraction and optimization method based on multi-modal visual recognition

By quantifying environmental conditions and employing a multi-level comparison mechanism, combined with convolutional network modules to extract file box information, the problem of recognition accuracy in complex scenarios was solved, achieving high robustness and continuously optimized recognition results.

CN122313488APending Publication Date: 2026-06-30STATE GRID BEIJING ELECTRIC POWER CO +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID BEIJING ELECTRIC POWER CO
Filing Date
2026-03-27
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

Existing technologies have low accuracy in recognizing information from file boxes in complex real-world scenarios, and struggle to overcome interference factors such as changes in lighting, reflections, shadows, fading fonts, and localized dirt.

Method used

By sensing and quantifying environmental conditions such as lighting and angle of the file box, and combining convolutional network modules to extract texture and character features of the identification area, a multi-level comparison mechanism is used to build a sample model for optimization training, thereby improving recognition accuracy.

Benefits of technology

It improves the robustness and accuracy of recognition in complex environments, and achieves a steady improvement in long-term recognition performance through continuous feedback to optimize the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122313488A_ABST
    Figure CN122313488A_ABST
Patent Text Reader

Abstract

This invention relates to the field of archival management and addresses the problem that traditional visual recognition solutions lack the means to cope with interference factors such as changes in lighting, reflections, and non-destructive factors in complex real-world scenarios, leading to difficulties in ensuring the accuracy of the final data. Specifically, it is a method for intelligent extraction and optimization of archival box information based on multimodal visual recognition. This invention senses and quantifies the environmental conditions of the archival box, such as lighting and angles, and combines the dual extraction of texture and character features of the identification area by a convolutional network extraction module, as well as the comparison and selection mechanism of the archival information comparison module. The system can effectively overcome interference caused by adverse factors such as reflections, shadows, font fading, and local dirt. At the same time, it also evaluates the accuracy of the recognition results and constructs a sample model for judging accurate and incorrect judgments, which is continuously fed back to the convolutional network and comparison module for model retraining, achieving a steady improvement in recognition performance over long-term use.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of archives management, and more specifically, to a method for intelligent extraction and optimization of archive box information based on multimodal visual recognition. Background Technology

[0002] With the continuous deepening of information technology construction, archives management is undergoing a transformation from traditional physical management to digital and intelligent management. In this process, quickly and accurately converting the information on physical archive boxes into structured digital information is the primary step in achieving efficient archives retrieval, inventory, and management. Currently, the technological development in this field has mainly gone through the following stages;

[0003] The earliest and most basic method was manual data entry. Staff would visually inspect the information on the spine or cover of the file box and manually input it into the computer system.

[0004] To improve efficiency, many modern file boxes are labeled with one-dimensional or two-dimensional codes containing basic identification information. However, this is completely ineffective for a large number of historical files, files without codes, or files with damaged or missing labels.

[0005] In recent years, text detection and recognition models have been applied in this field. By using optical character recognition technology to directly read text information in images, it is possible to achieve automated collection of direct text information, get rid of dependence on specific labels, and have a wider range of applications. However, its recognition performance is highly dependent on image quality. In actual warehouse environments, factors such as the placement angle of file boxes, uneven lighting, reflection, shadows, bending and deformation of the box ridge, fading of handwriting, and stains can all lead to a sharp drop in recognition accuracy.

[0006] To address the aforementioned technical problems, this application proposes a solution. Summary of the Invention

[0007] This invention, by sensing and quantifying environmental conditions such as lighting and angle of the file box, combined with the dual extraction of texture and character features of the identification area by the convolutional network extraction module, and the comparison and selection mechanism of the file information comparison module, can effectively overcome interference caused by adverse factors such as reflection, shadow, font fading, and local dirt. Compared with traditional methods that rely solely on character recognition, it exhibits stronger robustness and higher recognition accuracy when facing complex real-world scenarios. At the same time, it also evaluates the accuracy of the recognition results and constructs a sample model for judging accuracy and error, which is continuously fed back to the convolutional network and comparison module for model retraining, achieving a steady improvement in recognition performance over long-term use. This solves the problem that traditional visual recognition schemes lack the means to deal with interference factors such as lighting changes, reflection, and non-destructive factors in complex real-world scenarios, making it difficult to guarantee the accuracy of the final data. Therefore, this invention proposes an intelligent extraction and optimization method for file box information based on multimodal visual recognition.

[0008] The objective of this invention can be achieved through the following technical solutions:

[0009] A method for intelligent extraction and optimization of file box information based on multimodal visual recognition includes the following steps:

[0010] Step 1: Collect environmental parameters to obtain the real-time ambient color temperature and ambient brightness of the file storage location;

[0011] Step 2: Acquire the image of the document identification area and correct the image of the identification area based on the ambient color temperature;

[0012] Step 3: Based on ambient brightness, extract features of different weights from the image of the marked area, and compare the extracted features with the feature library at multiple levels to confirm the specific information of the marked area of ​​the archive step by step.

[0013] Step 4: Based on the comparison results in Step 3, the final recognition result is confirmed through a precise comparison of the complete image.

[0014] Step 5: Determine the accuracy of the recognition based on the recognition results, and generate a sample library of accurate or incorrect judgments based on the judgment results. Use the sample library to optimize and train the system again.

[0015] As a preferred embodiment of the present invention, it also includes an environmental parameter generation module, a convolutional network preparation module, an archive information comparison module, a result generation module, and a result feedback optimization module;

[0016] The environmental parameter generation module can collect environmental information about the location of the file box and obtain environmental information.

[0017] The convolutional network extraction module acquires images of the marked areas on the file box through a camera device, identifies the marked areas through a convolutional neural network, extracts texture and character features, and simultaneously sends the texture and character features to the file information comparison module.

[0018] The archive information comparison module compares the texture and character features separately, selects the best one based on the degree of comparison approximation, generates a high-confidence result, and sends the high-confidence result to the result generation module.

[0019] The result generation module obtains high-confidence results and completes the archive information based on the high-confidence results to obtain complete information recognition results;

[0020] The result feedback optimization module acquires the information recognition results and can judge the accuracy of the information recognition results. It also performs statistical analysis on the judgment results to obtain accurate judgment models and incorrect judgment models. The result feedback optimization module sends the accurate judgment models and incorrect judgment models to the convolutional network extraction module and the archive information comparison module for model training.

[0021] In a preferred embodiment of the present invention, the environmental parameter generation module collects environmental information through a sensor capable of recognizing color temperature and brightness, specifically collecting environmental color temperature and ambient brightness.

[0022] In a preferred embodiment of the present invention, the convolutional network extraction module acquires an image of the marked area on the file box through a camera device, and simultaneously imports the acquired ambient color temperature and ambient brightness into an environmental interference model to obtain the degree of mark distortion, wherein the degree of mark distortion includes RGB variation and feature loss.

[0023] In a preferred embodiment of the present invention, when the convolutional network extraction module performs texture and character feature recognition, it first restores the RGB image of the marked area using a color overlay model based on the ambient color temperature to obtain the color-restored marked area image. Then, it analyzes the ambient brightness using a big data model to obtain the recognition satisfaction level of the current ambient brightness. The recognition satisfaction level is then compared with a set threshold level to obtain a high recognition level, a medium recognition level, and a low recognition level.

[0024] In a preferred embodiment of the present invention, the convolutional network extraction module extracts features from the color-restored identification area image according to the recognition level. If the recognition level is high, an accurate recognition scheme is adopted; if the recognition level is medium or low, a fuzzy recognition scheme of different degrees is adopted.

[0025] In a preferred embodiment of the present invention, when the convolutional network extraction module adopts a precise recognition scheme, it uses precise comparison. If a low-degree fuzzy recognition scheme is adopted, a set association ratio X1 is selected to perform feature recognition on the color-restored marker area image. When the overlap of features is greater than the set ratio X1, the feature recognition is determined to be complete. If a high-degree fuzzy recognition scheme is adopted, a set association ratio X2 is selected to perform feature recognition on the color-restored marker area image. When the overlap of features is greater than the set ratio X2, the feature recognition is determined to be complete, where X1 > X2.

[0026] In a preferred embodiment of the present invention, the archive information comparison module selects texture features, compares the texture features with texture identifiers in the feature library to obtain the comparison overlap, and selects texture identifiers with a comparison overlap greater than a set threshold as a first screening result.

[0027] The archive information comparison module obtains the character identifiers of the formation in the first screening result, and compares the character features with the character identifiers again to obtain the comparison overlap. The character identifiers with the comparison overlap greater than a set threshold are selected as the second screening result. At the same time, the archive information comparison module sorts the comparison overlap in the second screening result and selects the i group with the highest comparison overlap as the high confidence result, where i is a constant set by the user.

[0028] In a preferred embodiment of the present invention, the result generation module obtains i sets of high-confidence results and retrieves i sets of comprehensive archive information from the database based on the i sets of high-confidence results. The color-restored identification region image obtained by the convolutional network extraction module is compared with the i sets of comprehensive archive information one by one, and the set with the highest comparison overlap is taken as the final recognition result.

[0029] Compared with the prior art, the beneficial effects of the present invention are:

[0030] 1. This invention provides key information for subsequent recognition by sensing and quantifying the environmental conditions such as lighting and angle of the file box. Combined with the dual extraction of texture and character features of the identification area by the convolutional network extraction module, and the comparison and selection mechanism of the file information comparison module, the system can effectively overcome the interference caused by adverse factors such as reflection, shadow, font fading, and local dirt. Compared with the traditional method that relies solely on character recognition, it shows stronger robustness and higher recognition accuracy when facing complex real-world scenarios.

[0031] 2. This invention not only outputs the final result, but also evaluates the accuracy of the recognition result and constructs a sample model for accurate and erroneous judgments based on this. This model is continuously fed back to the convolutional network and the comparison module for model retraining, enabling the system to learn from recognition practice, continuously correct the weights of feature extraction and the thresholds of comparison judgment, and continuously optimize for the unique style of a specific archive, effectively suppressing the recurrence of similar errors, thereby achieving a steady improvement in recognition performance over long-term use. Attached Figure Description

[0032] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:

[0033] Figure 1 This is a system block diagram of the present invention;

[0034] Figure 2 This is a system flowchart of the present invention. Detailed Implementation

[0035] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0036] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0037] Example 1

[0038] Figure 1 This is a system block diagram of the present invention. Figure 2 The system flowchart of the present invention is shown below. Figure 1 - Figure 2As shown, the intelligent extraction and optimization method for file box information based on multimodal visual recognition includes the following steps:

[0039] Step 1: Collect environmental parameters to obtain the real-time ambient color temperature and ambient brightness of the archive storage location, providing a corrected data basis for subsequent archive information feature extraction;

[0040] Step 2: Acquire the image of the document identification area and correct the image of the identification area based on the ambient color temperature to obtain the color-restored image of the identification area, thereby avoiding the RGB deviation caused by lighting differences from affecting the final texture recognition result;

[0041] Step 3: Based on ambient brightness, feature extraction is performed on the image of the marked area with different weights. The extracted features are compared with the feature library at multiple levels to confirm the specific information of the archive marked area step by step. Thus, under different lighting conditions, the archive information is automatically associated based on visibility, which can ensure the accuracy of feature extraction to the greatest extent. When the feature recognition is insufficient, the recognition standard can be lowered and the degree of association can be increased to ensure the recognition effect.

[0042] Step 4: Based on the comparison results in Step 3, the final recognition result is confirmed through a precise comparison of the complete image.

[0043] Step 5: Determine the accuracy of the recognition based on the recognition results, and generate a sample library of accurate or incorrect judgments based on the judgment results. Use the sample library to optimize and train the system again, so that the system can be continuously optimized during use and effectively suppress the recurrence of similar errors.

[0044] Example 2

[0045] Please see Figure 1 - Figure 2 As shown, the intelligent extraction and optimization method for archive box information based on multimodal visual recognition includes an environmental parameter generation module, a convolutional network preparation module, an archive information comparison module, a result generation module, and a result feedback optimization module.

[0046] The environmental parameter generation module can collect environmental information about the location of the file box and obtain environmental information. The specific items collected include ambient color temperature and ambient brightness. After obtaining the ambient color temperature and ambient brightness, the environmental parameter generation module sends them to the convolutional network extraction module.

[0047] The convolutional network extraction module acquires images of the marked areas on the file box through a camera device. At the same time, it imports the acquired ambient color temperature and ambient brightness into the environmental interference model to obtain the degree of marking distortion, which includes RGB variation and feature loss. The module then uses a convolutional neural network to identify the marked areas, extract texture and character features, and simultaneously sends the texture and character features to the file information comparison module.

[0048] Specifically, when the convolutional network extraction module performs texture and character feature recognition, it first uses the ambient color temperature to perform a color overlay model to restore the RGB image of the marked area, and then analyzes the ambient brightness through a big data model to obtain the recognition satisfaction level of the current ambient brightness. The recognition satisfaction level is then compared with the set threshold level to obtain the high recognition level, medium recognition level, and low recognition level.

[0049] The convolutional network extraction module extracts features from the color-restored label area image based on the recognition level. Specifically, if the recognition level is high, an accurate recognition scheme is used; if the recognition level is medium or low, a fuzzy recognition scheme of different degrees is used.

[0050] When the convolutional network extraction module adopts the precise recognition scheme, it does not perform feature association supplementation on the color-restored label area image. Instead, it uses precise comparison. If a low-level fuzzy recognition scheme is adopted, a set association ratio X1 is selected to perform feature recognition on the color-restored label area image. If a high-level fuzzy recognition scheme is adopted, a set association ratio X2 is selected to perform feature recognition on the color-restored label area image. Where X1 > X2, that is, when using the association ratio X1, if the degree of feature overlap is greater than the set ratio X1, the feature recognition is considered complete. When using the association ratio X2, if the degree of feature overlap is greater than the set ratio X2, the feature recognition is considered complete.

[0051] The features identified by the convolutional network extraction module include texture and character features;

[0052] The archive information comparison module selects texture features, compares the texture features with texture identifiers in the feature library, obtains the comparison overlap, and selects texture identifiers with a comparison overlap greater than a set threshold as the first screening result;

[0053] The archive information comparison module obtains the character identifiers of the formation in the first screening result, and compares the character features with the character identifiers again to obtain the comparison overlap. The character identifiers with a comparison overlap greater than a set threshold are selected as the second screening result. At the same time, the archive information comparison module sorts the comparison overlap in the second screening result and selects the i group with the highest comparison overlap as the high confidence result, where i is a constant set by the user.

[0054] The archive information comparison module sends high-confidence results to the result generation module;

[0055] The result generation module acquires i sets of high-confidence results and retrieves i sets of comprehensive archival information from the database based on these results. It then compares the color-restored label region image obtained by the convolutional network extraction module with the i sets of comprehensive archival information one by one. The set with the highest overlap is taken as the final recognition result, resulting in a complete information recognition result. The result is then saved for subsequent operations of the archival management robot.

[0056] Example 3

[0057] Please see Figure 1 - Figure 2 As shown, the result feedback optimization module acquires the information recognition results and displays the machine-recognized archival information alongside the original image through a human-computer interaction interface. Humans then judge and label the correctness of each result, forming high-quality training sample pairs with accurate labels. The judgment results are statistically analyzed to obtain an accurate judgment model and a misjudgment model. The accurate judgment model is used to summarize and solidify the feature combinations and decision paths leading to successful recognition. The misjudgment model, on the other hand, collects typical types of recognition errors, including character segmentation errors, misjudgment of similar character shapes, and interference from complex backgrounds. It analyzes the root causes of each error and their corresponding feature manifestations. The two models are encapsulated into a structured knowledge package. The result feedback optimization module sends the accurate judgment model and the misjudgment model to the convolutional network extraction module and the archival information comparison module in real-time and periodically for targeted backpropagation training. This optimizes the network weights, enabling better extraction of robust features against interference in the future. Simultaneously, it dynamically adjusts the internal comparison algorithm and selection threshold to optimize the ability to distinguish easily confused characters, thereby improving the accuracy of decision-making.

[0058] Thresholds, preset values, preset ranges, etc. are set for result comparison and analysis to determine whether they are good or bad. The value of these thresholds is determined by a combination of large-scale model analysis of sample data and human experience. They can also be adjusted appropriately based on seasonal or common-sense influences.

[0059] Furthermore, the settings for weighting ratios, influence factors, etc., are based on the magnitude of each parameter's influence on the results. The specific values ​​are allocated to ultimately reflect the impact on the results. The settings for input and storage are also determined by a combination of large-scale model analysis of sample data and human experience. Appropriate adjustments can also be made based on seasonal or rational influence conditions.

[0060] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to any specific implementation. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

[0061] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0062] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present invention.

Claims

1. A method for intelligent extraction and optimization of file box information based on multimodal visual recognition, characterized in that, include: Step 1: Collect environmental parameters to obtain the real-time ambient color temperature and ambient brightness of the file storage location; Step 2: Acquire the image of the document identification area and correct the image of the identification area based on the ambient color temperature; Step 3: Based on ambient brightness, extract features of different weights from the image of the marked area, and compare the extracted features with the feature library at multiple levels to confirm the specific information of the marked area of ​​the archive step by step. Step 4: Based on the comparison results in Step 3, the final recognition result is confirmed through a precise comparison of the complete image. Step 5: Determine the accuracy of the recognition based on the recognition results, and generate a sample library of accurate or incorrect judgments based on the judgment results. Use the sample library to optimize and train the system again.

2. The method for intelligent extraction and optimization of file box information based on multimodal visual recognition according to claim 1, characterized in that, It also includes an environmental parameter generation module, a convolutional network preparation module, an archive information comparison module, and a result generation module; The environmental parameter generation module can collect environmental information about the location of the file box and obtain environmental information. The convolutional network extraction module acquires images of the marked areas on the file box through a camera device, identifies the marked areas through a convolutional neural network, extracts texture and character features, and simultaneously sends the texture and character features to the file information comparison module. The archive information comparison module compares the texture and character features separately, selects the best one based on the degree of comparison approximation, generates a high-confidence result, and sends the high-confidence result to the result generation module. The result generation module obtains high-confidence results and completes the archive information based on the high-confidence results to obtain complete information recognition results.

3. The method for intelligent extraction and optimization of file box information based on multimodal visual recognition according to claim 2, characterized in that, The environmental parameter generation module collects environmental information through sensors that can identify color temperature and brightness. The specific items collected include ambient color temperature and ambient brightness.

4. The method for intelligent extraction and optimization of file box information based on multimodal visual recognition according to claim 2, characterized in that, When the convolutional network extraction module performs texture and character feature recognition, it first uses the ambient color temperature to perform a color overlay model to restore the RGB image of the marked area, and then analyzes the ambient brightness using a big data model to obtain the recognition satisfaction level of the current ambient brightness. The recognition satisfaction level is then compared with a set threshold level to obtain a high recognition level, a medium recognition level, and a low recognition level. The convolutional network extraction module extracts features from the color-restored identification area image according to the recognition level. If the recognition level is high, an accurate recognition scheme is adopted; if the recognition level is medium or low, a fuzzy recognition scheme of different degrees is adopted.

5. The method for intelligent extraction and optimization of file box information based on multimodal visual recognition according to claim 4, characterized in that, When the convolutional network extraction module adopts a precise recognition scheme, it uses precise comparison. If a low-level fuzzy recognition scheme is adopted, a set association ratio X1 is selected to perform feature recognition on the color-restored marker area image. When the overlap of features is greater than the set ratio X1, the feature recognition is considered complete. If a high-level fuzzy recognition scheme is adopted, a set association ratio X2 is selected to perform feature recognition on the color-restored marker area image. When the overlap of features is greater than the set ratio X2, the feature recognition is considered complete, where X1 > X2.

6. The method for intelligent extraction and optimization of file box information based on multimodal visual recognition according to claim 2, characterized in that, The archive information comparison module selects texture features, compares the texture features with texture identifiers in the feature library, obtains the comparison overlap, and selects texture identifiers with a comparison overlap greater than a set threshold as a first screening result. The archive information comparison module obtains the character identifiers of the formation in the first screening result, and compares the character features with the character identifiers again to obtain the comparison overlap. The character identifiers with the comparison overlap greater than a set threshold are selected as the second screening result. At the same time, the archive information comparison module sorts the comparison overlap in the second screening result and selects the i group with the highest comparison overlap as the high confidence result, where i is a constant set by the user.

7. The method for intelligent extraction and optimization of file box information based on multimodal visual recognition according to claim 2, characterized in that, The result generation module obtains i sets of high-confidence results and retrieves i sets of comprehensive archive information from the database based on the i sets of high-confidence results. It then compares the color-restored identification region image obtained by the convolutional network extraction module with the i sets of comprehensive archive information one by one, and takes the set with the highest degree of overlap as the final recognition result.

8. The method for intelligent extraction and optimization of file box information based on multimodal visual recognition according to claim 1, characterized in that, It also includes a result feedback optimization module, which acquires the information recognition results and can judge the accuracy of the information recognition results. It then performs statistical analysis on the judgment results to obtain an accurate judgment model and an incorrect judgment model. The result feedback optimization module sends the accurate judgment model and the incorrect judgment model to the convolutional network extraction module and the archive information comparison module for model training.