Image recognition-based archive management method and system, and electronic device

WO2026193996A1PCT designated stage Publication Date: 2026-09-24SHANGHAI TAIYU INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/088053
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-21
Filing Date
2025-04-09
Publication Date
2026-09-24

Smart Images

  • Figure CN2025088053_24092026_PF_FP_ABST
    Figure CN2025088053_24092026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of archive information security screening, and in particular to an image recognition-based archive management method and system, and an electronic device. The method comprises: on the basis of a preset digital database, acquiring a picture set to be recognized corresponding to an archive to be managed, and performing region division and marking on pictures to be recognized in said picture set to obtain marked regions to be recognized; on the basis of said marked regions, acquiring images to be recognized, and sequentially inputting said images into a preset training model to obtain a feature set to be processed corresponding to said image; sequentially performing classification and recognition on features to be processed in said feature set to obtain marking information corresponding to said features; and on the basis of a result type, performing security classification marking on said images to obtain security-classification-marked images. The present application can quickly determine the type of an archive to be managed, thereby avoiding manual recognition operations and improving the type recognition efficiency of said archive.
Need to check novelty before this filing date? Find Prior Art

Description

A method, system, and electronic device for managing archives using image recognition. Technical Field

[0001] This application relates to the technical field of archival information security screening, and in particular to an image recognition archival management method, system, and electronic device. Background Technology

[0002] The field of archival management encompasses the collection, organization, preservation, retrieval, and reuse of archives. It ensures the security and accessibility of historical data by providing structured information management strategies. With the rapid development of information technology and the advancement of archival digitization, archival management has expanded from traditional physical storage to the digital processing of electronic archives. Relevant archival management systems place particular emphasis on the digital protection of archives, retrieval efficiency, and data security.

[0003] To improve the efficiency of records management, related technologies utilize AI technologies such as machine learning to enhance the intelligent processing capabilities of records management, automatically classify and archive records, identify and extract key information, and optimize search algorithms to improve retrieval efficiency.

[0004] However, the archives to be managed include classified and unclassified archives. Classified and unclassified archives are stored together and there is no clear label on the catalog, making it impossible to directly distinguish between classified and unclassified archives. Furthermore, due to the large amount of digital resources in the archives to be managed, manual identification alone is inefficient. Summary of the Invention

[0005] To facilitate the automatic identification of files to be managed and quickly distinguish the confidentiality type of the files to be managed, this application provides an image recognition file management method, system, and electronic device.

[0006] In a first aspect, this application provides an image recognition-based file management method, comprising the following steps: obtaining a set of images to be recognized corresponding to the files to be managed based on a preset digital database; dividing and marking the images to be recognized in the set of images to be recognized into regions to obtain marked regions to be recognized; obtaining images to be recognized based on the marked regions to be recognized, and sequentially inputting the images to be recognized into a preset training model to obtain a set of features to be processed corresponding to the images to be recognized, the set of features to be processed including at least one feature to be processed; sequentially classifying and recognizing the features to be processed in the set of features to be processed to obtain marking information corresponding to the features to be processed, the marking information including result type; marking the images to be recognized with a security classification based on the result type to obtain a security-marked image; and managing and storing the files to be managed based on the security-marked image.

[0007] By adopting the above technical solution, the images to be identified corresponding to the archives to be managed are divided into regions and marked to obtain the marked regions to be identified. Then, the images to be identified corresponding to the marked regions are identified by a preset training model to obtain a set of features to be processed. By classifying and identifying the features to be processed in the set of features to be processed, different result types are obtained. Based on the result types, the images to be identified are marked to obtain security classification images. This improves the management of archives to be managed and facilitates the automatic identification of archives to be managed. By obtaining the result types of security classification images, the confidentiality type of archives to be managed can be quickly distinguished, thereby avoiding manual identification operations and improving the efficiency of type identification of archives to be managed.

[0008] In one embodiment, the result type includes at least one piece of data. The process involves sequentially classifying and identifying the features to be processed in the feature set to obtain the corresponding label information. This includes the following steps: determining whether the feature to be processed has label information; if the feature to be processed does not have label information, then sequentially performing key feature comparison on the features to be processed to obtain the corresponding category data set, where the category data set includes at least one category data item, and the category data corresponds one-to-one with the feature to be processed; determining whether the category data in the category data set are the same; if the category data in the category data set are the same, then using one of the category data items as the result type corresponding to the image to be recognized; if the category data in the category data set are different, then obtaining the label category based on the category data and a preset label category, and using the label category as the result type of the image to be recognized.

[0009] By adopting the above technical solution, the category data set is obtained by sequentially comparing key features of the features to be processed, and the category data is judged to obtain the result type corresponding to the current image to be identified. The category data corresponding to each feature to be processed is marked accurately step by step, and the confidential area of ​​the archives to be managed is accurately found, thereby improving the management security of the archives to be managed.

[0010] In one embodiment, key feature comparisons are performed sequentially on the features to be processed to obtain the corresponding category data set, including the following steps: determining whether the feature to be processed is a processable feature; if the feature to be processed is a processable feature, key feature comparisons are performed based on the feature to be processed to obtain the category data corresponding to the feature to be processed; if the feature to be processed is not a processable type, the feature to be processed is segmented to obtain the corresponding branch features, and key feature comparisons are performed on the branch features respectively to obtain the corresponding branch result type, and the branch result type is marked according to a preset security level to obtain the category data corresponding to the feature to be processed.

[0011] By adopting the above technical solution, it is determined whether the feature to be processed is a processable feature. If so, key feature comparison is performed based on the feature to be processed to obtain the category data corresponding to the feature to be processed. If not, feature segmentation is required to obtain branch features. Key feature comparison is performed on each branch feature to obtain the corresponding branch result type. The branch result type is marked according to the preset security level to obtain the category data corresponding to the feature to be processed. Then, the category data is obtained in different ways according to the different features to be processed, which can improve the accuracy of obtaining the category data corresponding to the feature to be processed, accurately find the confidential area in the archive to be managed, and thus improve the management security of the archive to be managed.

[0012] In one embodiment, determining whether a feature to be processed is a processable feature includes the following steps: obtaining a feature identifier based on the feature to be processed, and determining whether the feature identifier is unique, wherein the feature identifier represents the type of the feature to be processed; if the feature identifier is unique, then the feature to be processed is determined to be a processable feature; if the feature identifier is not unique, then the feature to be processed is determined to be a processable feature.

[0013] By adopting the above technical solution, for complex images with multiple feature identifiers to be processed, it is necessary to obtain branch features and perform key feature comparison on each branch feature to obtain the corresponding branch result type. The branch result type is marked according to a preset security level to obtain the category data corresponding to the feature to be processed. In this way, it is possible to accurately filter out the feature to be processed that has mixed types. By obtaining the corresponding category data in different ways, the accuracy of obtaining the category data of the feature to be processed can be improved.

[0014] In one embodiment, the images to be identified in the set of images to be identified are divided into regions and marked to obtain the marked regions to be identified. This includes the following steps: sequentially scanning the images to be identified according to a preset interval based on a preset scanning route to obtain several sets of scanning regions. The sets of scanning regions include the scanning regions and the data types corresponding to the scanning regions; taking the continuous scanning regions with the same data type as the regions to be determined, and marking the regions to be determined to obtain the marked regions to be identified.

[0015] By adopting the above technical solution, the set of images to be identified is scanned sequentially to obtain the corresponding set of scanned areas. Then, the scanned areas with the same data type and continuous scans are taken as the areas to be determined and marked to obtain the marked areas to be identified. This facilitates the subsequent analysis of the marked areas to be identified, improves the acquisition efficiency of the images to be identified, facilitates the subsequent processing of the images to be identified, and can quickly distinguish the confidentiality type of the files to be managed, thereby improving the management efficiency of the files to be managed.

[0016] In one embodiment, obtaining an image to be identified based on a region to be identified includes the following steps: matching a corresponding preprocessing model in a preprocessing database based on the region to be identified, wherein the preprocessing database stores several sets of preprocessing models, which are models used to perform clear processing on the region to be identified; generating a set of images to be identified based on the preprocessing models, wherein the set of images to be identified includes at least one image to be identified.

[0017] By adopting the above technical solution and processing based on the preprocessing model, clear images to be identified are obtained, thereby forming a corresponding set of images to be identified, which facilitates subsequent operations, improves the efficiency of obtaining classified images, and enhances the management security of archives to be managed.

[0018] In one embodiment, after obtaining the feature identifier based on the feature to be processed, the method further includes the following steps: determining whether the feature identifier corresponding to the feature to be processed is a text identifier; if the feature identifier is a text identifier, obtaining several feature words to be identified based on the feature to be processed using feature recognition technology; filtering the feature words to be identified in a preset database and determining whether corresponding comparison feature words can be obtained in the preset database; if corresponding comparison feature words can be obtained in the preset database, filtering the feature words to be identified in the feature set to be processed to obtain the filtering position, the filtering position representing the position of the comparison feature words in the image to be identified; obtaining the corresponding comparison result based on the comparison feature words, and marking the feature to be processed based on the comparison result and the filtering position to obtain the marking information.

[0019] By adopting the above technical solution, the identification of each feature to be processed is reduced by pre-labeling based on the feature words to be identified, thereby improving the efficiency of acquiring classified images and thus improving the efficiency of the work of protecting and keeping the archives to be managed.

[0020] Secondly, this application provides an image recognition-based file management system, comprising: an image acquisition module, which acquires a set of images to be recognized corresponding to the files to be managed based on a preset digital database, divides and marks the images to be recognized in the set of images to be recognized into regions to obtain marked regions to be recognized; a feature extraction module, which acquires images to be recognized based on the marked regions to be recognized, and sequentially inputs the images to be recognized into a preset training model to obtain a set of features to be processed corresponding to the images to be recognized, the set of features to be processed including at least one feature to be processed; a classification and identification module, which sequentially classifies and identifies the features to be processed in the set of features to be processed to obtain marking information corresponding to the images to be recognized, the marking information including a result type; and a marking generation module, which marks the images to be recognized with a security classification based on the result type to obtain a security-marked image, and manages and stores the files to be managed based on the security-marked image.

[0021] By adopting the above technical solution, the image acquisition module acquires images to be identified from the digitized paper archives according to set rules, divides and marks the images to be identified, and generates marked regions to be identified, ensuring that the acquired images meet the requirements of subsequent processing. The feature extraction module is responsible for extracting the set of features to be processed corresponding to various types of images to be identified, so that the classification and identification module can classify and identify the images to be identified based on the features to be processed. The mark generation module marks the images to be identified based on the mark information, which facilitates the subsequent scanning of the digitized paper archives, can quickly distinguish the confidentiality type of the archives to be processed, and improves the efficiency of the work of managing the security and confidentiality of the archives.

[0022] In one embodiment, the security classification marker includes a hyperlink marker, and the marker information also includes image features and location information. It also includes a result output module, which is used to display the result type, image features and location information in a list form on the display screen for staff to view.

[0023] Thirdly, this application provides an electronic device including a processor and a memory coupled to each other, wherein the memory stores a computer program that can run on the processor; when the computer program is executed by the processor, it implements the file management method of image recognition as described in the first aspect. Attached Figure Description

[0024] Figure 1 is a block diagram of an image recognition file management method provided in an embodiment of this application;

[0025] Figure 2 is a block diagram of a method for obtaining a region to be identified according to an embodiment of this application;

[0026] Figure 3 is a schematic diagram of the preset scanning route provided in an embodiment of this application;

[0027] Figure 4 is another scanning schematic diagram of the preset scanning route provided in the embodiment of this application;

[0028] Figure 5 is a block diagram of the result type acquisition method provided in the embodiments of this application;

[0029] Figure 6 is a schematic diagram of the structure of the file management system provided in this embodiment;

[0030] Figure 7 is a structural block diagram of the electronic device provided in this embodiment.

[0031] Figure labeling: 10, image acquisition module; 20, feature extraction module; 30, classification and identification module; 40, label generation module; 50, result output module; 61, processor; 62, memory; 63, computer program. Detailed Implementation

[0032] To better understand the purpose, technical solutions, and advantages of this application, it has been described and illustrated below with reference to the accompanying drawings and embodiments. However, those skilled in the art should understand that this application can be implemented without these details. In some cases, to avoid obscuring various aspects of this application due to unnecessary description, well-known methods, processes, systems, components, and / or circuits already described at a higher level will not be elaborated upon. It will be apparent to those skilled in the art that various modifications can be made to the embodiments disclosed in this application, and the general principles defined in this application can be applied to other embodiments and application scenarios without departing from the principles and scope of this application. Therefore, this application is not limited to the illustrated embodiments, but conforms to the broadest scope consistent with the scope of protection claimed in this application.

[0033] This application discloses an image recognition-based document management method applied to a document management system. Specifically, the document management system marks the documents to be managed with a security classification. Subsequently, by querying the paper-based digitized database and parsing the file path, a list containing information such as the classified document number and its storage location is output, which facilitates subsequent stripping and separate storage operations by document management personnel.

[0034] As shown in Figure 1, the image recognition-based file management method includes the following steps:

[0035] S100: Based on a preset digital database, obtain the set of images to be identified corresponding to the archives to be managed, divide and mark the images to be identified in the set of images to be identified, so as to obtain the marked areas to be identified.

[0036] The system includes a pre-defined digital database that stores the digital data corresponding to paper archives. This database can be created by scanning paper archives into PDF format or by manually inputting the data. Paper archives with official seals can be directly obtained via PDF scanning. The archives to be managed represent those requiring security classification and marking. The set of images to be identified represents the set of images corresponding to the archives to be managed in the pre-defined digital database. The marked regions represent the data formed by dividing and marking each image to be identified. The same image may have multiple marked regions, and each marked region obtained from the image to be identified does not overlap. Furthermore, the marked regions are laid out to cover the image to be identified.

[0037] Referring to Figure 2, it should be noted that dividing and labeling the images in the image set to be identified to obtain the labeled regions includes the following steps:

[0038] S110, sequentially scans the image to be identified according to a preset scan route within a preset interval to obtain several sets of scan areas.

[0039] S120: Select consecutive scan areas with the same data type as the region to be determined, and mark the region to be determined to obtain the region to be identified.

[0040] The preset interval represents the size of the pre-defined scanning area. The size of the preset interval is set according to the size of the image to be recognized. In order to facilitate the recognition later, the size of the preset interval is definitely smaller than the size of the image to be recognized.

[0041] The preset scanning route is the scanning route for the image to be recognized. The preset scanning route is determined based on the size of the preset interval. When the borders of the preset interval are all smaller than the borders of the image to be recognized, the scanning route may include scanning sequentially along the short border of the image to be recognized or scanning sequentially along the long border of the image to be recognized. When the borders of the preset interval are all equal to one border of the image to be recognized, the scanning proceeds to the opposite side based on the equal border of the image to be recognized.

[0042] For example, as shown in Figure 3, there is an image B to be identified, a preset interval A1, and a preset interval A2. When the preset interval in this embodiment is A1, the preset scanning route scans the image B to be identified according to the symbol F1 in the figure until the entire image B to be identified is scanned. When the preset interval in this embodiment is A2, the preset scanning route scans the image B to be identified according to the symbol F2 in the figure until the entire image B to be identified is scanned.

[0043] For example, when the preset interval in this embodiment is shown as preset interval A3, the preset scanning route scans the image B to be identified according to the label F3 in Figure 4.

[0044] It should be noted that the preset scanning route is not limited to the scanning method provided in this application. Other scanning methods can also be used. For example, when the preset interval is preset interval A3, the image to be identified can be scanned sequentially along the short side of the image to be identified, as long as the scanning of the image to be identified can be completed quickly.

[0045] The scan area set includes the scan area and the data type corresponding to the scan area. In order to facilitate the acquisition of the image to be identified, scan areas with the same data type and continuous scan areas are regarded as the same undetermined area, and feature area marker boxes are added to the undetermined area. Each marker box is assigned a unique ID to facilitate subsequent tracking and processing. The position coordinates of the marker box, the image number to which it belongs, and other information are recorded in a special metadata file to ensure the traceability of the data.

[0046] It should be noted that after obtaining the set of images to be identified based on the files to be managed, the images to be identified need to be numbered according to the order of the paper documents in the files to be managed.

[0047] Furthermore, the acquisition of the image set to be identified utilizes an efficient image acquisition script to batch acquire the image set from the database of digitized paper archives in a parallel processing manner. During the acquisition process, deep learning-based object detection algorithms, such as the improved CRAFT (Character Region Awareness for Text Detection) algorithm, are employed for text detection. Each image in the image set to be identified is automatically analyzed to obtain the undetermined regions, which include, but are not limited to, text regions, seal regions, etc. Feature region marker boxes are precisely added to these undetermined regions to obtain the marked regions to be identified.

[0048] It should be noted that the specific marking method can use hyperlink markers to mark the location coordinates and image number of the area to be marked. Clicking the hyperlink marker will display the corresponding feature area marking box.

[0049] S200: Obtain the image to be identified based on the region to be identified, and input the image to be identified into the preset training model in sequence to obtain the set of features to be processed corresponding to the image to be identified.

[0050] The image to be recognized represents the image that requires feature extraction, and the image to be recognized is the image from which the marked region to be recognized is extracted. The image to be recognized can be directly used for subsequent processing. The preset training model is the model for feature extraction of the image to be recognized. The set of features to be processed includes at least one feature to be processed. The feature to be processed represents the feature obtained by recognizing the image to be recognized. The feature to be processed can be text features and seal image features.

[0051] It should be noted that the preset training model is a pre-trained deep neural network model that combines the advantages of convolutional neural networks (CNN) and recurrent neural networks (RNN), and uses the connectionist temporal classification (CTC) algorithm for transcription to generate a comprehensive and accurate set of features to be processed.

[0052] For example, the extraction of features from printed text involves alternating operations using multiple layers of CNN convolutional and pooling layers to extract features such as font, font size, and layout. For instance, convolutional kernels of different scales are used to extract features from font strokes, and pooling operations are used to reduce feature dimensionality while preserving key feature information. The layout format of the text is determined by analyzing the statistical features of character spacing and line spacing.

[0053] For example, the extraction of features from a seal image can be achieved by using convolutional operations in a CNN, combined with color space transformation and texture analysis algorithms, to extract features such as the seal's shape, color, and texture. For instance, edge detection algorithms can be used to identify the seal's outline shape, color features can be extracted through color histogram analysis, and texture features can be extracted using methods such as gray-level co-occurrence matrix.

[0054] For example, in the extraction of handwritten text features, CNN is used to extract local features such as stroke length, angle, and curvature. Then, a bidirectional long short-term memory network (Bi-LSTM) is used to model the stroke sequence, learn writing style features, and fully capture the time series information and contextual dependencies in the handwriting process.

[0055] The process of obtaining the image to be identified based on the marked region includes the following steps:

[0056] S210, Match the corresponding preprocessing model in the preprocessing database based on the region to be identified.

[0057] S220 generates a set of images to be recognized based on a preprocessing model.

[0058] The preprocessing database stores several sets of preprocessing models, which are used to clarify the regions to be identified. The set of images to be identified includes at least one image that can be used for subsequent recognition. Since some images may be unclear during scanning, image recognition and classification techniques are used to automatically assign different preprocessing models based on the characteristics of different sets of target images. By analyzing the image features corresponding to the regions to be identified, such as color and texture, the image type is determined, and a suitable preprocessing model is selected.

[0059] It's important to note that these image preprocessing models comprehensively utilize techniques such as image denoising, background removal, and occlusion removal. For example, for printed text images, Gaussian filtering and other denoising techniques are used to remove noise. For images with security seals, image segmentation techniques are used to separate the seal from the background before background removal. For handwritten text images, morphological processing methods are used to remove occlusions. For seal images, a specialized seal contour enhancement algorithm is employed, using edge detection and contour extraction techniques to highlight the seal's outline while removing background interference.

[0060] S300, classify and identify the features to be processed in the feature set to be processed in turn to obtain the label information corresponding to the features to be processed.

[0061] The labeling information includes result type, features to be identified, and corresponding location information. The features to be processed in the feature set are classified and identified, specifically based on a Bidirectional Long Short-Term Memory (Bi-LSTM) network model, combined with attention mechanisms and transfer learning techniques, to efficiently and accurately classify the feature set. This model is pre-trained on a large-scale labeled dataset.

[0062] It should be noted that specific situations require specific analysis; therefore, the training data for this model can be adjusted according to the scheme of this embodiment to improve classification accuracy and generalization ability. The specific classification process is as follows, and the result types include secret, confidential, top secret, and unclassified types. The model automatically determines the result type of the feature to be identified by learning the features to be identified in the feature set to be processed.

[0063] S400: Based on the result type, classify the image to be identified by security level to obtain a security-classified image; and manage and store the archive to be managed based on the security-classified image.

[0064] Among them, the security classification marked image represents an image with marked information. The archive management system can obtain the result type, features to be identified and corresponding location information of the image to be identified by scanning the marked information of the security classification marked image corresponding to the image to be identified. It can also quickly find the location of classified archives, which can facilitate the subsequent stripping and separate storage operations of archive management personnel.

[0065] It should be noted that the classification of the image to be recognized based on the result type is achieved by using image annotation technology based on the set of features to be processed and the classification criteria corresponding to the model. Specifically, hyperlink markers can be used to mark the features to be recognized and the result type corresponding to the image. Clicking on a hyperlink marker will display the corresponding marking information.

[0066] Referring to Figure 5, in one embodiment, the result type includes at least one piece of data. The process involves sequentially classifying and identifying the features to be processed in the feature set to obtain the label information corresponding to the features to be processed, including the following steps:

[0067] S310, Determine whether the feature to be processed has labeling information.

[0068] S320, if the feature to be processed does not have labeling information, then key feature comparisons are performed on the features to be processed in turn to obtain the corresponding category data set.

[0069] S330, determine whether the category data in the category data set are the same.

[0070] S340, if the category data in the category data set are the same, then one of the category data will be used as the result type corresponding to the image to be identified.

[0071] S350, if the category data in the category data set are different, then obtain the label category based on the category data set and the preset label category, and use the label category as the result type of the image to be recognized.

[0072] The category dataset includes at least one category, and there is a one-to-one correspondence between the category data and the feature to be processed. The category data represents the result type corresponding to the feature to be processed. Preset label categories include top secret, confidential, secret, and unclassified types. The label category represents the type that is matched first when matching within the preset label categories.

[0073] Specifically, the process of obtaining a tag category based on the category data set and preset tag categories includes the following steps: Data of the preset tag type is obtained sequentially and used as a comparison category. This comparison category is then matched against the category data set. If a match is found, the comparison category is adopted as the tag category, and the matching process ends. If no match is found, the next data of the preset tag type is obtained sequentially until all data of the preset tag type is matched.

[0074] It should be noted that if the feature to be processed has labeling information, the labeling information is identified, and the feature to be identified and the corresponding location information of the labeling information are obtained and identified. The feature to be identified obtained above is skipped in the feature set to be processed, thereby reducing the identification of the entire image to be identified.

[0075] It should be noted here that, for the extraction of marked information, for the areas where the "secret level" and "confidentiality level" marking information are located on the paper documents, we first use an attention-based object detection model to accurately locate them, and then use CNN to extract features such as the shape and size of the characters in the text, as well as format features such as the text alignment and font color.

[0076] In one embodiment, key feature comparisons are performed sequentially on the features to be processed to obtain the corresponding category data set, including the following steps:

[0077] S321, determine whether the feature to be processed is a processable feature.

[0078] S322, If the feature to be processed is a processable feature, then perform key feature comparison based on the feature to be processed to obtain the category data corresponding to the feature to be processed.

[0079] S323, if the feature to be processed is not a processable type, then the feature to be processed is segmented to obtain the corresponding branch features, and key feature comparison is performed on the branch features respectively to obtain the corresponding branch result type.

[0080] S324, mark the branch result type according to the preset security level to obtain the category data corresponding to the feature to be processed.

[0081] Specifically, the processable feature representation is a feature that can be compared with key features. Determining whether the feature to be processed is a processable feature includes the following steps:

[0082] S321-1, Obtain feature identifiers based on the features to be processed, and determine whether the feature identifiers are unique.

[0083] S321-2, If the feature identifier is unique, then the feature to be processed is determined to be a processable feature.

[0084] S321-3 If the feature identifier is not unique, then the feature to be processed is determined to be a non-processable feature.

[0085] Among them, the feature identifier represents the type of the feature to be processed. The feature identifier can be a text identifier or a seal image, etc.

[0086] In steps S323-S324, if the feature to be processed is not a processable type, the feature to be processed is segmented to obtain the corresponding branch features, and key feature comparisons are performed on each branch feature to obtain the corresponding branch result type. The branch result types are marked according to a preset security level to obtain the category data corresponding to the feature to be processed.

[0087] For cases where the same feature contains multiple identifiers, such as printed text and seals, and involves different security classifications, a multi-branch network structure is adopted to extract and classify features independently for different types of feature regions. Then, a fusion strategy is used to comprehensively analyze the classification results from each branch and process them according to their respective classification attributes. Bounding boxes only label the feature regions of their current classification attribute, and hyperlink markers only label the image features and classification criteria of their respective classification attributes, ensuring the accuracy and interpretability of the classification results.

[0088] This embodiment determines the type of feature identifier. When the feature identifier only contains printed text, it can be directly detected and recognized, automatically identifying the classification type of the feature to be processed. For example, it automatically acquires the marking information of the feature to be processed. When an image file with a classification seal is detected and recognized, the classification type of the feature to be processed can be automatically identified. When the feature identifier is handwritten text, handwritten text image files in digitized outputs can be detected and recognized, automatically identifying the classification type of the feature to be processed, and assisting in the removal of classified documents. In addition, for some paper archives containing classified information, the classified information includes identifiers such as secret level and confidentiality level, with confidentiality level including confidential and top secret. When the identified classified information is secret level, the paper archive is determined to be unclassified; when the identified classified information is both secret level and confidentiality level, the paper archive is determined to be classified.

[0089] It should be noted that the specific nature of classified information can be determined based on the level of confidentiality, which will not be elaborated on here.

[0090] In one embodiment, after obtaining the feature identifier based on the feature to be processed, the following steps are further included:

[0091] S410, determine whether the feature identifier corresponding to the feature to be processed is a text identifier.

[0092] S420, if the feature identifier is a text identifier, then based on the feature to be processed, several feature words to be identified are obtained using feature recognition technology.

[0093] S430: Filter the feature words to be identified in the preset database and determine whether the corresponding comparison feature words can be obtained in the preset database.

[0094] S440, if the corresponding comparison feature words can be obtained in the preset database, the feature words to be identified are filtered in the feature set to be processed to obtain the filtering position.

[0095] S450: Obtain the corresponding comparison results based on the comparison feature words, and label the features to be processed based on the screening positions to obtain labeling information.

[0096] Among them, the feature words to be identified represent the feature words obtained based on the features to be processed. These feature words are key feature words. Several sets of feature words are stored in the preset database. Each feature word has a corresponding level of density. The screening position represents the position of the feature words in the image to be identified.

[0097] Specifically, the feature words to be identified are filtered in the feature set to be processed to obtain the filtering positions with the same density level, and the feature words to be identified at the filtering positions are marked according to the comparison results to obtain the marking information.

[0098] The implementation principle is as follows:

[0099] A set of images to be identified corresponding to the files to be managed is obtained from a pre-set digital database. These images are then divided into regions and labeled to obtain labeled regions. Images to be identified are obtained based on these labeled regions and sequentially input into a pre-set training model to obtain a set of features to be processed. These features are then classified and identified sequentially to obtain corresponding labeling information. Based on the result type, the images are classified into security levels to obtain security-marked images. The files to be managed are then managed and stored based on these security-marked images.

[0100] This application also discloses an image recognition-based file management system.

[0101] As shown in Figure 6, the image recognition-based file management system includes an image acquisition module 10, a feature extraction module 20, a classification and identification module 30, and a label generation module 40. The image acquisition module 10 acquires a set of images to be identified corresponding to the files to be managed based on a preset digital database. It then divides and labels the images in the set to be identified into regions to obtain the labeled regions. The feature extraction module 20 acquires the images to be identified based on the labeled regions and sequentially inputs these images into a preset training model to obtain a set of features to be processed corresponding to each image. The set of features to be processed includes at least one feature to be processed. The classification and identification module 30 sequentially classifies and identifies the features to be processed in the set of features to be processed to obtain label information corresponding to the images to be identified. The label information includes the result type. The label generation module 40 assigns a security classification to the images to be identified based on the result type to obtain a security-marked image. The files to be managed are then managed and stored based on the security-marked image.

[0102] The other functions performed in the image acquisition module 10, feature extraction module 20, classification and identification module 30, and tag generation module 40, as well as the technical details of each function, are the same as or similar to the corresponding features in the image recognition file management method described above, so they will not be repeated here.

[0103] The security classification markers include hyperlinks, and the marker information also includes image features and location information. The image recognition file management system also includes a result output module 50 and an image preprocessing module. The result output module 50 displays the result type, image features, and location information in a list format on the screen for staff to view. The image preprocessing module assigns appropriate image preprocessing models according to different image types to perform denoising, background removal, and occlusion removal in feature areas, thereby improving image quality.

[0104] Specifically, the image acquisition module 10 is responsible for acquiring images from the digitized paper archives according to set rules, corresponding to the function of calling the image acquisition script, ensuring that the acquired images are complete and meet the requirements of subsequent processing. The feature extraction module 20 is built based on CRNN and is responsible for extracting features of various types of images to form an image feature set. The classification and identification module 30 is equipped with a Bi-LSTM model to classify and identify images based on the image feature set and determine their security level. The tag generation module 40 completes image tagging, generates hyperlink markers, and outputs a report containing auxiliary stripping information based on the results output module 50, facilitating archive management operations.

[0105] It's important to note that CRNN is a deep learning technique that combines Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs). The CNN part utilizes the sliding kernels of convolutional layers across the data space to efficiently extract local features from the input data. The input data can be two-dimensional images or multi-dimensional signals; stacking layers allows for the extraction of multi-scale features, from basic edges to complex structures. Pooling layers downsample the convolution results, reducing data dimensionality, preserving key features, and enhancing the model's robustness to data transformations. The RNN part, especially when using Long Short-Term Memory (LSTM) networks or Gated Recurrent Units (GRUs), effectively handles temporal dependencies through gating mechanisms. At each time step, the hidden state is updated based on the current input and the hidden state from the previous time step, achieving continuous memorization and transmission of sequence information. Bi-LSTM, or Bidirectional Long Short-Term Memory network, is a deep learning model that excels in sequence data processing, effectively capturing both forward and backward contextual information.

[0106] This application also discloses an electronic device.

[0107] Referring to FIG7, the electronic device includes a processor 61 and a memory 62 coupled to each other, wherein the memory 62 stores a computer program 63 that can run on the processor 61.

[0108] When computer program 63 is executed by processor 61, it implements an image recognition file management method.

[0109] The memory 62 may be a ROM or other type of static storage device capable of storing static information and instructions, a random access memory 62, or other type of dynamic storage device capable of storing information and instructions. It may also be an electrically erasable programmable read-only memory 62, a read-only optical disc or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), a magnetic disk storage medium, or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. In some embodiments, the memory 62 may be an internal storage unit.

[0110] Processor 61 can be a central processing unit 61, a general-purpose processor 61, a data signal processor 61, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It is used to run program code stored in memory 62 or process data.

[0111] Processor 61 and memory 62 are connected via a bus. The bus may include a pathway for transmitting information between the aforementioned components. The bus may be a peripheral interconnect standard bus or an extended industry standard structure bus, etc. Buses can be classified as address buses, data buses, control buses, etc. For ease of illustration, only one thick line is used to represent it in Figure 7, but this does not indicate that there is only one bus or one type of bus.

[0112] Figure 7 only shows an electronic device with memory 62, processor 61, and bus. Those skilled in the art will understand that the structure shown in Figure 7 does not constitute a limitation on the electronic device; it can be a bus-type structure or a star-type structure. The electronic device may also include more or fewer components than shown, or combine certain components, or deploy different components. Other existing or future electronic devices are applicable and should be included within the scope of protection, and are incorporated herein by reference.

[0113] It should be understood that although the steps in the flowcharts in the accompanying drawings are shown sequentially as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise expressly stated herein, there is no strict order in which these steps are performed, and they may be performed in other orders.

[0114] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.

Claims

1. A file management method based on image recognition, comprising: Based on a preset digital database, obtain a set of images to be identified corresponding to the archives to be managed. Divide and mark the images to be identified in the set of images to be identified to obtain the marked regions to be identified. The image to be identified is obtained based on the region to be identified, and the image to be identified is sequentially input into a preset training model to obtain a set of features to be processed corresponding to the image to be identified. The set of features to be processed includes at least one feature to be processed. The features to be processed in the set of features to be processed are classified and identified sequentially to obtain the labeling information corresponding to the features to be processed, and the labeling information includes the result type; The image to be identified is classified based on the result type to obtain a classified image, and the file to be managed is managed and stored based on the classified image.

2. The image recognition-based file management method according to claim 1, characterized in that, The result type includes at least one piece of data. The step of sequentially classifying and identifying the features to be processed in the set of features to be processed to obtain the label information corresponding to the features to be processed includes: Determine whether the feature to be processed contains any marking information; If the feature to be processed does not have labeling information, then key feature comparison is performed on the feature to be processed in sequence to obtain the corresponding category data set. The category data set includes at least one category data, and the category data corresponds one-to-one with the feature to be processed. Determine whether the category data in the category data set are the same; If the category data in the category data set are the same, then one of the category data items will be used as the result type corresponding to the image to be identified; If the category data in the category data set are different, then the marker category is obtained based on the category data set and the preset marker category, and the marker category is used as the result type of the image to be identified.

3. The image recognition-based file management method according to claim 2, characterized in that, The step of sequentially performing key feature comparisons on the features to be processed to obtain the corresponding category data set includes: Determine whether the feature to be processed is a processable feature; If the feature to be processed is a processable feature, then key feature comparison is performed based on the feature to be processed to obtain the category data corresponding to the feature to be processed; If the feature to be processed is not a processable type, the feature to be processed is segmented to obtain the corresponding branch features, and key feature comparison is performed on the branch features to obtain the corresponding branch result type. The branch result type is marked according to a preset security level to obtain the category data corresponding to the feature to be processed.

4. The image recognition-based file management method according to claim 3, characterized in that, The step of determining whether the feature to be processed is a processable feature includes: Based on the feature to be processed, a feature identifier is obtained, and it is determined whether the feature identifier is unique, wherein the feature identifier represents the type of the feature to be processed; If the feature identifier is unique, then the feature to be processed is determined to be a processable feature; If the feature identifier is not unique, then the feature to be processed is determined to be a non-processable feature.

5. The image recognition-based file management method according to claim 1, characterized in that, The step of dividing and marking the images to be identified in the set of images to be identified, in order to obtain the marked regions to be identified, includes: The image to be identified is scanned sequentially according to a preset interval and a preset scanning route to obtain several sets of scanning regions. The sets of scanning regions include the scanning regions and the data types corresponding to the scanning regions. The scanned regions that are of the same data type and are continuous are taken as regions to be determined, and the regions to be determined are marked to obtain the regions to be identified.

6. The image recognition-based file management method according to claim 1, characterized in that, The step of obtaining the image to be identified based on the identified marker region includes: The corresponding preprocessing model is matched in the preprocessing database based on the region to be identified. The preprocessing database stores several sets of preprocessing models, which are used to perform clear processing on the region to be identified. A set of images to be identified is generated based on the preprocessing model, and the set of images to be identified includes at least one image to be identified.

7. The image recognition-based file management method according to claim 4, characterized in that, After obtaining the feature identifier based on the feature to be processed, the method further includes: Determine whether the feature identifier corresponding to the feature to be processed is a text identifier; If the feature identifier is a text identifier, then based on the feature to be processed, several feature words to be identified are obtained using feature recognition technology; The feature words to be identified are filtered in a preset database, and it is determined whether the corresponding comparison feature words can be obtained in the preset database; If a corresponding matching feature word can be obtained from the preset database, the feature word to be identified is filtered in the feature set to be processed to obtain the filtering position, and the filtering position represents the position of the matching feature word in the image to be identified; The corresponding comparison results are obtained based on the comparison feature words, and the comparison results are used to mark the features to be processed based on the filtering positions to obtain marking information.

8. An image recognition-based file management system, comprising: Image acquisition module (10): The image acquisition module (10) acquires a set of images to be identified corresponding to the archives to be managed based on a preset digital database, divides and marks the images to be identified in the set of images to be identified, so as to obtain the marked areas to be identified. The feature extraction module (20) obtains the image to be identified based on the region to be identified and sequentially inputs the image to be identified into a preset training model to obtain the set of features to be processed corresponding to the image to be identified. The set of features to be processed includes at least one feature to be processed. The classification and identification module (30) sequentially classifies and identifies the features to be processed in the set of features to be processed in order to obtain the label information corresponding to the image to be identified, and the label information includes the result type; The tag generation module (40) performs security level tagging on the image to be identified based on the result type to obtain a security level tag image, and manages and stores the file to be managed based on the security level tag image.

9. The image recognition-based file management system according to claim 8, characterized in that, The security level marker includes a hyperlink marker, and the marker information also includes image features and location information. It also includes a result output module (50), which is used to display the result type, image features and location information in a list form on the display screen.

10. An electronic device comprising a processor (61) and a memory (62) coupled to each other, wherein the memory (62) stores a computer program (63) capable of running on the processor (61); When the computer program (63) is executed by the processor (61), it implements the image recognition file management method as described in any one of claims 1-7.