File retrieval method and system based on deep learning and storage medium
By identifying and extracting text and key fragments from archive images, and using a pre-trained multimodal archive search model for search, the problem of insufficient comprehensive and accurate archive search in the existing technology is solved, and efficient and accurate archive search and personalized reading recommendations are achieved.
Patent Information
- Application Number
- CN202510595542.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-06
AI Technical Summary
The existing deep learning-based archive retrieval methods ignore the multi-dimensional characteristics of archive information, resulting in incomplete and accurate retrieval, making it difficult for users to quickly find the content they need, affecting user experience and retrieval efficiency.
By obtaining archival image data, identifying text information to form electronic text data, extracting key fragment text data, and inputting it into the pre-trained archival search model for multimodal search, outputting matching search sets and correlations, generating reading recommendation guides, and updating the archival search system.
A multimodal and complete archive knowledge system has been realized, retrieval efficiency and accuracy have been improved, and a personalized reading path is provided for users, which has improved user experience and reduced information overload and reading confusion.
Smart Images

Figure CN120104844A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of archive retrieval, deep learning, and image processing technology, and in particular to an archive retrieval method, system, and storage medium based on deep learning. Background Art
[0002] With the rapid development of information technology, the electronic and digitalization of archives has become a general trend. However, when faced with massive and diverse archival data, traditional archive retrieval mainly relies on paper archives and manual classification retrieval, which often has problems such as low efficiency, non-intuitive information display, and poor user experience. With the popularization of electronic text materials, how to efficiently retrieve and use these materials has become an urgent problem to be solved. Although existing technologies have also tried to introduce digital means, they often lack intelligent retrieval and reading recommendation functions and cannot provide users with personalized reading guides.
[0003] The application of deep learning technology in the fields of image processing and natural language processing has achieved remarkable results, providing strong support for the automation and intelligence of archive retrieval. However, the existing archive retrieval methods based on deep learning still have some obvious shortcomings. For example, an electronic archive management method recorded in the patent with publication number CN112329669B discloses that the scanned image is preprocessed, cut, matched with text, keyword recognized and classified, and a training set is constructed by marking the keywords and their effective features in the training text, and then the training set is used to train the neural network to obtain the keywords of the document, and the classification of the document image to be identified is obtained according to the keywords, and the classification information is marked on the document image to be identified. However, it mainly focuses on the text recognition of archive images, but ignores the multi-dimensional characteristics of archive information, that is, archives not only contain text information, but also may contain audio, video and other types of information. This single-dimensional processing method limits the comprehensiveness and accuracy of archive information retrieval. Users often find it difficult to quickly find the required content from a large amount of archive information, which not only affects the user experience, but also limits the retrieval efficiency of the archive retrieval system.
[0004] This application is directed to establishing an archive retrieval method, system and storage medium based on deep learning to solve the above-mentioned problems. Summary of the invention
[0005] In order to achieve the above-mentioned purpose and other advantages according to the present invention, the first purpose of the present invention is to provide an archive retrieval method based on deep learning, comprising the following steps: Obtaining archival image data to be retrieved; Recognizing the text information in the archive image data to form electronic text data; Extracting key segments from the electronic text data to obtain key segment text data; Input the key segment text data into a pre-trained archive retrieval model for retrieval, and output a retrieval set matching the key segment and its relevance; Wherein, the search set is an electronic text data document, an audio data document, and a video data document that matches the key segment text data; generating a reading recommendation guide for the search set according to the relevance; Wherein, the reading recommendation guide at least includes the relevance and the recommended reading order; The reading recommendation guide is associated with the archive image data, and updated and stored in the archive retrieval system.
[0006] Furthermore, the relevance is a relevance threshold outputted from matching the search set with the key segment text data.
[0007] Furthermore, the correlation degree threshold calculation method includes: The relevance threshold between the key segment text data and the search set is calculated by cosine similarity.
[0008] Furthermore, the step of generating a reading recommendation guide for the search set according to the relevance includes: sorting the search set according to the relevance; Generate a reading recommendation guide based on the sorting results.
[0009] Furthermore, the step of sorting the search set according to the relevance includes: Obtaining the relevance of the search set to the key segment text data; Determining whether the relevances in the search set belong to the same threshold range; If the correlations belong to the same threshold range, the video data files, the audio data files, and the electronic text data files are sorted in the order of one another; If the relevance does not belong to the same threshold range, the video data document, the audio data document, and the electronic text data document are recommended in descending order of the relevance.
[0010] Furthermore, the step of identifying the text information in the archive image data to form electronic text data includes: Preprocessing the archival image data to obtain preprocessed archival image data; Converting the pre-processed archival image data into machine-readable character encoding data using OCR technology; The character encoding data is subjected to data recognition processing to form electronic text data.
[0011] Furthermore, the step of extracting key segments from the electronic text data includes: removing stop words from the electronic text; Performing word segmentation on the electronic text to split it into individual words / phrases; Performing part-of-speech tagging on the words / phrases; Calculating the TF-IDF value of the word / phrase to count the distribution frequency of the word / phrase in the electronic text; Convert the words / phrases into word vectors to capture the semantic relationship between the words / phrases; Input the word vector into a classifier for binary classification to distinguish key segments from non-key segments; The key segments are subjected to syntactic analysis, adjacent key segments are merged, and repeated key segments are removed to obtain key segment text data.
[0012] Furthermore, the method further includes the step of constructing the archive retrieval model: Acquire electronic text data documents, audio data documents, and video data documents in the archive retrieval database to form an archive data set; Performing language processing on the electronic text data document to extract key features of the text; Building a text processing model using deep learning technology based on key features of the text to understand and analyze the semantics and content in the electronic text data document; Extracting key features of the audio data document using an audio processing algorithm; Building an audio processing model based on key features of the audio data document, converting the audio data document into text to extract and identify semantics and content in the audio data document; Extracting key frame images and key features from the video data document using image processing technology; Building a video processing model using deep learning technology based on key frame images and key features in the video data document, converting the video data document into text to extract and identify semantics and content in the video data document; Through multimodal learning technology, the outputs of the text processing model, audio processing model, and video processing model are integrated to form a unified archive retrieval model.
[0013] The second object of the present invention is to provide an archive retrieval system based on deep learning, comprising the following modules: An image acquisition module, used to acquire archive image data to be retrieved; A text recognition module, used to recognize text information in the archive image data to form electronic text data; A key segment extraction module is used to extract key segments from the electronic text data to obtain key segment text data; An archive retrieval module, used to input the key segment text data into a pre-trained archive retrieval model for retrieval, and output a retrieval set matching the key segment and a relevance; Wherein, the search set is an electronic text data document, an audio data document, and a video data document that matches the key segment text data; A reading guide generating module, used for generating a reading recommendation guide for the search set according to the relevance; wherein the reading recommendation guide at least includes the relevance and the recommended reading order; The updating module is used to associate the reading recommendation guide with the archive image data, and update and store it in the archive retrieval system.
[0014] The third object of the present invention is to provide a readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, an archive retrieval method based on deep learning is implemented.
[0015] Compared with the prior art, the present invention has the following beneficial effects: The present invention relates to an archive retrieval method, system and storage medium based on deep learning. First, the archive image data is converted into electronic text using image recognition technology, and then the key fragments are extracted and deep retrieval is performed using a pre-trained archive retrieval model, which can accurately match various types of documents related to the key fragments (including electronic text, audio, video, etc.), forming a multimodal and complete archive knowledge system, and improving the efficiency of retrieval. According to the relevance of the retrieval results, a reading recommendation guide containing a recommended reading order is generated, which provides users with a personalized reading path, helps users find the required information faster, enables users to read in a logical order and priority, improves reading efficiency, and avoids the problem of information overload and reading confusion for users. Then, the reading recommendation guide is associated with the archive image data, and updated and stored in the archive retrieval system, realizing the intelligent retrieval of archives, promoting the sharing of archive information, and providing strong support for subsequent archive research, analysis and utilization.
[0016] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention and implement it according to the contents of the specification, the following is a detailed description of the preferred embodiments of the present invention in conjunction with the accompanying drawings. The specific implementation of the present invention is given in detail by the following embodiments and their accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 This is a flow chart of the deep learning-based archive retrieval method of this application; Figure 2 The flowchart of recognizing text information in archive image data and forming electronic text data as described in Example 1; Figure 3 This is a flow chart of extracting key segments from electronic text data as described in Example 1; Figure 4 A flow chart of generating a reading recommendation guide for the search set according to relevance as described in Example 1; Figure 5 This is a flow chart of sorting the search set according to relevance described in Example 1; Figure 6 The flowchart of constructing the archive retrieval model described in Example 1; Figure 7 This is a schematic diagram of the deep learning-based archive retrieval method described in Example 1; Figure 8 Schematic diagram of the archive retrieval system based on deep learning in Example 2; Fig. 9 This is a schematic diagram of a computer-readable storage medium in Example 3. DETAILED DESCRIPTION
[0018] The present invention is further described below in conjunction with the accompanying drawings and specific implementation methods. It should be noted that, under the premise of no conflict, the various embodiments or technical features described below can be arbitrarily combined to form a new embodiment.
[0019] In the subsequent description, the suffixes such as "module", "component" or "unit" used to represent elements are only used to facilitate the description of the present invention, and have no specific meanings. Therefore, "module", "component" or "unit" can be used in a mixed manner. Example 1
[0020] The present invention provides a deep learning-based archive retrieval method, such as Figure 1 , Figure 7 As shown, the specific steps include: S101, obtaining archive image data to be retrieved; S102, identifying the text information in the archive image data to form electronic text data; S103, extracting key segments from the electronic text data to obtain key segment text data; S104, inputting the key segment text data into a pre-trained archive retrieval model for retrieval, and outputting a retrieval set matching the key segment and its relevance; Wherein, the search set is an electronic text data document, an audio data document, and a video data document that matches the key segment text data; S105, generating a reading recommendation guide for the search set according to the relevance; Wherein, the reading recommendation guide at least includes the relevance and the recommended reading order; S106, associating the reading recommendation guide with the archive image data, and updating and storing the updated data in the archive retrieval system.
[0021] In some embodiments, the archival image data described in step S101 is derived from paper archives, screenshots or photos, databases or network retrieval, etc., and is obtained by real-time shooting, scanning or uploading.
[0022] Specifically, it should be understood that when the archival image data originates from paper archives, the paper archives are scanned and digitized through a high-precision scanner or high-definition camera equipment, and the paper archive data is converted into archival image data to generate high-quality image files to ensure the clarity and readability of the text information.
[0023] When the archival image data is derived from photos or screenshots of historical photos, meeting records, news reports, etc., the archival image data is imported through an upload function.
[0024] When the archival image data is sourced from an existing digital archive of an institution such as a library or museum or retrieved from a search engine network, the archival image data is imported through an upload function.
[0025] In a preferred embodiment, the archival image data can also be obtained by manual input or upload by the user.
[0026] In some embodiments, the text information in the archive image data is recognized in step S102 to form electronic text data, such as Figure 2 As shown, the steps include: S1021, preprocessing the archive image data to obtain preprocessed archive image data; S1022, using OCR technology to convert the pre-processed archive image data into machine-readable character encoding data; S1023, performing data recognition processing on the character encoding data to form electronic text data.
[0027] During the scanning or uploading process, the archival image data is affected by the device pixel accuracy, lighting conditions, environmental factors, etc., and is prone to image interference, such as noise, models, background clutter, etc., which affects the accuracy and precision of recognition.
[0028] In a preferred embodiment, the preprocessing of the archival image data in step S1021 includes denoising, binarization, and image enhancement of the archival image data, which effectively removes noise, blur, background clutter, and other problems in the image, making the text information in the image clearer, thereby improving the recognition accuracy of the OCR technology.
[0029] In a preferred embodiment, the data recognition processing of the character encoding data in step S1023 includes character segmentation, feature extraction, character recognition, and text correction processing of the character encoding data, so that the OCR technology can more efficiently recognize the text information in the image, reduce unnecessary calculations and resource consumption, ensure the accuracy and completeness of the recognition results, and reduce misrecognition and missed recognition.
[0030] In some embodiments, the key segments of the electronic text data are extracted in step S103, such as Figure 3 As shown, the steps include: S1031, removing stop words in the electronic text; S1032, performing word segmentation processing on the electronic text to split it into individual words / phrases; S1033, performing part-of-speech tagging on the words / phrases; S1034, calculating the TF-IDF value of the word / phrase to count the distribution frequency of the word / phrase in the electronic text; S1035, converting the words / phrases into word vectors to capture the semantic relationship between the words / phrases; S1036, inputting the word vector into a classifier for binary classification to distinguish key segments from non-key segments; S1037, performing syntactic analysis on the key segments, merging adjacent key segments and removing duplicate key segments to obtain key segment text data.
[0031] Key fragments usually contain important information and core ideas in the text, so extracting key fragments helps improve the accuracy of retrieval. Users can locate the required information more accurately, reducing false detection and missed detection. By extracting key fragments, this application reduces the amount of text to be retrieved, allowing users to find the required information in a shorter time, thereby improving the retrieval speed.
[0032] In some embodiments, the relevance in step S104 is a relevance threshold outputted by matching the search set with the key segment text data.
[0033] In a preferred embodiment, the correlation threshold calculation method includes: The relevance threshold between the key segment text data and the search set is calculated by cosine similarity.
[0034] In some embodiments, the step S105 generates a reading recommendation guide for the search set according to the relevance, such as Figure 4 As shown, the steps include: S1051, sorting the search set according to the relevance; S1502, generating a reading recommendation guide based on the sorting results.
[0035] In a preferred embodiment, the search set is sorted according to the relevance in step S1051, such as Figure 5 As shown, the steps include: S10511, obtaining the relevance of the search set and the key segment text data; S10512, determining whether the relevances in the search set belong to the same threshold range; S10513a, if the correlations belong to the same threshold range, sorting the video data documents, the audio data documents, and the electronic text data documents in the order of one another; S10513b: If the relevance does not belong to the same threshold range, the video data document, the audio data document, and the electronic text data document are recommended in descending order of the relevance.
[0036] In a preferred embodiment, the threshold range specifically includes: When the correlation threshold is 85-100, it is the best matching range; When the correlation threshold is 70-84, it is an excellent matching range; When the correlation threshold is 55-69, it is a qualified matching range; When the correlation degree threshold is lower than 54, it is a low matching range.
[0037] For example, the present application uses the case where the correlations do not belong to the same threshold range as an example, which is as follows: The key segment text data extracted is “Silk Road”; Inputting the key segment text data "Silk Road" into the pre-trained archive retrieval model for retrieval; The search set and relevance output by the model are as follows: Electronic text data document: "Silk Road Historical Documents", the relevance threshold is 85.
[0038] Audio data file: Recording of a lecture on the Ancient Silk Road and Eurasian Civilization, with a relevance threshold of 69.
[0039] Video data file: Silk Road documentary clip, relevance threshold is 78.
[0040] Since the relevance thresholds of the three audio data documents are not in the same threshold range, the generated reading recommendation guidelines are as follows: It is recommended that users first watch the electronic text data document with the highest relevance threshold, and then watch the video data document with the second highest relevance threshold, so as to obtain the most intuitive reproduction of the Silk Road scene; finally, users can choose whether to listen to the audio data document as supplementary archival material, listen to the recordings of lectures on the Ancient Silk Road and Eurasian Civilization, and enhance their in-depth understanding of the Silk Road lectures.
[0041] For example, the present application uses the case where the correlations belong to the same threshold range as an example, which is as follows: The extracted key fragment text is “AI technology development trend”; Input the key fragment text “AI technology development trend” into the pre-trained archive retrieval model for retrieval; The search set and relevance output by the model are as follows: Electronic text data document: "Analysis of the latest progress and trends in AI technology", with a relevance threshold of 85; Audio data files: recordings of interviews with AI technology experts, with a relevance threshold of 87; Video data file: video showing examples of AI technology applications, with a relevance threshold of 86; Since the relevance thresholds of the three audio data documents are in the same threshold range, the generated reading recommendation guidelines are as follows: It is recommended that users watch the video data document first, then the audio data document, and finally the electronic text data document.
[0042] In some embodiments, the reading recommendation guide described in step S105 also includes the name and type of the search set.
[0043] In some embodiments, the reading recommendation guide is associated with the archival image data in step S106. Specifically, the association can be performed by adding a hyperlink, generating an identifier, generating an annotation, a reading note or a reading introduction.
[0044] In a preferred embodiment, the present application generates a unique identifier for each copy of the archival image data, and adds the unique identifier corresponding to the archival image data to the part of the reading recommendation guide related to the archival image data. When the user views the reading recommendation guide, the corresponding archival image data can be found through the unique identifier.
[0045] In some embodiments, Figure 6 As shown, the deep learning-based archive retrieval method further includes the step of constructing the archive retrieval model: S107, acquiring electronic text data documents, audio data documents, and video data documents in the archive retrieval database to form an archive data set; S108, performing language processing on the electronic text data document to extract key features of the text; S109, constructing a text processing model using deep learning technology based on key features of the text to understand and analyze the semantics and content in the electronic text data document; S110, extracting key features of the audio data document using an audio processing algorithm; S111, constructing an audio processing model based on key features of the audio data document, converting the audio data document into text, so as to extract and identify semantics and content in the audio data document; S112, extracting key frame images and key features in the video data document using image processing technology; S113, constructing a video processing model using deep learning technology based on key frame images and key features in the video data document, converting the video data document into text, so as to extract and identify semantics and content in the video data document; S114, through multimodal learning technology, the outputs of the text processing model, the audio processing model, and the video processing model are integrated to form a unified archive retrieval model.
[0046] This application integrates various types of archival data such as electronic text, audio and video into a unified archival retrieval model. First, various data documents are collected from the archival retrieval database to form a rich archival data set. Then, the electronic text data is language processed to extract the key features in the text, and a text processing model is constructed based on these features. This model can deeply understand and analyze the semantics and content in the electronic text data documents. At the same time, the audio processing algorithm and video processing technology are used to extract the key features of the audio and video data documents, and the audio processing model and the video processing model are constructed respectively, so as to extract and identify the semantics and content therein, providing a solid foundation for accurate retrieval. Finally, through multimodal learning technology, the outputs of the text processing model, the audio processing model and the video processing model are organically integrated to form a unified archival retrieval model. This model can fully understand and utilize various types of information in the archival data, provide users with more accurate and convenient retrieval services, significantly improve the efficiency and accuracy of archival retrieval, and reduce the labor cost of archival retrieval and management. Example 2
[0047] The present invention provides an archive retrieval system based on deep learning, such as Figure 8 As shown, it includes the following modules: An image acquisition module, used to acquire archive image data to be retrieved; A text recognition module, used to recognize text information in the archive image data to form electronic text data; A key segment extraction module is used to extract key segments from the electronic text data to obtain key segment text data; An archive retrieval module, used to input the key segment text data into a pre-trained archive retrieval model for retrieval, and output a retrieval set matching the key segment and a relevance; Wherein, the search set is an electronic text data document, an audio data document, and a video data document that matches the key segment text data; A reading guide generating module, used for generating a reading recommendation guide for the search set according to the relevance; wherein the reading recommendation guide at least includes the relevance and the recommended reading order; The updating module is used to associate the reading recommendation guide with the archive image data, and update and store it in the archive retrieval system.
[0048] The archive retrieval system involved in this application integrates image acquisition, text recognition, key fragment extraction, archive retrieval module, reading guide generation and system update module, realizing the intelligent recognition from efficient collection of archive image data to electronic text data. First, the system efficiently collects archive images to be retrieved through the image acquisition module. Then, the text information in these images is accurately converted into electronic text data using text recognition technology. On this basis, the key fragment extraction module further filters out the core information in the text to ensure the accuracy of the retrieval. Next, the system inputs these key fragments into the archive retrieval model that has been deeply learned and pre-trained. The model can cross multiple data formats such as electronic text, audio and video, perform efficient and comprehensive retrieval, quickly output a retrieval set that is highly matched with the key fragment, and provide a relevance score. In order to further enhance the user experience, the system is also equipped with a reading guide generation module, which intelligently arranges the recommended reading order based on the relevance of the retrieval set and generates a recommended reading guide. Finally, the update module closely links the reading guide with the original archival image data to achieve dynamic updating of the system and continuous optimization of personalized services. It not only significantly improves the efficiency and accuracy of archival retrieval and greatly reduces management costs, but also greatly enriches the user experience by providing intelligent and personalized retrieval and reading services. Example 3
[0049] The embodiment of the present invention also provides a computer readable storage medium. Fig. 9 As shown, program instructions are stored thereon, and when the program instructions are executed, the archive retrieval method based on deep learning as recorded in the above-mentioned embodiment 1 is implemented.
[0050] Among them, the program instructions are stored in a computer-readable storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.) or on a network, including a number of computer program instructions to enable a computing device (which can be a personal computer, server, or network device, etc.) to execute the above method according to the implementation mode of the present application.
[0051] Through the description of the above implementation modes, it is easy for those skilled in the art to understand that the example implementation modes described here can be implemented by software, or by software combined with necessary hardware. Although the implementation modes of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and implementation modes, and they can be fully applied to various fields suitable for the present invention. For those familiar with the art, other modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to the specific details and the legends shown and described here.
[0052] The apparatus, electronic device, non-volatile computer storage medium and method provided in the embodiments of this specification correspond to each other, and therefore, the apparatus, electronic device and non-volatile computer storage medium also have similar beneficial technical effects as the corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the corresponding apparatus, electronic device and non-volatile computer storage medium will not be repeated here.
[0053] Those skilled in the art also know that, in addition to implementing the controller in a purely computer-readable program code, the controller can be made to implement the same function in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered as a hardware component, and the devices for implementing various functions included therein can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules for implementing the method and structures within the hardware component.
[0054] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0055] For the convenience of description, the above devices are described in terms of functions and are divided into various units. Of course, when implementing one or more embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0056] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, the embodiments of this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware.
[0057] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0058] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0059] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0060] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0061] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0062] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0063] The specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media including storage devices.
[0064] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0065] The above description is only an embodiment of this specification and is not intended to limit one or more embodiments of this specification. For those skilled in the art, one or more embodiments of this specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this specification shall be included in the scope of the claims of one or more embodiments of this specification.
Claims
1. A deep learning-based archive retrieval method, characterized in that: The specific steps include: Obtaining archival image data to be retrieved; Recognizing the text information in the archive image data to form electronic text data; Extracting key segments from the electronic text data to obtain key segment text data; Input the key segment text data into a pre-trained archive retrieval model for retrieval, and output a retrieval set matching the key segment and its relevance; Wherein, the search set is an electronic text data document, an audio data document, and a video data document that matches the key segment text data; generating a reading recommendation guide for the search set according to the relevance; Wherein, the reading recommendation guide at least includes the relevance and the recommended reading order; The reading recommendation guide is associated with the archive image data, and updated and stored in the archive retrieval system.
2. The deep learning-based archive retrieval method according to claim 1, characterized in that: The relevance is a relevance threshold outputted when the search set matches the key segment text data.
3. The deep learning-based archive retrieval method according to claim 2, characterized in that: The correlation threshold calculation method comprises: The relevance threshold between the key segment text data and the search set is calculated by cosine similarity.
4. The deep learning-based archive retrieval method according to claim 1, characterized in that: The step of generating a reading recommendation guide for the search set according to the relevance comprises: sorting the search set according to the relevance; Generate a reading recommendation guide based on the sorting results.
5. The deep learning-based archive retrieval method according to claim 4, characterized in that: The step of sorting the search set according to the relevance comprises: Obtaining the relevance of the search set to the key segment text data; Determining whether the relevances in the search set belong to the same threshold range; If the correlations belong to the same threshold range, the video data files, the audio data files, and the electronic text data files are sorted in the order of one another; If the relevance does not belong to the same threshold range, the video data document, the audio data document, and the electronic text data document are recommended in descending order of the relevance.
6. The deep learning-based archive retrieval method according to claim 1, characterized in that: The step of identifying the text information in the archive image data to form electronic text data includes: Preprocessing the archival image data to obtain preprocessed archival image data; Converting the pre-processed archival image data into machine-readable character encoding data using OCR technology; The character encoding data is subjected to data recognition processing to form electronic text data.
7. The deep learning-based archive retrieval method according to claim 1, characterized in that: The step of extracting key segments from the electronic text data includes: removing stop words from the electronic text; Performing word segmentation on the electronic text to split it into individual words / phrases; Performing part-of-speech tagging on the words / phrases; Calculating the TF-IDF value of the word / phrase to count the distribution frequency of the word / phrase in the electronic text; Convert the words / phrases into word vectors to capture the semantic relationship between the words / phrases; Input the word vector into a classifier for binary classification to distinguish key segments from non-key segments; The key segments are subjected to syntactic analysis, adjacent key segments are merged, and repeated key segments are removed to obtain key segment text data.
8. The deep learning-based archive retrieval method according to claim 1, characterized in that: It also includes the steps of constructing the archive retrieval model: Acquire electronic text data documents, audio data documents, and video data documents in the archive retrieval database to form an archive data set; Performing language processing on the electronic text data document to extract key features of the text; Building a text processing model using deep learning technology based on key features of the text to understand and analyze the semantics and content in the electronic text data document; Extracting key features of the audio data document using an audio processing algorithm; Building an audio processing model based on key features of the audio data document, converting the audio data document into text to extract and identify semantics and content in the audio data document; Extracting key frame images and key features from the video data document using image processing technology; Building a video processing model using deep learning technology based on key frame images and key features in the video data document, converting the video data document into text to extract and identify semantics and content in the video data document; Through multimodal learning technology, the outputs of the text processing model, audio processing model, and video processing model are integrated to form a unified archive retrieval model.
9. A deep learning-based archive retrieval system, characterized in that: Includes the following modules: An image acquisition module, used to acquire archive image data to be retrieved; A text recognition module, used to recognize text information in the archive image data to form electronic text data; A key segment extraction module is used to extract key segments from the electronic text data to obtain key segment text data; An archive retrieval module, used to input the key segment text data into a pre-trained archive retrieval model for retrieval, and output a retrieval set matching the key segment and a relevance; Wherein, the search set is an electronic text data document, an audio data document, and a video data document that matches the key segment text data; A reading guide generating module, used for generating a reading recommendation guide for the search set according to the relevance; wherein the reading recommendation guide at least includes the relevance and the recommended reading order; The updating module is used to associate the reading recommendation guide with the archive image data, and update and store it in the archive retrieval system.
10. A readable storage medium having a computer program stored thereon, characterized in that: The computer program is executed by the processor as the deep learning-based archive retrieval method as described in any one of claims 1-8.
Citation Information
Patent Citations
An electronic record management method
CN112329669B
Personal knowledge management method and system based on content retrieval
CN116521626A
Archive data retrieval method, system and device
CN119271630A
Archive information retrieval method based on multi-modal model
CN119759949A
Archive information extraction management method and system based on multi-modal learning
CN119939120A