An intelligent retrieval method for multi-type files in the field of digital culture
By introducing intelligent file search methods in the field of digital culture, and using knowledge graphs and label forest structures for multi-level recursive query, the problem of inefficient data management in the field of digital culture is solved, and efficient and accurate data retrieval and intelligent classification are achieved.
Patent Information
- Application Number
- CN202411710533.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-27
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-11-27
AI Technical Summary
It is difficult to effectively classify and manage user data in the digital culture field. The existing search algorithm can only simply query folders or file names, resulting in inefficient data management and it takes a long time to find specific content.
An intelligent file search method is proposed, which conducts multi-level recursive query based on folder levels, file content and tags. By establishing a knowledge graph in the field of digital culture, the files are labeled, and the label forest structure is generated, and the search conditions input by users are extended and recursively searched.
Efficient and accurate data retrieval is realized, and users can quickly find the required data in complex file structures. The system is intelligently classified and managed through tags and recursive algorithms, which is suitable for large-scale data management scenarios.
Smart Images

Figure CN119719036B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an intelligent retrieval method, electronic device and storage medium for multiple types of files in the field of digital culture. Background Art
[0002] With the rapid increase in data volume, user data in the digital cultural field (including folders, file contents, and data tags) is difficult to classify and manage effectively. Especially in scenarios where cultural enterprises and institutions need to process a large amount of different types of data, users often need to efficiently query and retrieve folders and their contents. Existing search algorithms usually only support simple folder or file name queries, resulting in inefficient data management and a long time to find specific content. Summary of the invention
[0003] In order to overcome the above defects, this application is proposed, that is, an intelligent file retrieval method, which can perform multi-level recursive queries based on folder levels, file contents and tags to achieve efficient and accurate data retrieval. Users can quickly find the required data in a complex file structure through simple query operations, and can perform detailed classification and intelligent retrieval based on file content and tag information.
[0004] Specifically, the present application provides, on the one hand, an intelligent retrieval method for multiple types of files in the field of digital culture, including:
[0005] Establish a knowledge graph in the field of digital culture;
[0006] Through the cultural domain knowledge graph, the files stored by the user are labeled, and a label forest structure of each file is generated by mapping;
[0007] Acquire the search conditions input by the user, expand the search conditions in combination with the digital culture field knowledge graph, and generate a tag domain to be selected;
[0008] The user selects the candidate tag domain to obtain a search tag set, wherein each element in the search tag set includes a search weight attribute;
[0009] Taking the specified node set in the tag forest structure as the starting point set, recursively searching the tag forest structure corresponding to the files and folders under all the nodes in the starting point set, and adding the location of the file hit by the search tag set to the search results;
[0010] Based on the search results, a search result tree structure is formed and presented to the user, and the digital culture field knowledge graph is updated.
[0011] In one embodiment, establishing a knowledge graph in the digital culture field includes:
[0012] According to the knowledge norms and constraints in the digital culture field, define the entities, relationships and attributes of the digital culture field knowledge graph;
[0013] Entity extraction, relationship extraction and attribute extraction are performed on different types of digital culture field sample data provided by users to form the digital culture field knowledge graph.
[0014] In one embodiment, the files stored by the user are labeled through the cultural domain knowledge graph, and a label forest structure of each file is mapped and generated, including:
[0015] After the user stores the file in the storage medium, the file is identified by label according to the file type, and a label forest structure of the file is generated by mapping the digital cultural field knowledge graph, wherein the label forest structure is composed of at least one label tree, each of which points to a semantic subset, and the internal nodes of the label tree are semantically related;
[0016] Among them, when the label obtained after the label identification cannot be mapped in the digital culture field knowledge graph, a new entity based on the label is created in the digital culture field knowledge graph, and corresponding relationships and attributes are formed.
[0017] In one embodiment, the tag identification is performed on the file according to the different file types, and the tag forest structure of the file is generated by mapping the digital culture field knowledge graph, including:
[0018] For a text file, the text file is converted into a text feature vector through a natural language processing method to calculate the text relevance between the text feature vector and the digital cultural field knowledge graph; when the text relevance meets the set text retrieval threshold, the text feature vector is mapped to generate a label forest structure of the text file;
[0019] For an image file, feature extraction is performed through a machine learning method to generate an image feature vector representing the image file, and the image association degree between the image feature vector and the digital culture field knowledge graph is calculated; when the image association degree meets a specified image retrieval threshold, the image feature vector is mapped to generate a label forest structure of the image file;
[0020] For an audio file, sampling is performed according to a specified audio sampling standard to obtain an audio sampling set, the audio sampling set is converted into a text set through speech recognition, and the audio relevance between the text set and the digital culture field knowledge graph is calculated; by comparing the audio relevance with a specified audio retrieval threshold, the text set is compressed into an audio feature vector, and a label forest structure of the audio file is generated by mapping;
[0021] For video files, frame sampling is performed according to the specified frame sampling standard to obtain a sampled frame set, the visual features and changes in the time dimension of each frame are extracted to form a video sampling feature set, and the video correlation between each element in the video sampling set and the retrieval label domain is calculated; if the video correlation meets the specified video retrieval threshold, the element is retained, otherwise it is eliminated to obtain a video feature vector; the video feature vector is mapped to generate a label forest structure of the video file.
[0022] In one embodiment, the search condition is expanded in combination with the digital culture field knowledge graph to generate a tag domain to be selected, including:
[0023] Acquire the search condition, wherein the search condition includes at least one of a text condition, a picture condition, an audio condition, and a video condition;
[0024] According to different file types, feature extraction is performed on the search conditions to obtain an initial sequence of search conditions;
[0025] The initial sequence of search conditions is sorted according to the semantic weight and occurrence frequency of each element in the initial sequence of search conditions.
[0026] According to the sorting, an expansion threshold is calculated, and each element in the initial sequence of the retrieval conditions is expanded according to the threshold in combination with the digital culture field knowledge graph to generate a tag domain to be selected.
[0027] In one embodiment, taking the specified node set in the tag forest structure as the starting point set, recursively searching the tag forest structure corresponding to the files and folders under all the nodes in the starting point set, and adding the location of the file hit by the search tag set to the search results, including:
[0028] sorting the search tag set according to the search weight attribute;
[0029] Taking each element in the starting point set as initial input, recursively searching in parallel the tag forest structure pointed to by the file or folder corresponding to the element;
[0030] When searching the tag forest structure pointed to by each file or folder, multiple threads are started for parallel processing, wherein each thread is corresponding to searching a tag tree in the tag forest structure;
[0031] When no tag is hit during the recursive search process, the search for the folder and its files or the tag forest structure corresponding to the folder is stopped; when a tag is hit, the recursive search for the tag forest structure corresponding to the folder or file continues, and the location information of the file or folder is added to the search results.
[0032] In one embodiment, a retrieval result tree structure is formed based on the retrieval result, presented to the user, and the digital culture field knowledge graph is updated, including:
[0033] Extracting features of the file or folder pointed to by the search result to obtain summary information of the file or folder;
[0034] The search results are presented to the user in a hierarchical tree structure, and each node in the hierarchical structure includes the summary information.
[0035] In one embodiment, taking a specified node set in the tag forest structure as a starting point set, recursively searching the tag forest structures corresponding to files and folders under all nodes in the starting point set, and adding the locations of the files hit by the search tag set to the search results, further comprising:
[0036] Setting a semantic relevance threshold, and in the recursive search process, comparing the tag forest structure pointed to by the file or folder corresponding to the element with the search tag set for similarity;
[0037] When the result of the similarity comparison meets the semantic relevance threshold, the location of the file hit by the search tag set is added to the search result.
[0038] The second aspect of the present application provides an electronic device, comprising at least one processor and a memory, wherein the memory stores computer-executable instructions, and the at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the intelligent retrieval method for multiple types of files in the digital culture field as described in any one of the first aspects.
[0039] The third aspect of the present application provides a computer-readable storage medium storing a plurality of program codes, wherein the program codes are suitable for being loaded and run by a processor to execute the intelligent retrieval method for multiple types of files in the digital culture field as described in any one of the first aspects.
[0040] The above one or more technical solutions in this application have at least one or more of the following beneficial effects:
[0041] (1) Multi-level recursive deep search: This algorithm supports recursive query and can go deep into multi-level folder structures to ensure that files in each level can be searched, greatly improving the depth and comprehensiveness of the query.
[0042] (2) Multi-dimensional combined query: The algorithm supports multi-dimensional combined query by combining folder name, file content and tags, achieving a more accurate and flexible retrieval function than traditional single-dimensional query.
[0043] (3) Semantic analysis and intelligent expansion: Through natural language processing technology, the system can parse file content and generate semantic information. Combined with tag-driven queries, the system can intelligently recommend related files based on the file content and tag relationships.
[0044] (4) Efficient data classification and management: The system realizes intelligent classification and efficient management of data through labels and recursive algorithms, which is particularly suitable for large-scale data management scenarios of cultural enterprises and institutions.
[0045] In general, through this solution, users can perform intelligent retrieval in complex file structures, which not only improves query efficiency but also enables better management and utilization of digital cultural data. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The disclosure of the present application will become easier to understand with reference to the accompanying drawings. It is easy for those skilled in the art to understand that these drawings are only for illustrative purposes and are not intended to limit the scope of protection of the present application. Among them:
[0047] Figure 1 A flow chart showing an intelligent retrieval method for multiple types of files in the digital culture field according to an embodiment of the present invention is shown.
[0048] Figure 2 A schematic diagram of a knowledge graph in the digital culture field according to an embodiment is shown.
[0049] Figure 3 A schematic diagram of a file system label forest structure according to an embodiment of the present invention is shown.
[0050] Figure 4 A schematic diagram showing the results obtained according to an embodiment of the present patent is shown. DETAILED DESCRIPTION
[0051] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It is also necessary to explain that, for ease of description, only the parts related to the relevant invention are shown in the accompanying drawings.
[0052] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0053] In a first aspect, the present invention provides an intelligent retrieval method for multiple types of files in the field of digital culture.
[0054] Figure 1 A flowchart of a method for intelligently retrieving multiple types of files in the digital culture field according to an embodiment of the present invention is shown. In one embodiment, the method for intelligently retrieving multiple types of files in the digital culture field includes:
[0055] S101: Establish a knowledge graph in the field of digital culture.
[0056] In one embodiment, establishing a knowledge graph in the digital culture field includes:
[0057] According to the knowledge norms and constraints in the digital culture field, define the entities, relationships and attributes of the digital culture field knowledge graph;
[0058] Entity extraction, relationship extraction and attribute extraction are performed on different types of digital culture field sample data provided by users to form the digital culture field knowledge graph.
[0059] Figure 2 A schematic diagram of a knowledge graph in the digital culture field according to an embodiment is shown.
[0060] like Figure 2 As shown in the figure, Dream of the Red Chamber 201, Cao Xueqin 202, Lin Daiyu 203, National Ballet of China 204, and Tianqiao Theater 205 are all entities. These entities can have multiple attributes. For example, the attributes of Cao Xueqin 202 may include Qing Dynasty, novelist, etc. The relationship between entities is as follows Figure 2 As shown, Cao Xueqin 202 created Dream of the Red Chamber 201, the Central Ballet Company 204 performed Dream of the Red Chamber 201 at the Tianqiao Theater 205, the protagonist of Dream of the Red Chamber 201 is Lin Daiyu 203, and so on.
[0061] It is worth noting that the field of digital culture is different from general fields and has many proprietary terms or connotation relationships. Therefore, in the process of constructing the knowledge graph in the field of digital culture, we can use dictionary files in this field to enrich the comprehensiveness of the knowledge graph in the field of digital culture, thereby improving the accuracy and efficiency of feature recognition and file search.
[0062] S102: Labeling the files stored by the user through the cultural field knowledge graph, and mapping and generating a label forest structure for each of the files.
[0063] In one embodiment, the files stored by the user are labeled through the cultural domain knowledge graph, and a label forest structure of each file is mapped and generated, including:
[0064] After the user stores the file in the storage medium, the file is identified by label according to the file type, and a label forest structure of the file is generated by mapping the digital cultural field knowledge graph, wherein the label forest structure is composed of at least one label tree, each of which points to a semantic subset, and the internal nodes of the label tree are semantically related;
[0065] Among them, when the label obtained after the label identification cannot be mapped in the digital culture field knowledge graph, a new entity based on the label is created in the digital culture field knowledge graph, and corresponding relationships and attributes are formed.
[0066] Figure 3 A schematic diagram of a file system label forest structure according to an embodiment of the present invention is shown.
[0067] like Figure 3 As shown, in a file system, a first folder 31 may contain a second folder 311 and a first file 312, and the second folder 311 may contain a second file 3111, a third file 3112, and a fourth file 3113. The first file 312, the second file 3111, the third file 3112, and the fourth file 3113 may be any of text, image, audio, and video type files.
[0068] For example, after the first file 312 is stored in the first folder 31, the system will perform a tagging process on the first file 312 and extract the tags contained in the first file 312. These tags may belong to different concepts, thus forming different tag trees, such as Figure 3 As shown, four label trees can be extracted from the first file 312, and these four label trees together constitute the label forest structure of the first file 312. For example, the first file 312 can be an introduction short film of Dream of Red Mansions, from which label trees with four themes, Cao Xueqin, Lin Daiyu, Tianqiao Theater, and Central Ballet Company, can be extracted. Taking the label tree of Cao Xueqin as an example, this label tree can include label entities such as Qing Dynasty and novelist. A file corresponds to multiple label trees, and these label trees converge together to form the label forest structure corresponding to this file.
[0069] Each entity of the tag tree in the tag forest structure needs to have a corresponding node in the construction of the digital culture field knowledge graph. After extracting the tag information of a file, if there is no corresponding entity in the digital culture field knowledge graph, then a new entity based on this tag tree can be created in the digital culture field knowledge graph, and the corresponding attributes and relationships can be brought into the digital culture field knowledge graph. For example, a tag "Grand View Garden" is extracted from the second file 312, and there is no entity about "Grand View Garden" in the digital culture field knowledge graph. Then, a new "Grand View Garden" entity can be created in the digital culture field knowledge graph, and its attributes and the relationship between it and other entities in the digital culture field knowledge graph can be brought into it.
[0070] Let's look at the first folder 31 and the second folder 311. For a folder, except for the name of the folder, it does not have any semantic elements. Therefore, the label forest structure corresponding to a folder should actually be the collection of label forest structures corresponding to all files in the folder directory, as well as the label forest structure obtained by expanding the folder name (since the folder name is generally short, this forest may have only one label tree). In order to increase the comprehensiveness of the search results, the label forest structure corresponding to a folder can be the collection of all files and folders recursively in several layers under the directory. Of course, the merged label forest structure can have a reorganization and deduplication operation to reduce the redundancy of data and improve the efficiency of later retrieval.
[0071] For example, the folder name of the second folder 311 is "Red Mansion Dream Live Video", then firstly, the label of "Red Mansion Dream Live Video" can be extracted to obtain the label entity "Red Mansion Dream", and then the corresponding label forest structure can be obtained by expanding it. Obviously, since only one entity "Red Mansion Dream" is extracted, there may be only one label tree in its corresponding label forest structure.
[0072] On the basis of this label forest structure, the collection of label forest structures corresponding to the second file 3111, the third file 3112, and the fourth file 3113 in the second folder can also be included. The label forest structure corresponding to the first folder 31 can be the label forest structure corresponding to the folder name itself, plus the collection of label forest structures corresponding to the first file 312, the second file 3111, the third file 3112, and the fourth file 3113. Obviously, since the contents of these files are all related to "Dream of Red Mansions", there is a high probability that there will be intersections between the label forest structures. Therefore, after a simple merge, the merged label forest structure needs to be reorganized, deduplicated, and other operations. After these operations are completed, each file or folder in the file system has a corresponding label forest structure.
[0073] In one embodiment, according to different file types, the file is labeled and identified, and the label forest structure of the file is generated by mapping the digital culture field knowledge graph, including:
[0074] For a text file, the text file is converted into a text feature vector through a natural language processing method to calculate the text relevance between the text feature vector and the digital cultural field knowledge graph; when the text relevance meets the set text retrieval threshold, the text feature vector is mapped to generate a label forest structure of the text file;
[0075] For an image file, feature extraction is performed through a machine learning method to generate an image feature vector representing the image file, and the image association degree between the image feature vector and the digital culture field knowledge graph is calculated; when the image association degree meets a specified image retrieval threshold, the image feature vector is mapped to generate a label forest structure of the image file;
[0076] For an audio file, sampling is performed according to a specified audio sampling standard to obtain an audio sampling set, the audio sampling set is converted into a text set through speech recognition, and the audio relevance between the text set and the digital culture field knowledge graph is calculated; by comparing the audio relevance with a specified audio retrieval threshold, the text set is compressed into an audio feature vector, and a label forest structure of the audio file is generated by mapping;
[0077] For video files, frame sampling is performed according to the specified frame sampling standard to obtain a sampled frame set, the visual features and changes in the time dimension of each frame are extracted to form a video sampling feature set, and the video correlation between each element in the video sampling set and the retrieval label domain is calculated; if the video correlation meets the specified video retrieval threshold, the element is retained, otherwise it is eliminated to obtain a video feature vector; the video feature vector is mapped to generate a label forest structure of the video file.
[0078] Specifically, the digital culture domain knowledge graph can be an already constructed proprietary domain knowledge graph, an incomplete proprietary domain knowledge graph, or even a brand new proprietary domain knowledge graph. Since the concepts, entities, and terms of a field are constantly evolving with the times, the digital culture domain knowledge graph cannot be fixed and needs to be constantly enriched and updated. Therefore, after labeling different types of data files, it is very likely that the identified labels cannot find corresponding entities in the digital culture domain knowledge graph. At this time, the entities, attributes, and relationships corresponding to the labels can be expanded to the digital culture domain knowledge graph.
[0079] After labeling different files, the purpose of comparing them with the knowledge graph in the field of digital culture is to filter out some labels that are not relevant to the field of digital culture, thereby improving the efficiency of retrieval.
[0080] S103: Acquire the search conditions input by the user, expand the search conditions in combination with the digital culture field knowledge graph, and generate a tag domain to be selected.
[0081] In one embodiment, the search condition is expanded in combination with the digital culture field knowledge graph to generate a tag domain to be selected, including:
[0082] Acquire the search condition, wherein the search condition includes at least one of a text condition, a picture condition, an audio condition, and a video condition;
[0083] According to different file types, feature extraction is performed on the search conditions to obtain an initial sequence of search conditions;
[0084] The initial sequence of search conditions is sorted according to the semantic weight and occurrence frequency of each element in the initial sequence of search conditions.
[0085] According to the sorting, an expansion threshold is calculated, and each element in the initial sequence of the retrieval conditions is expanded according to the threshold in combination with the digital culture field knowledge graph to generate a tag domain to be selected.
[0086] For example, when a user searches for a file, they can select search conditions and can search for videos or videos and text at the same time. Users can also search for files by file. For example, users can directly enter a video file by dragging and dropping. The system will extract features from the file and obtain the initial sequence of search conditions. For example, if a user directly enters an excerpt from "Grandma Liu Enters the Grand View Garden" in "A Dream of Red Mansions" as a search condition, the system will automatically extract features from the text and obtain the initial sequence of search conditions.
[0087] Furthermore, the excerpt of "Grandma Liu Visits the Grand View Garden" in "A Dream of Red Mansions" may contain a large portion of the main character "Grandma Liu", so "Grandma Liu" should be ranked higher in the initial sequence of search conditions, so that the search results can be closer to the user's search intention.
[0088] S104: The user selects the candidate tag domain to obtain a search tag set, wherein each element in the search tag set includes a search weight attribute.
[0089] Specifically, when searching, users can input their own information or select the tag items provided by the system, and generate search conditions based on the tag items selected by the user. It is worth noting that when selecting tags, users can also set different priorities for different tags, so that the search content that users are concerned about can obtain higher priority searches.
[0090] S105: Taking the specified node set in the tag forest structure as the starting point set, recursively search the tag forest structure corresponding to the files and folders under all the nodes in the starting point set, and add the location of the file hit by the search tag set to the search results.
[0091] In one embodiment, taking the specified node set in the tag forest structure as the starting point set, recursively searching the tag forest structure corresponding to the files and folders under all the nodes in the starting point set, and adding the location of the file hit by the search tag set to the search results, including:
[0092] sorting the search tag set according to the search weight attribute;
[0093] Taking each element in the starting point set as initial input, recursively searching in parallel the tag forest structure pointed to by the file or folder corresponding to the element;
[0094] When searching the tag forest structure pointed to by each file or folder, multiple threads are started for parallel processing, wherein each thread is corresponding to searching a tag tree in the tag forest structure;
[0095] When no tag is hit during the recursive search process, the search for the folder and its files or the tag forest structure corresponding to the folder is stopped; when a tag is hit, the recursive search for the tag forest structure corresponding to the folder or file continues, and the location information of the file or folder is added to the search results.
[0096] Specifically, the entire recursive search process is actually a search process for the label forest structure mapped by the file system. Figure 3 For example, in the recursive search process, the label forest structure corresponding to the first folder 31 is first searched. If it matches, the label forest structure corresponding to the second folder 311 and the first file 312 under it is recursively searched; if the label forest structure corresponding to the first folder 31 does not match at all, then the search for the label forest structure of the first folder 31 and all the folders or files under it is stopped.
[0097] Still Figure 3For example, in another example, there are 4 tag trees in the tag forest structure corresponding to the second file 312, then 4 threads can be started to search the tag forest structure corresponding to the second file 312 in parallel, thereby improving the retrieval efficiency.
[0098] S106: Based on the search results, a search result tree structure is formed, presented to the user, and the digital culture field knowledge graph is updated.
[0099] In one embodiment, a retrieval result tree structure is formed based on the retrieval result, presented to the user, and the digital culture field knowledge graph is updated, including:
[0100] Extracting features of the file or folder pointed to by the search result to obtain summary information of the file or folder;
[0101] The search results are presented to the user in a hierarchical tree structure, and each node in the hierarchical structure includes the summary information.
[0102] For example, Figure 4 FIG. 1 shows a schematic diagram of the results obtained according to an embodiment of the present invention. Figure 4 As shown, after the search, the search results include two first-level directory folders (first folder 41 and second folder 42), the first folder 41 contains the first image file 411 and the second image file 412; the second folder 42 contains the third folder 421, the fourth folder 422 and the second text file 423; the third folder contains the second image file 4211 and the first video file 4212; the fourth folder contains the first audio file 4221 and the second video file 4222. It can be seen that the above tree-shaped directory format is presented to the user, which can not only establish a mapping relationship with the architecture of the file system, but also facilitate the user to browse and search. In addition, when the user clicks or floats the mouse on the third folder, the summary 40 will automatically appear on the user graphical interface, which can facilitate the user to quickly understand the main content of the files contained in the third folder, increase the user's comprehensibility and operability, and obtain a richer user experience.
[0103] In one embodiment, the method further comprises:
[0104] Setting a semantic relevance threshold, and in the recursive search process, comparing the tag forest structure pointed to by the file or folder corresponding to the element with the search tag set for similarity;
[0105] When the result of the similarity comparison meets the semantic relevance threshold, the location of the file hit by the search tag set is added to the search result.
[0106] Specifically, during the search process, whether the label forest structure corresponding to the file or folder meets the search conditions entered by the user is judged by comparing the similarity between the label forest structure corresponding to the file or folder and the search label set generated by the search conditions entered by the user. The user can set a relevance threshold in advance. When the search conditions need to be stricter, the relevance threshold can be set larger; on the contrary, if the search conditions are more relaxed, the relevance threshold can be set smaller. The calculation of similarity can be implemented by using similarity calculation methods such as cosine similarity and Euclidean similarity, which are more commonly used in the field of machine learning.
[0107] The second aspect of the present invention also relates to an electronic device, comprising at least one processor and a memory, wherein the memory stores computer-executable instructions, and the at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the intelligent retrieval method for multiple types of files in the digital culture field as described in any one of the first aspects.
[0108] The third aspect of the present invention also relates to a computer-readable storage medium, in which multiple program codes are stored, characterized in that the program code is suitable for being loaded and run by a processor to execute the intelligent retrieval method for multiple types of files in the digital culture field as described in any one of the first aspects.
[0109] The functions of each module in each system of the embodiment of the present invention can refer to the corresponding description in the above method, which will not be repeated here.
[0110] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are contradictory.
[0111] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention belong.
[0112] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.
[0113] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0114] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0115] In addition, each functional unit in each embodiment of the present invention may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The storage medium may be a read-only memory, a disk or an optical disk, etc.
[0116] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of various changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. An intelligent retrieval method for multiple types of files in the field of digital culture, characterized in that: include: Establish a knowledge graph in the field of digital culture; Through the cultural domain knowledge graph, the files stored by the user are labeled and mapped to generate a label forest structure for each file, including: After the user stores the file in the storage medium, the file is identified by label according to the file type, and the label forest structure of the file is generated by mapping the digital culture domain knowledge graph, wherein the label forest structure is composed of at least one label tree, each label tree points to a semantic subset, and the internal nodes of the label tree are semantically related; wherein, when the label obtained after the label identification cannot be mapped in the digital culture domain knowledge graph, a new entity based on the label is created in the digital culture domain knowledge graph, and corresponding relationships and attributes are formed; Acquire the search conditions input by the user, expand the search conditions in combination with the digital culture field knowledge graph, and generate a tag domain to be selected; The user selects the candidate tag domain to obtain a search tag set, wherein each element in the search tag set includes a search weight attribute; Taking the specified node set in the tag forest structure as the starting point set, recursively searching the tag forest structure corresponding to the files and folders under all the nodes in the starting point set, and adding the location of the file hit by the search tag set to the search results; Based on the search results, a search result tree structure is formed and presented to the user, and the digital culture field knowledge graph is updated.
2. The method according to claim 1, characterized in that Establish a knowledge graph in the field of digital culture, including: According to the knowledge norms and constraints in the digital culture field, define the entities, relationships and attributes of the digital culture field knowledge graph; Entity extraction, relationship extraction and attribute extraction are performed on different types of digital culture field sample data provided by users to form the digital culture field knowledge graph.
3. The method according to claim 1, characterized in that: The step of performing label recognition on the file according to the different file types and generating a label forest structure of the file by mapping the digital culture domain knowledge graph includes: For a text file, the text file is converted into a text feature vector through a natural language processing method to calculate the text relevance between the text feature vector and the digital cultural field knowledge graph; when the text relevance meets the set text retrieval threshold, the text feature vector is mapped to generate a label forest structure of the text file; For an image file, feature extraction is performed through a machine learning method to generate an image feature vector representing the image file, and the image association degree between the image feature vector and the digital culture field knowledge graph is calculated; when the image association degree meets a specified image retrieval threshold, the image feature vector is mapped to generate a label forest structure of the image file; For an audio file, sampling is performed according to a specified audio sampling standard to obtain an audio sampling set, the audio sampling set is converted into a text set through speech recognition, and the audio relevance between the text set and the digital culture field knowledge graph is calculated; by comparing the audio relevance with a specified audio retrieval threshold, the text set is compressed into an audio feature vector, and a label forest structure of the audio file is generated by mapping; For video files, frame sampling is performed according to the specified frame sampling standard to obtain a set of sampled frames, and the visual features and changes in the time dimension of each frame are extracted to form a video sampling feature set. The video correlation between each element in the video sampling set and the retrieval label domain is calculated; if the video correlation meets the specified video retrieval threshold, the element is retained, otherwise it is eliminated to obtain a video feature vector; the video feature vector is mapped to generate a label forest structure of the video file.
4. The method according to claim 1, characterized in that: The search conditions are expanded in combination with the digital culture field knowledge graph to generate a tag domain to be selected, including: Acquire the search condition, wherein the search condition includes at least one of a text condition, a picture condition, an audio condition, and a video condition; According to different file types, feature extraction is performed on the search conditions to obtain an initial sequence of search conditions; Sorting the initial sequence of search conditions according to the semantic weight and occurrence frequency of each element in the initial sequence of search conditions; According to the sorting, an expansion threshold is calculated, and each element in the initial sequence of the retrieval conditions is expanded according to the threshold in combination with the digital culture field knowledge graph to generate a tag domain to be selected.
5. The method according to claim 1, characterized in that Taking the specified node set in the tag forest structure as the starting point set, recursively searching the tag forest structure corresponding to the files and folders under all the nodes in the starting point set, and adding the location of the file hit by the search tag set to the search results, including: sorting the search tag set according to the search weight attribute; Taking each element in the starting point set as initial input, recursively searching in parallel the tag forest structure pointed to by the file or folder corresponding to the element; When searching the tag forest structure pointed to by each file or folder, multiple threads are started for parallel processing, wherein each thread is corresponding to searching a tag tree in the tag forest structure; When no tag is hit during the recursive search process, the search for the folder and its files or the tag forest structure corresponding to the folder is stopped; when a tag is hit, the recursive search for the tag forest structure corresponding to the folder or file continues, and the location information of the file or folder is added to the search results.
6. The method according to claim 1, characterized in that According to the search results, a search result tree structure is formed and presented to the user, and the digital culture field knowledge graph is updated, including: Extracting features of the file or folder pointed to by the search result to obtain summary information of the file or folder; The search results are presented to the user in a hierarchical tree structure, and each node in the hierarchical tree structure includes the summary information.
7. The method according to claim 5, characterized in that Also includes: Setting a semantic relevance threshold, and in the recursive search process, comparing the tag forest structure pointed to by the file or folder corresponding to the element with the search tag set for similarity; When the result of the similarity comparison meets the semantic relevance threshold, the location of the file hit by the search tag set is added to the search result.
8. An electronic device, characterized in that: It includes at least one processor and a memory, wherein the memory stores computer-executable instructions, and the at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the intelligent retrieval method for multiple types of files in the digital culture field as described in any one of claims 1 to 7.
9. A computer-readable storage medium storing a plurality of program codes, characterized in that: The program code is suitable for being loaded and run by a processor to execute the intelligent retrieval method for multiple types of files in the digital culture field according to any one of claims 1 to 7.
Citation Information
Patent Citations
Knowledge reasoning-based big data service label extension method and system
CN111737400A
File label classification method and device
CN111930944A