A digital file management method, system, device and medium
By obtaining the entity data of the archives, determining the entity relationship and common entities, establishing the archive index and performing related storage, the problem that the archives cannot be retrieved in an overall manner in the existing technology is solved, and efficient retrieval of the archives is achieved.
Patent Information
- Application Number
- CN202310751353.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-21
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-06-21
AI Technical Summary
In the prior art, when digital archives are stored in the form of Word, PDF, etc., it is impossible to retrieve the entire archive library, resulting in inefficient search.
By obtaining the entity data of the archive, determining the entity relationship and common entities, establishing the archive index, and performing associated storage to form an archive library, and using entity data, entity relationships, archive index and common entities for related storage to form an archive library.
It realizes efficient retrieval of the entire archive library and improves the efficiency of archive retrieval.
Smart Images

Figure CN117009616B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a digital file management method, system, device and medium. Background Art
[0002] Digitalization of archives is a new form of archival information that emerged with the development of computer technology, scanning technology, scanning linear CCD technology, OCR technology, digital photography technology (recording, video), database technology, multimedia technology, and storage technology. It converts archival resources on various carriers into digital archival information, stores them in digital form, connects them in a networked form, and uses a computer system for management to form an orderly structured archival information database, providing timely utilization and realizing resource sharing.
[0003] Currently, most digital archives used in various fields are mainly in the form of Word and PDF, that is, archival files are stored in the form of Word, PDF, etc., and a text-based index file library is formed. Based on the aforementioned file library, since digital archives are stored independently, when retrieving archival files, only individual archival files can be retrieved, and the retrieval of the entire archive library cannot be realized, thus reducing the efficiency of archival retrieval. Summary of the Invention
[0004] The present invention provides a digital file management method, system, device and medium to solve the defects in the prior art.
[0005] The present invention provides a digital file management method, including:
[0006] Obtaining entity data corresponding to multiple files to be processed;
[0007] Determining the entity relationship between entities according to the entity data, and performing text segmentation on the files to be processed according to the entity data to obtain file indexes;
[0008] Determining the common entities corresponding to the multiple files to be processed according to the entity data;
[0009] Associatively storing the entity data, the entity relationship, the file indexes and the common entities to form an archive library.
[0010] According to the digital file management method provided by the present invention, the obtaining of the entity data corresponding to multiple files to be processed includes:
[0011] Extracting entities from multiple files to be processed according to a preset meta-database, where the preset meta-database includes archival entity names;
[0012] Extracting attribute information corresponding to the entities in the files to be processed;
[0013] Determine the file class diagram according to the entity and the attribute information;
[0014] Correspondingly, the determining the entity relationship between entities according to the entity data includes:
[0015] Determine the entity relationship between entities according to the context information of the file to be processed and the entity data, and obtain the edges between the file class diagrams based on the entity relationship;
[0016] The associatively storing the entity data, the entity relationship, the file index, and the common entity includes:
[0017] Associatively store the file class diagrams according to the edges between the file class diagrams.
[0018] According to a digital file management method provided by the present invention, the obtaining the file index by text segmentation of the file to be processed according to the entity data includes:
[0019] Segment the text in the file to be processed according to the entity data into index units including index words and entity type words;
[0020] Use the index unit to obtain a data index file, and obtain an inverted index file based on the position information of the index words in the index unit in the multiple files to be processed. The file index includes the data index file and the inverted index file.
[0021] According to a digital file management method provided by the present invention, the determining the common entity corresponding to the multiple files to be processed according to the entity data includes:
[0022] Determine the common entities among the files to be processed according to the entity data, so as to obtain an initial common entity set;
[0023] Perform synonym replacement on the attribute information corresponding to the entities in the initial common entity set to obtain the first attribute;
[0024] Delete the entities with duplicate attributes according to the attribute information corresponding to the entities and the first attribute to obtain the final common entity set;
[0025] Correspondingly, the associatively storing the entity data, the entity relationship, the file index, and the common entity includes:
[0026] Associatively store the entities according to the final common entity set.
[0027] A digital file management method provided by the present invention, extracting attribute information corresponding to the entity from the file to be processed includes:
[0028] Using a pre-trained entity attribute extraction model to extract attribute information of the entity in the file to be processed;
[0029] Wherein, the pre-trained entity attribute extraction model is a convolutional neural network model and is trained based on training files and corresponding labels.
[0030] A digital file management method provided by the present invention, determining the entity relationship between entities according to the context information and entity data of the file to be processed includes:
[0031] Performing sentence content parsing and vectorization processing on the file to be processed to obtain word vectors;
[0032] Using bidirectional LSTM to perform forward and backward context learning on the word vectors to obtain word vectors including context information;
[0033] Using an attention mechanism to determine the importance of each word vector including context information in the file difference detection task to obtain a weight vector;
[0034] Obtaining a fusion result of lexical-level features by multiplying the word vectors including context information with the weight vector and using it as sentence-level features;
[0035] Classifying the sentence-level features through a classifier to obtain corresponding difference categories, and using the difference categories as entity relationships.
[0036] A digital file management method provided by the present invention, after associatively storing the entity data, the entity relationship, the file index, and the common entity to form a file library, the method further includes:
[0037] Obtaining a retrieval condition;
[0038] Obtaining multiple files corresponding to the entity from the file library according to the entity corresponding to the retrieval condition.
[0039] The present invention also provides a digital file management system, including:
[0040] An entity acquisition module, configured to acquire entity data corresponding to multiple files to be processed;
[0041] A relationship and index acquisition module, configured to determine the entity relationship between entities according to the entity data, and perform text segmentation on the file to be processed according to the entity data to obtain a file index;
[0042] A common entity acquisition module, configured to determine a common entity corresponding to the multiple to-be-processed files according to the entity data;
[0043] An associated storage module, configured to perform associated storage on the entity data, the entity relationships, the file indexes, and the common entity to form a file library.
[0044] The present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of any of the above digital file management methods are implemented.
[0045] The present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above digital file management methods are implemented.
[0046] The digital file management method, system, device, and medium provided by the present invention extract entities, entity relationships, establish file indexes, and acquire common entities from files, and perform associated storage on the obtained data to obtain a fused file library. What is stored in this file library is not the entire file, but the entities, entity relationships, common entities, and file indexes corresponding to the entities in the files. In subsequent file retrieval, according to the retrieval conditions, multiple associated files can be obtained from the fused file library through the common entities, thereby realizing the retrieval of the entire file library and improving the file retrieval efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0048] Figure 1 is a flowchart of the digital file management method provided by the present invention;
[0049] Figure 2 is a structural diagram of the digital file management system provided by the present invention;
[0050] Figure 3 is a structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0052] Before explaining the digital file management method provided by the present invention, the professional terms involved are explained first. Among them, the Unified Modeling Language (UML) is a standard language for describing, visualizing, and documenting products of object-oriented systems, and it is a non-patented third-generation modeling and specification language. UML is a modeling tool for object-oriented design, independent of any specific programming language. UML uses a set of graphical symbols to describe software models. These graphical symbols are simple, intuitive, and standardized, making it relatively easy for developers to learn and master. The described software models can be intuitively understood and read, and due to their standardization, the accuracy and consistency of the models can be guaranteed. UML mainly includes use case diagrams, static diagrams, behavior diagrams, interaction diagrams, and implementation diagrams. The present invention selects the class diagram in the static diagram to associate and store files in combination with the characteristics of digital files (entity and attribute information in the text). The digital file management method of the present invention will be explained below with reference to the accompanying drawings.
[0053] Figure 1 is a schematic flowchart of the digital file management method provided by the present invention; as Figure 1 shown, a digital file management method includes the following steps:
[0054] S101, obtain entity data corresponding to multiple files to be processed.
[0055] In this step, the named entity recognition algorithm is used, and entity data included in multiple files to be processed is extracted according to the meta database, where the meta database contains multiple entity names available for matching. Among them, the named entity recognition algorithm is a commonly used named entity recognition algorithm, such as LSTM+CRF, CNN+CRF, BERT+(LSTM)+CRF, BiLSTM+CRF, HMM, attention model, transfer learning, etc. The present invention is not specifically limited thereto.
[0056] More specifically, in this embodiment, the process of extracting entity data using the named entity recognition algorithm includes the steps:
[0057] Text extraction: Extract text from the files to be processed in different formats. For example, directly obtain text from files in word format, and use OCR recognition to identify text in files in pdf format to obtain text.
[0058] Text segmentation: Perform word segmentation on the text to obtain a set of words.
[0059] Entity matching: Match each word in the set of words with the entity names in the metadata, and save the words that match successfully as the entities of the files to be processed into the entity list.
[0060] Entity annotation: Annotate the entities in the entity list. The results obtained by annotation include at least: the type of the entity, the start offset and the end offset of the entity in the text.
[0061] Entity attribute establishment: Extract the text related to the entity from the text according to the entity annotation results as the attributes of the entity. Finally, the obtained entity data includes the entity and its corresponding attributes.
[0062] Construct a UML class diagram based on the obtained entities and entity attributes, and store the file information in the form of a UML class diagram.
[0063] S102. Determine the entity relationships between entities according to the entity data, and perform text segmentation on the files to be processed according to the entity data to obtain file indexes.
[0064] In this step, use the entity relationship extraction algorithm to model the text information in the files to be processed, and automatically extract the semantic relationships between entities from the text as entity relationships. Among them, the entity relationship extraction algorithm is a commonly used entity relationship extraction algorithm, such as supervised feature-based and kernel function-based entity relationship extraction, semi-supervised Bootstrapping, pipeline Pipeline, joint learning Joint Learning, etc. The present invention does not limit this.
[0065] At the same time, perform text segmentation on the files to be processed according to the entity names to form different index units, obtain the database index file and the inverted index file according to different index units, and the file index is composed of the database index file and the inverted index file for subsequent document retrieval.
[0066] S103. Determine the common entities corresponding to the multiple files to be processed according to the entity data.
[0067] In this step, compare the entity data between the files to be processed, and store the same entities into the common entity set. It should be noted that steps S102 and S103 can be executed simultaneously.
[0068] S104. Associatively store the entity data, the entity relationships, the file indexes, and the common entities to form a file library.
[0069] In this step, the entity data, entity relationships, file indexes, and common entities obtained previously are associatively stored, and finally a merged file library is obtained. Specifically, for each entity, perform an N:N association according to the obtained entity relationships; associate each entity with its corresponding common entity; and associate each entity with its corresponding index.
[0070] According to the digital file management method provided by the embodiments of the present invention, by performing entity extraction, entity relationship extraction, file index establishment, and common entity acquisition on files, and associatively storing the obtained data, a merged file library is obtained. What is stored in this file library is not the entire file, but the entities, entity relationships, common entities, and file indexes corresponding to the files. In the subsequent file retrieval process, according to the retrieval conditions, multiple related files can be obtained from the merged file library through the common entities, thereby realizing the retrieval of the entire file library and improving the file retrieval efficiency.
[0071] Further, on the basis of the above embodiments, the obtaining of the entity data corresponding to multiple files to be processed includes:
[0072] Extract entities from multiple files to be processed according to a preset meta-database, and the preset meta-database includes file entity names.
[0073] Extract attribute information corresponding to the entities in the files to be processed.
[0074] Determine a file class diagram according to the entities and the attribute information.
[0075] Correspondingly, the determining of the entity relationships between entities according to the entity data includes:
[0076] Determine the entity relationships between entities according to the context information of the files to be processed and the entity data, and obtain the edges between the file class diagrams based on the entity relationships.
[0077] The associatively storing of the entity data, the entity relationships, the file indexes, and the common entities includes:
[0078] Associatively store the file class diagrams according to the edges between the file class diagrams.
[0079] In this embodiment, entities and entity attributes in the file to be processed are obtained through named entity recognition, and a UML class diagram corresponding to the file is constructed based on the entities and entity attributes. After extracting the entity relationships, edges are constructed for each UML diagram according to the entity relationships to realize the associated storage of multiple files, without storing the entire file.
[0080] Among them, determining the entity relationship between entities according to the context information and entity data of the file to be processed, and obtaining the edges between the file class diagrams based on the entity relationship includes:
[0081] Parse the sentence content of the file to be processed and perform vectorization processing to obtain word vectors.
[0082] Use bidirectional LSTM to perform forward and backward context learning on the word vectors to obtain word vectors including context information.
[0083] Use the attention mechanism to determine the importance of each word vector including context information in the file difference detection task to obtain a weight vector.
[0084] Specifically, input the word vectors including context information into the Attention layer to obtain weight scores. The weight scores represent the importance of the words corresponding to the word vectors in the file difference detection task. In addition, the sum of the weight scores is 1, thus indicating that the attention is distributed over all input words.
[0085] Multiply the word vectors including context information by the weight vector to obtain a fusion result of lexical-level features and use it as sentence-level features.
[0086] Classify the sentence-level features through a classifier to obtain the corresponding difference categories, and use the difference categories as entity relationships. Among them, the classifier is a conventional classifier (such as a Softmax classifier), which is not limited in this regard. The difference categories can be divided into the same and different, or can be further subdivided. The present invention does not limit this.
[0087] It should be noted that the file difference detection task includes the above-mentioned word vector conversion, bidirectional LSTM, Attention layer, and classifier. Inputting the text into the model corresponding to the file difference detection task can obtain the similarities and differences between files, and label the relationships between entities according to the similarities and differences.
[0088] According to the digital file management method provided by the embodiments of the present invention, by associatively storing the entities, entity attributes, and entity relationships in a UML class diagram, compared with the traditional whole file storage, the subsequent retrieval efficiency can be effectively improved, and multiple related files can be retrieved through entity retrieval. In addition, through the bidirectional LSTM and the attention mechanism, the relationship between entities can be determined by combining the context information and the importance of word vectors in the file difference detection task, improving the accuracy of the associated storage.
[0089] Further, on the basis of the above embodiments, the obtaining of the file index by segmenting the to-be-processed file according to the entity data includes:
[0090] Segmenting the text in the to-be-processed file according to the entity data includes index units of index words and entity type words.
[0091] Using the index unit to obtain a data index file, and obtaining an inverted index file based on the position information of the index words in the index unit in the multiple to-be-processed files, where the file index includes the data index file and the inverted index file.
[0092] In this embodiment, the text in the to-be-processed file is segmented so that the segmented text includes index words and entity type words, and the segmented text is the index unit.
[0093] The specific segmentation process includes: searching for entities according to the entity data, and if an entity is found, outputting the entity type word and offset of the entity according to the annotation result of the entity data (that is, the type of the entity, the start offset and the end offset of the entity in the text). Further, it is judged whether there is a superclass for the output entity type. If there is a superclass, all entity type words corresponding to the superclass entity type to the root node and the related offsets also need to be output to complete the output of all entity type words. Among them, the index words are indexed according to the general database establishment method to obtain the index words.
[0094] Using the above-mentioned index unit to form an index file, which is the database index file, that is, the forward index file.
[0095] At the same time, taking the index word as the center, the information of the same index word appearing in different files can be merged and stored to form an inverted index file.
[0096] According to the digital file management method provided by the embodiments of the present invention, by segmenting the text into index units that do not include index words and entity type words, and then forming a data index file and an inverted index file based on the index unit, thus supporting forward indexing and reverse indexing.
[0097] Further, based on the above embodiments, determining the common entities corresponding to the multiple files to be processed according to the entity data includes:
[0098] Determining the common entities among the files to be processed according to the entity data, so as to obtain an initial set of common entities.
[0099] Performing synonym replacement on the attribute information corresponding to the entities in the initial set of common entities, so as to obtain the first attribute.
[0100] Deleting the entities with duplicate attributes according to the attribute information corresponding to the entities and the first attribute, so as to obtain the final set of common entities.
[0101] Correspondingly, associatively storing the entity data, the entity relationships, the file indexes, and the common entities includes:
[0102] Associatively storing the entities according to the final set of common entities.
[0103] In this embodiment, after obtaining the initial set of common entities, synonym replacement is performed on the attributes of each entity in the set of common entities to obtain the first attribute; for the entities with the same first attribute, one is left and the others are deleted, that is, only one entity is retained.
[0104] According to the digital file management method provided by the embodiments of the present invention, using synonym replacement to check the duplication of entity attributes can effectively avoid the same attributes caused by different words, avoid duplicate attributes, and help establish a more efficient index. After the above UML data format and synonym duplication checking, digital file fusion can be achieved, so that the fused digital files can be efficiently stored and retrieved for objects.
[0105] Further, based on the above embodiments, extracting the attribute information corresponding to the entities from the files to be processed includes:
[0106] Using a pre-trained entity attribute extraction model to extract attribute information from the entities in the files to be processed.
[0107] Among them, the pre-trained entity attribute extraction model is a convolutional neural network model, and is trained based on training files and corresponding labels.
[0108] In this embodiment, the extraction of entity attributes is implemented through a CNN model. Specifically, training samples are first constructed: multiple groups of files that have completed the entity attribute establishment step are collected. These multiple groups of files are used as training inputs, and the entity attributes corresponding to the entities are used as training outputs. Among them, the convolutional neural network used is a conventional convolutional neural network, including an input layer, a convolutional layer, a Relu non-linear activation layer, a pooling layer, a fully connected layer, and an output layer.
[0109] An entity attribute extraction model is obtained through training with the convolutional neural network. The entire training process includes two stages: forward propagation network training and backward propagation network training. In forward propagation network training, each entity in the training files is processed through convolution and pooling to extract feature vectors, and the obtained feature vectors are converted into one-dimensional vectors and input into the fully connected layer. The classifier then obtains the recognition result, that is, the output vector. Each value of the output vector represents the probability that the established attribute matches the corresponding entity. Backward propagation network training is as follows: when the output result of the forward propagation network training does not match the corresponding attributes and entities in the expected output, the random gradient descent optimization algorithm is used for backward propagation network training to update the parameters of the convolutional layer.
[0110] The entity attribute extraction model is used to extract the text related to the entity in the current text as the attribute information of the entity.
[0111] Further, on the basis of the above embodiment, after the entity data, the entity relationship, the file index, and the common entity are associated and stored to form a file library, the method further includes:
[0112] Obtain a retrieval condition.
[0113] Based on the entity corresponding to the retrieval condition, obtain multiple files corresponding to the entity from the file library.
[0114] Specifically, a general retrieval method can be used for retrieval, such as Language Integrated Query (LINQ), or the method in the class of NoSQL for retrieval.
[0115] According to the digital file management method provided by the embodiment of the present invention, based on the obtained file library, multiple highly relevant associated files can be quickly obtained.
[0116] Next, the digital file management system provided by the present invention will be described. The digital file management system described below can be mutually corresponding and referred to with the digital file management method described above.
[0117] Figure 2 is a schematic structural diagram of the digital file management system provided by the present invention; as Figure 2 shown, a digital file management system includes:
[0118] The entity acquisition module 201 acquires entity data corresponding to multiple files to be processed.
[0119] In this module, using the named entity recognition algorithm, entity data contained in multiple files to be processed is extracted according to the meta database, where the meta database contains multiple entity names available for matching. Among them, the named entity recognition algorithm is a common named entity recognition algorithm, such as LSTM+CRF, CNN+CRF, BERT+(LSTM)+CRF, BiLSTM+CRF, HMM, attention model, transfer learning, etc., and the present invention is not specifically limited thereto.
[0120] More specifically, in this embodiment, the process of extracting entity data using the named entity recognition algorithm includes the steps of:
[0121] Text extraction: Extract text from files to be processed in different formats. For example, directly obtain text from word format files, and use OCR recognition to recognize text in pdf format files to obtain text.
[0122] Text tokenization: Perform tokenization processing on the text to obtain a set of words.
[0123] Entity matching: Match each word in the set of words with the entity names in the metadata, and save the words that match successfully as the entities of the files to be processed into the entity list.
[0124] Entity annotation: Annotate the entities in the entity list, and the results obtained by annotation at least include: the type of the entity, the start offset and end offset of the entity in the text.
[0125] Entity attribute establishment: Extract the text related to the entity in the text according to the entity annotation result as the attribute of the entity. Finally, the obtained entity data includes the entity and the attribute corresponding to the entity.
[0126] Construct a UML class diagram based on the obtained entities and entity attributes, and store the file information in the form of a UML class diagram.
[0127] The relationship and index acquisition module 202 determines the entity relationships between entities according to the entity data, and performs text segmentation on the files to be processed according to the entity data to obtain file indexes.
[0128] In this module, the text information in the file to be processed is modeled using an entity relationship extraction algorithm, and the semantic relationships between entities are automatically extracted from the text as entity relationships. Among them, the entity relationship extraction algorithm is a commonly used entity relationship extraction algorithm, such as supervised feature-based and kernel function-based entity relationship extraction, semi-supervised Bootstrapping, pipeline, joint learning, etc. The present invention does not limit this.
[0129] Meanwhile, the file to be processed is segmented according to the entity names to form different index units, and database index files and inverted index files are obtained according to different index units. The file index is composed of the database index file and the inverted index file and is used for subsequent document retrieval.
[0130] There is a common entity acquisition module 203, which determines the common entities corresponding to the multiple files to be processed according to the entity data.
[0131] In this module, the entity data between the files to be processed is compared, and the same entities are stored in the common entity set. It should be noted that the relationship and index acquisition module 202 and the common entity acquisition module 203 can be executed simultaneously.
[0132] The associated storage module 204 performs associated storage on the entity data, the entity relationships, the file index, and the common entities to form a file library.
[0133] In this module, the entity data, entity relationships, file index, and common entities obtained above are stored in association, and finally a fused file library is obtained. Specifically, for each entity, N:N association is performed according to the obtained entity relationships; each entity is associated with its corresponding common entity; and each entity is associated with its corresponding index.
[0134] According to the digital file management system provided by the embodiments of the present invention, by performing entity extraction, entity relationship extraction, file index establishment, and common entity acquisition on files, and performing associated storage on the data obtained above, a fused file library is obtained. What is stored in this file library is not the entire file, but the entities corresponding to the files, entity relationships, common entities, and file indexes corresponding to the entities. In the subsequent file retrieval process, multiple related files can be obtained from the fused file library through the common entities according to the retrieval conditions, thereby realizing the retrieval of the entire file library and improving the file retrieval efficiency.
[0135] Figure 3 Illustrates a schematic diagram of the entity structure of an electronic device, as Figure 3As shown in the figure, the electronic device may include: a processor 310, a communications interface 320, a memory 330, and a communication bus 340. Among them, the processor 310, the communications interface 320, and the memory 330 complete communication with each other through the communication bus 340. The processor 310 may call logic instructions in the memory 330 to execute a digital file management method, which includes: obtaining entity data corresponding to multiple files to be processed; determining entity relationships between entities according to the entity data, and performing text segmentation on the files to be processed according to the entity data to obtain file indexes; determining common entities corresponding to the multiple files to be processed according to the entity data; and associatively storing the entity data, the entity relationships, the file indexes, and the common entities to form a file library.
[0136] In addition, when the logic instructions in the above-mentioned memory 330 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0137] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the digital file management method provided by the above-mentioned various methods. The method includes: obtaining entity data corresponding to multiple files to be processed; determining entity relationships between entities according to the entity data, and performing text segmentation on the files to be processed according to the entity data to obtain file indexes; determining common entities corresponding to the multiple files to be processed according to the entity data; and associatively storing the entity data, the entity relationships, the file indexes, and the common entities to form a file library.
[0138] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the digital file management method provided above, and the method includes: obtaining entity data corresponding to multiple files to be processed; determining entity relationships between entities according to the entity data, and performing text segmentation on the files to be processed according to the entity data to obtain file indexes; determining common entities corresponding to the multiple files to be processed according to the entity data; and associatively storing the entity data, the entity relationships, the file indexes, and the common entities to form a file library.
[0139] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.
[0140] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0141] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or equivalently replace some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A digital file management method, characterized in that, Including: Obtaining entity data corresponding to multiple files to be processed; Determining entity relationships between entities according to the entity data, and performing text segmentation on the files to be processed according to the entity data to obtain file indexes; Determining common entities corresponding to the multiple files to be processed according to the entity data; Associatively storing the entity data, the entity relationships, the file indexes, and the common entities to form a file library; The obtaining entity data corresponding to multiple files to be processed includes: Extracting entities from multiple files to be processed according to a preset meta database, where the preset meta database includes file entity names; Extracting attribute information corresponding to the entities in the files to be processed; Determining a file class diagram according to the entities and the attribute information; Correspondingly, the determining entity relationships between entities according to the entity data includes: Determining entity relationships between entities according to the context information of the files to be processed and the entity data, and obtaining edges between the file class diagrams based on the entity relationships; The associatively storing the entity data, the entity relationships, the file indexes, and the common entities includes: Associatively storing the file class diagrams according to the edges between the file class diagrams.
2. The digital file management method according to claim 1, characterized in that, The performing text segmentation on the files to be processed according to the entity data to obtain file indexes includes: Segmenting the text in the files to be processed according to entity data into index units including index words and entity type words; Obtaining a data index file by using the index units, and obtaining an inverted index file based on the position information of the index words in the index units in the multiple files to be processed, where the file indexes include the data index file and the inverted index file.
3. The digital file management method according to claim 1, wherein The determining common entities corresponding to the multiple files to be processed according to the entity data includes: Determining entities common to the files to be processed according to the entity data to obtain an initial set of common entities; Performing synonym replacement on the attribute information corresponding to the entities in the initial set of common entities to obtain first attributes; Deleting entities with duplicate attributes according to the attribute information corresponding to the entities and the first attributes to obtain a final set of common entities; Correspondingly, the associatively storing the entity data, the entity relationships, the file indexes, and the common entities includes: Associatively storing entities according to the final set of common entities.
4. The digital file management method according to claim 1, characterized in that The extracting attribute information corresponding to the entities in the files to be processed includes: Using a pre-trained entity attribute extraction model to extract attribute information of the entities in the files to be processed; Wherein, the pre-trained entity attribute extraction model is a convolutional neural network model and is trained based on training files and corresponding labels.
5. The digital file management method according to claim 1, wherein The determining entity relationships between entities according to the context information of the files to be processed and the entity data includes: Performing sentence content parsing and vectorization processing on the files to be processed to obtain word vectors; Using a bidirectional LSTM to perform forward and backward context learning on the word vectors to obtain word vectors including context information; Using an attention mechanism to determine the importance of each word vector including context information in the file difference detection task to obtain a weight vector; Multiplying the word vector including context information by the weight vector to obtain a fusion result of lexical-level features and using it as sentence-level features; Classifying the sentence-level features through a classifier to obtain corresponding difference categories and using the difference categories as entity relationships.
6. The digital file management method according to any one of claims 1-5, characterized in that After the associative storage of the entity data, the entity relationships, the file index, and the common entities to form a file library, the method further includes: Obtaining a retrieval condition; Obtaining multiple files corresponding to the entity from the file library according to the entity corresponding to the retrieval condition.
7. A digital file management system, characterized in that, Including: An entity acquisition module for acquiring entity data corresponding to multiple files to be processed; A relationship and index acquisition module for determining entity relationships between entities according to the entity data and performing text segmentation on the files to be processed according to the entity data to obtain a file index; A common entity acquisition module for determining common entities corresponding to the multiple files to be processed according to the entity data; An associative storage module for associatively storing the entity data, the entity relationships, the file index, and the common entities to form a file library; The acquisition of entity data corresponding to multiple files to be processed includes: Extracting entities from multiple files to be processed according to a preset meta-database, where the preset meta-database includes file entity names; Extracting attribute information corresponding to the entity in the file to be processed; Determining a file class diagram according to the entity and the attribute information; Correspondingly, the determination of entity relationships between entities according to the entity data includes: Determining entity relationships between entities according to the context information and entity data of the file to be processed, and obtaining edges between the file class diagrams based on the entity relationships; The associative storage of the entity data, the entity relationships, the file index, and the common entities includes: Associatively storing the file class diagrams according to the edges between the file class diagrams.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the digital file management method according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it implements the digital file management method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Method and system for enhancing file entity association degree based on knowledge graph
CN111753099A
Data query system and method based on knowledge graph and terminal equipment
CN113761213A