A file information data management system and method based on big data
By introducing detection nodes and machine-selected intelligent database modes into the archive storage system, a three-party association database was built, which solved the problems of low storage efficiency, high retrieval difficulty and weak archive callability in archive data management, and realized flexible storage and association call of archive data.
Patent Information
- Application Number
- CN202510163087.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-14
AI Technical Summary
The prior art has problems in the management of archive data, such as low storage efficiency, difficulty in retrieval and weak file callability. Especially after the number of archives increases, it is difficult to achieve flexible storage and associated calls.
By introducing detection nodes and machine-selected intelligent database entry modes into the archive storage system, using computer programs to analyze and process the archive information, and building a three-party association library to realize flexible storage and association calls of archives.
It improves the storage efficiency and retrieval difficulty of archive data, enhances the callability of archives, and realizes flexible storage and associated calls of archive data.
Smart Images

Figure CN119621666B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of archive information management, and in particular to an archive information data management system and method based on big data. Background Art
[0002] Big data archive information management is a management method that uses big data technology to efficiently collect, store, process, analyze and utilize massive and diverse archive data; it improves the efficiency, accuracy and value of archive management; currently archive management has mostly shifted from traditional offline paper archives to online digital archive management, which has improved the storage limitations and retrieval difficulty of traditional paper archives. However, with the increase in the number of archives, this model still has great limitations. The digital stacking of archive data will cause deep redundancy of archive data while the requirements for storage space are becoming more stringent, which is not conducive to the flexible storage of archive data, and to a certain extent lacks the associated call storage between archive data, which is not conducive to the management of archive data, resulting in low efficiency of archive storage, great retrieval difficulty and weak archive callability; therefore, there is a lack of an efficient archive data storage management method to achieve flexible storage and associated call of archive data. Summary of the invention
[0003] The purpose of the present invention is to provide an archival information data management system and method based on big data to solve the problems raised in the prior art.
[0004] To achieve the above object, the present invention provides the following technical solutions:
[0005] A method for managing archive information data based on big data, the method comprising the following steps:
[0006] A detection node is preset based on the archive repository port, and the detection node triggers sending an inquiry event to the management port according to a preset judgment condition;
[0007] The detection node determines the archive storage mode according to the query event feedback information; the archive storage mode includes a manual storage mode and a machine-selected intelligent storage mode; the machine-selected intelligent storage mode is to perform storage analysis and processing on the archive information through a computer program;
[0008] When the archive entry mode is machine-selected intelligent entry, the machine-selected intelligent analysis model is enabled to analyze and store the archive information entered into the archive; the machine-selected intelligent analysis model performs information node analysis and processing on the archive entered into the archive; a three-party association library is constructed based on the analysis and processing results of the archive information nodes entered into the archive, and the archive entered into the archive is classified and stored and the associated archive information is called and stored.
[0009] Further, the detection node is an archive information flow detection device, which is configured with an archive information flow detection program; the archive information flow detection program judges the real-time archive information flow to be stored according to the preset judgment conditions; when the preset judgment conditions are met, an inquiry event is sent to the management port; wherein the detection node can be an actual detection device or a virtual device, and its function is to be configured with the archive information flow detection program;
[0010] The preset judgment condition is to make a threshold judgment on the number of archives to be stored in the preset period; the preset period is the average time required for manual storage of a single archive information; the threshold is manually set;
[0011] When the detection node detects that the number of files to be stored is greater than or equal to a threshold value within a preset period, an inquiry event is sent to the management port; the inquiry event is to send a message to the management port asking whether to change the file storage mode.
[0012] Furthermore, the detection node determines whether to change the archive storage mode according to the feedback information of the management port for the inquiry event;
[0013] If the management port's feedback information for the inquiry event is negative, the latent trigger mode is entered, and the archive storage mode is not changed; the latent trigger mode is used to send the inquiry event to the management port again by setting a latent trigger period; or within the latent trigger period, if the detection node receives new active feedback information from the management port as positive, the detection node determines the archive storage mode as the machine-selected intelligent storage mode; if within the latent trigger period, the detection node receives new active feedback information from the management port as negative, the archive storage mode is not changed, and the latent period is re-timed;
[0014] If the management port's feedback information for the inquiry event is yes, the detection node determines the archive storage mode as the machine-selected intelligent storage mode;
[0015] If the management port does not give feedback information to the inquiry, a timing cycle is set. If the management port still does not give feedback information when the timing cycle is met or exceeded, the detection node determines the archive storage mode as the machine-selected intelligent storage mode.
[0016] Furthermore, when the detection node determines that the archive storage mode is the machine-selected intelligent storage mode, the machine-selected intelligent analysis model is enabled to analyze and store the storage archive information;
[0017] The machine-selected intelligent analysis model retrieves the archive information to be stored, performs feature processing on the archive text, divides and summarizes the vocabulary of the archive text paragraphs, and obtains the document text feature words; performs text feature degree analysis on the archive text feature words, and determines the archive text key feature words; its calculation formula is:
[0018] ;
[0019] Among them, DC(m) is the text feature degree of the archival text feature word corresponding to sequence number m; n(m) is the word frequency of the archival text feature word corresponding to sequence number m in the archival text; N is the total word frequency of the archival text feature word; P is the total number of archives in the archive repository; p(m) is the number of archives in the archive repository that appear with the archival text feature word corresponding to sequence number m; x1 and x2 are weight coefficients;
[0020] The text feature degree analysis and calculation is performed based on the archival text feature words, and the archival text feature words are divided by setting the feature degree threshold; if the text feature degree of the archival text feature word is greater than or equal to the feature degree threshold, it is classified as the archival text key feature word; otherwise, it is classified as the archival text filling feature word; wherein, the text filling feature word refers to the vocabulary combined around the archival text key feature word to fill the archival text content;
[0021] Combine the key feature words of the archive text and the filling feature words of the archive text to perform node clustering processing, take the key feature words of the archive text as the central node, and obtain the cluster distribution of the text filling feature words as the filling node, and determine the node affinity distribution diagram of each key feature word of the archive text and the archival text filling feature word; by taking the central node as the center of the circle and setting the division radius, obtain the point cluster distribution of each central node and the corresponding affinity filling node; determine the point clusters corresponding to each central node and the corresponding filling node according to the node affinity distribution diagram, respectively, with the number of each central node and the corresponding affinity filling node, combined with the text feature degree of each central node corresponding to the key feature word of the archive text, use the bubble comparison calculation to analyze the archive text content distribution bias of adjacent point clusters, and determine the point cluster corresponding to the maximum archive text content distribution bias; its calculation formula is
[0022] ;
[0023] Among them, DB (g+1: g) is the comparison value of the archive text content distribution bias between the clusters with sequence numbers g+1 and g; DC (g+1) and DC (g) correspond to the text feature degree of the key feature words of the archive text of the central nodes of the clusters with sequence numbers g+1 and g, respectively; H (g+1) and H (g) correspond to the number of affinity filling nodes in the clusters with sequence numbers g+1 and g, respectively;
[0024] When performing bubble comparison analysis on the archive text content distribution bias comparison value of two adjacent point clusters, if the calculation result is greater than or equal to 1, the point cluster with a later order among the adjacent point clusters will be taken as the object of continued bubble comparison analysis, that is, g+1; if the calculation result is less than 1, the point cluster with a forward order among the adjacent point clusters will be taken as the object of continued bubble comparison analysis, that is, g.
[0025] Furthermore, according to the corresponding point clusters of the maximum archive text content distribution bias, the archives stored in the repository are subjected to association analysis; the node affinity distribution graphs corresponding to each archive in the repository are obtained respectively, and the corresponding point clusters of the maximum archive text content distribution bias corresponding to each archive are determined based on the archive text content distribution bias;
[0026] The feature processing is performed on the corresponding point clusters of the maximum archive text content distribution bias of the current archive and each archive in the storage repository, and the vectors are constructed with the central node of the corresponding point cluster and each affinity filling node respectively, and the vector synthesis processing is performed on the central node and affinity filling node vectors in each point cluster respectively to obtain the comprehensive feature vector of the corresponding point cluster; the correlation analysis is performed on the comprehensive feature vector of the corresponding point cluster of the maximum archive text content distribution bias of the current archive and each archive in the storage repository respectively, and the division threshold is set to determine the storage repository archive information associated with the current archive; its calculation formula is:
[0027] ;
[0028] Among them, AS (Z, Q) is the correlation between the archive corresponding to the entry sequence Z and the archive corresponding to the storage sequence Q in the storage library; Z (t) and Q (t) correspond to the maximum archive text content distribution bias corresponding to the cluster comprehensive feature vector in the archive corresponding to the entry sequence Z and the archive corresponding to the storage sequence Q in the storage library respectively;
[0029] By setting a division threshold, the files with the association analysis results greater than or equal to the threshold are judged as files that are associated with the stored files; otherwise, they are files that are not associated with the stored files.
[0030] Based on the node analysis data of the archives in the library and the association analysis data between the archives in the library and the archives stored in the library, a three-party association library is constructed to store information of the archives in the library;
[0031] The three-party association library includes a first storage library for entry sequence numbers, a second storage library for archive node information, and a third storage library for associated archive calls;
[0032] The first storage serial number storage repository is used to store the storage allocation serial number and file name information of the stored files and the storage location information of the files in the second file node information storage repository;
[0033] The second archive node information storage repository is used to store archive text feature keyword data and archive text filling feature word data corresponding to each central node and corresponding affinity filling node corresponding to the corresponding serial number archive in the first storage serial number storage repository; wherein the archive text feature keyword data and archive text filling feature word data stored in the second archive node information storage repository both record the archive text position information corresponding to each feature word;
[0034] The third associated archive call storage repository is used to store the allocation serial number and archive name information of the archives that are associated with the stored archives;
[0035] Based on the intelligent storage processing results of the archive machine selection, the storage information of the archives in the three-party associated libraries is displayed on the associated ports in real time.
[0036] A file information data management system based on big data, the system includes a file detection module, a storage mode determination module, a file information node analysis module and a correlation storage module;
[0037] The archive detection module triggers sending an inquiry event to the management port according to preset judgment conditions based on the preset detection node of the archive storage port; the storage mode determination module determines the archive storage mode according to the feedback information of the inquiry event; the archive storage mode includes a manual storage mode and a machine-selected intelligent storage mode; the machine-selected intelligent storage mode is to perform storage analysis and processing on the archive information through a computer program; when the archive storage mode is machine-selected intelligent storage, the archive information node analysis module enables the machine-selected intelligent analysis model to analyze and store the stored archive information; the machine-selected intelligent analysis model performs information node analysis and processing on the stored archives; the associated storage module constructs a three-party associated library based on the analysis and processing results of the stored archive information nodes, and classifies and stores the stored archives and calls and stores the associated archive information.
[0038] The archive detection module includes an entry detection node unit and an inquiry event sending unit;
[0039] The entry detection node unit is a file information flow detection device, which is equipped with a file information flow detection program; the file information flow detection program judges the real-time file information flow to be entered into the warehouse according to the preset judgment condition; the preset judgment condition is to make a threshold judgment on the number of files to be entered into the warehouse within a preset period; the preset period is the average time required for manual entry of a single file information; the threshold is manually set;
[0040] The inquiry event sending unit sends an inquiry event to the management port when the preset judgment condition is met; the inquiry event is to send to the management port whether to change the file storage mode.
[0041] The storage mode determination module includes a feedback information receiving unit and a storage mode judgment unit;
[0042] The feedback information receiving unit determines the change of the archive storage mode according to the feedback information of the management port for the inquiry event;
[0043] The storage mode judgment unit judges the change of storage mode for the feedback information of the management port; if the feedback information of the management port for the inquiry event is negative, the latent trigger mode is entered, and the archive storage mode is not changed; the latent trigger mode is used to send the inquiry event to the management port again by setting a latent trigger period; or within the latent trigger period, if the detection node receives new active feedback information from the management port as positive, the detection node determines the archive storage mode as the machine-selected intelligent storage mode; if within the latent trigger period, the detection node receives new active feedback information from the management port as negative, the archive storage mode is not changed, and the latent period is re-timed;
[0044] If the management port's feedback information for the inquiry event is yes, the detection node determines the archive storage mode as the machine-selected intelligent storage mode;
[0045] If the management port does not give feedback information to the inquiry, a timing cycle is set. If the management port still does not give feedback information when the timing cycle is met or exceeded, the detection node determines the archive storage mode as the machine-selected intelligent storage mode.
[0046] The archive information node analysis module includes an archive text feature analysis unit and an archive information node analysis unit;
[0047] When the archive text feature analysis unit determines that the archive storage mode is the machine-selected intelligent storage mode, the machine-selected intelligent analysis model is enabled to analyze and store the archive information to be stored; the archive information to be stored is retrieved, and the archive text is feature processed, and the archive text paragraph vocabulary is divided and summarized to obtain document text feature words; the archive text feature words are analyzed for text feature degree to determine the archive text key feature words; the archive text feature words are analyzed and calculated for text feature degree, and the archive text feature words are divided by setting a feature degree threshold; if the text feature degree of the archive text feature word is greater than or equal to the feature degree threshold, it is classified as the archive text key feature word; otherwise, it is the archive text filling feature word;
[0048] The archive information node analysis unit performs node clustering processing in combination with archive text key feature words and archive text filling feature words, takes the archive text key feature words as the central node, obtains the cluster distribution of the text filling feature words as the filling nodes, and determines the node affinity distribution diagram of each archive text key feature word and the archive text filling feature word; by taking the central node as the center of the circle and setting the division radius, obtains the point cluster distribution of each central node and the corresponding affinity filling node; determines the point clusters corresponding to each central node and the corresponding filling node according to the node affinity distribution diagram, respectively, uses the number of each central node and the corresponding affinity filling node, combined with the text feature degree of the archive text key feature words corresponding to each central node, uses bubble comparison calculation to analyze the archive text content distribution bias of adjacent point clusters, and determines the point cluster corresponding to the maximum archive text content distribution bias.
[0049] The association storage module includes an association archive analysis unit and a three-party association library;
[0050] The associated archive analysis unit performs an associated analysis on the archives stored in the storage repository according to the corresponding point clusters of the maximum archive text content distribution bias; obtains the node affinity distribution diagrams corresponding to each archive in the storage repository respectively, and analyzes based on the archive text content distribution bias to determine the corresponding point clusters of the maximum archive text content distribution bias corresponding to each archive; performs feature processing on the corresponding point clusters of the maximum archive text content distribution bias of the current stored archives and each archive in the storage repository, respectively constructs vectors with the central node of the corresponding point cluster and each affinity filling node, and respectively performs vector synthesis processing on the central node and affinity filling node vectors in each point cluster to obtain the comprehensive feature vector of the corresponding point cluster; performs an associated analysis on the comprehensive feature vector of the corresponding point cluster of the maximum archive text content distribution bias of the current stored archive and each archive in the storage repository respectively, and sets a partition threshold to determine the archive information of the storage repository that is associated with the current stored archive; by setting a partition threshold, the archive with an associated analysis result greater than or equal to the threshold is judged as an archive that is associated with the stored archive; otherwise, it is an archive that is not associated;
[0051] The three-party association library is constructed based on the node analysis data of the stored archives and the association analysis data of the stored archives and the archives stored in the storage library to store information of the stored archives;
[0052] The three-party association library includes a first storage library for entry sequence numbers, a second storage library for archive node information, and a third storage library for associated archive calls;
[0053] The first storage serial number storage repository is used to store the storage allocation serial number and file name information of the stored files and the storage location information of the files in the second file node information storage repository;
[0054] The second archive node information storage repository is used to store archive text feature keyword data and archive text filling feature word data corresponding to each central node and corresponding affinity filling node corresponding to the corresponding serial number archive in the first storage serial number storage repository; wherein the archive text feature keyword data and archive text filling feature word data stored in the second archive node information storage repository both record the archive text position information corresponding to each feature word;
[0055] The third associated archive call storage repository is used to store the allocation serial number and archive name information of the archives that are associated with the stored archives;
[0056] Based on the intelligent storage processing results of the archive machine selection, the storage information of the archives in the three-party associated libraries is displayed on the associated ports in real time.
[0057] Compared with the prior art, the present invention has the following beneficial effects:
[0058] The present invention combines archive storage and archive entry detection and analysis to realize a new mode function for archive entry storage; by setting detection nodes, the precondition is used to realize the judgment of the archive entry flow state, and based on the judgment, the query event is sent to determine the entry mode of the archive entry; based on this, the archive text node is analyzed and processed to realize the feature determination of the archive text, and a node point cluster distribution map is constructed to determine the feature bias of the archive text, and then determine the maximum feature text bias point cluster of the archive text; by analyzing the maximum feature text bias point cluster correlation between the archive entry archive and the storage library archive, the associated archive of the archive entry archive is determined; thus, a three-party associated library is established to realize the information storage management of the archive entry; the present invention improves the problems of low storage efficiency, great retrieval difficulty and weak archive callability of traditional digital archives to a certain extent, and realizes the flexible storage and associated call of archive data. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 This is a structural schematic diagram of an archive information data management system based on big data of the present invention;
[0060] Figure 2 The present invention is a flowchart of an archival information data management method based on big data. DETAILED DESCRIPTION
[0061] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0062] Example: Figure 1 As shown, the present invention provides a technical solution:
[0063] An archive information data management system based on big data, comprising an archive detection module, a storage mode determination module, an archive information node analysis module and an associated storage module;
[0064] Among them, the archive detection module triggers sending an inquiry event to the management port based on the preset detection node of the archive repository port according to the preset judgment conditions;
[0065] The storage mode determination module determines the archive storage mode according to the query event feedback information; the archive storage mode includes manual storage mode and machine selection intelligent storage mode;
[0066] The machine-selected intelligent storage mode is to analyze and process the archive information through computer programs; when the archive storage mode is machine-selected intelligent storage, the archive information node analysis module enables the machine-selected intelligent analysis model to analyze and store the archive information; the machine-selected intelligent analysis model performs information node analysis on the archives;
[0067] The associated storage module builds a three-party associated library based on the analysis and processing results of the incoming archive information nodes, classifies and stores the incoming archives, and calls and stores the associated archive information.
[0068] The archive detection module includes an entry detection node unit and an inquiry event sending unit;
[0069] The storage detection node unit is a file information flow detection device, which is equipped with a file information flow detection program; the file information flow detection program judges the real-time file information flow to be stored according to the preset judgment condition; the preset judgment condition is the threshold judgment of the number of files to be stored within the preset period; the preset period is the average time required for manual storage of a single file information; the threshold is set manually;
[0070] When the preset judgment conditions are met, the inquiry event sending unit sends an inquiry event to the management port; the inquiry event is to send to the management port whether to change the file storage mode.
[0071] The storage mode determination module includes a feedback information receiving unit and a storage mode determination unit;
[0072] The feedback information receiving unit determines the change of the archive storage mode according to the feedback information of the management port for the inquiry event;
[0073] The storage mode judgment unit judges the change of storage mode for the feedback information of the management port; if the feedback information of the management port for the inquiry event is negative, the latent trigger mode is entered, and the archive storage mode is not changed; the latent trigger mode is used to send the inquiry event to the management port again by setting the latent trigger period; or within the latent trigger period, if the detection node receives the new active feedback information of the management port as yes, the detection node determines the archive storage mode as the machine-selected intelligent storage mode; if within the latent trigger period, the detection node receives the new active feedback information of the management port as negative, the archive storage mode is not changed, and the latent period is re-timed;
[0074] If the management port's feedback information for the inquiry event is yes, the detection node determines the archive storage mode as the machine-selected intelligent storage mode;
[0075] If the management port does not give any feedback information to the inquiry, the timing cycle is set. If the management port still does not give any feedback information when the timing cycle is met or exceeded, the detection node determines the archive storage mode as the machine-selected intelligent storage mode.
[0076] The archive information node analysis module includes an archive text feature analysis unit and an archive information node analysis unit;
[0077] When the archive text feature analysis unit determines that the archive storage mode is the machine-selected intelligent storage mode, the machine-selected intelligent analysis model is enabled to analyze and store the archive information to be stored; the archive information to be stored is retrieved, and the archive text is feature processed, and the archive text paragraph vocabulary is divided and summarized to obtain the document text feature words; the archive text feature words are analyzed for text feature degree to determine the archive text key feature words; the text feature degree analysis and calculation are performed based on the archive text feature words, and the archive text feature words are divided by setting the feature degree threshold; if the text feature degree of the archive text feature word is greater than or equal to the feature degree threshold, it is classified as the archive text key feature word; otherwise, it is the archive text filling feature word;
[0078] The archive information node analysis unit performs node clustering processing in combination with the archive text key feature words and the archive text filling feature words, takes the archive text key feature words as the central node, obtains the cluster distribution of the text filling feature words as the filling nodes, and determines the node affinity distribution diagram of each archive text key feature word and the archive text filling feature word; by taking the central node as the center of the circle and setting the division radius, obtains the point cluster distribution of each central node and the corresponding affinity filling node; determines the point clusters corresponding to each central node and the corresponding filling node according to the node affinity distribution diagram, respectively, uses the number of each central node and the corresponding affinity filling node, combined with the text feature degree of the archive text key feature words corresponding to each central node, uses the bubble comparison calculation to analyze the archive text content distribution bias of adjacent point clusters, and determines the point cluster corresponding to the maximum archive text content distribution bias.
[0079] The associated storage module includes an associated archive analysis unit and a three-party associated library;
[0080] The associated archive analysis unit performs an associated analysis on the archives stored in the storage repository according to the corresponding point clusters of the maximum archive text content distribution bias; obtains the node affinity distribution diagrams corresponding to each archive in the storage repository respectively, and analyzes based on the archive text content distribution bias to determine the corresponding point clusters of the maximum archive text content distribution bias corresponding to each archive; performs feature processing on the corresponding point clusters of the maximum archive text content distribution bias of the current stored archives and each archive in the storage repository, respectively constructs vectors with the central node of the corresponding point cluster and each affinity filling node, and respectively performs vector synthesis processing on the central node and affinity filling node vectors in each point cluster to obtain the comprehensive feature vector of the corresponding point cluster; performs an associated analysis on the comprehensive feature vector of the corresponding point clusters of the maximum archive text content distribution bias of the current stored archives and each archive in the storage repository, and sets a partition threshold to determine the archive information of the storage repository that is associated with the current stored archive; by setting a partition threshold, the archives with an associated analysis result greater than or equal to the threshold are judged as archives that are associated with the stored archives; otherwise, they are archives that are not associated;
[0081] The tripartite association library is constructed based on the node analysis data of the archives in the library and the association analysis data between the archives in the library and the archives stored in the library to store information of the archives in the library;
[0082] The three-party association library includes a first storage library for entry sequence numbers, a second storage library for archive node information, and a third storage library for associated archive calls;
[0083] The first storage serial number storage repository is used to store the storage allocation serial number and file name information of the stored files and the storage location information of the files in the second file node information storage repository;
[0084] The second archive node information storage library is used to store archive text feature keyword data and archive text filling feature word data corresponding to each central node and corresponding affinity filling node corresponding to the corresponding serial number archive in the first storage serial number storage library; wherein the archive text feature keyword data and archive text filling feature word data stored in the second archive node information storage library both record the archive text position information corresponding to each feature word;
[0085] The third associated archive call storage repository is used to store the allocation sequence number and archive name information of the archives that are associated with the stored archives;
[0086] Based on the intelligent storage processing results of the archive machine selection, the storage information of the archives in the three-party associated library is displayed on the associated port in real time;
[0087] like Figure 2 As shown, the present invention provides another technical solution:
[0088] A method for managing archive information data based on big data, the method comprising the following steps:
[0089] Based on the preset detection node of the archive storage port, the detection node triggers the sending of an inquiry event to the management port according to the preset judgment condition; wherein the detection node is an archive information flow detection device, which is configured with an archive information flow detection program; wherein the archive information flow detection program judges the real-time archive information flow to be stored according to the preset judgment condition, and sends an inquiry event to the management port when the preset judgment condition is met; wherein the inquiry event is to send to the management port whether to change the archive storage mode;
[0090] The detection node determines the archive storage mode based on the query event feedback information;
[0091] The archive storage modes include manual storage mode and machine-selected intelligent storage mode;
[0092] Among them, the machine-selected intelligent storage mode is to use computer programs to store and analyze the archive information;
[0093] According to the feedback information of the management port on the inquiry event, the file storage mode change is determined;
[0094] If the management port's feedback information for the inquiry event is negative, the latent trigger mode is entered, and the archive storage mode is not changed; the latent trigger mode is used to send the inquiry event to the management port again by setting a latent trigger period; or within the latent trigger period, if the detection node receives new active feedback information from the management port as positive, the detection node determines the archive storage mode as the machine-selected intelligent storage mode; if within the latent trigger period, the detection node receives new active feedback information from the management port as negative, the archive storage mode is not changed, and the latent period is re-timed;
[0095] If the management port's feedback information for the inquiry event is yes, the detection node determines the archive storage mode as the machine-selected intelligent storage mode;
[0096] If the management port does not give feedback information to the inquiry, then by setting a timing cycle, if the management port still does not give feedback information when the timing cycle is met or exceeded, the detection node determines the archive storage mode as the machine-selected intelligent storage mode;
[0097] When the archive storage mode is machine-selected intelligent storage, the machine-selected intelligent analysis model is enabled to analyze and store the archive information in the storage;
[0098] The archive information to be stored is retrieved, and the archive text is processed by features, and the vocabulary of the archive text paragraphs is divided and summarized to obtain the document text feature words; the text feature degree of the archive text feature words is analyzed to determine the key feature words of the archive text; the calculation formula is:
[0099] ;
[0100] Perform text feature degree analysis and calculation based on archival text feature words, and classify archival text feature words by setting feature degree thresholds; if the text feature degree of an archival text feature word is greater than or equal to the feature degree threshold, it is classified as an archival text key feature word; otherwise, it is classified as an archival text filling feature word;
[0101] The machine-selected intelligent analysis model performs information node analysis and processing on the archived files;
[0102] Combine the key feature words of the archive text and the filling feature words of the archive text to perform node clustering processing, take the key feature words of the archive text as the central node, and obtain the cluster distribution of the text filling feature words as the filling node, and determine the node affinity distribution diagram of each key feature word of the archive text and the archival text filling feature word; by taking the central node as the center of the circle and setting the division radius, obtain the point cluster distribution of each central node and the corresponding affinity filling node; determine the point clusters corresponding to each central node and the corresponding filling node according to the node affinity distribution diagram, respectively, with the number of each central node and the corresponding affinity filling node, combined with the text feature degree of each central node corresponding to the key feature word of the archive text, use the bubble comparison calculation to analyze the archive text content distribution bias of adjacent point clusters, and determine the point cluster corresponding to the maximum archive text content distribution bias; its calculation formula is
[0103] ;
[0104] When performing bubble comparison analysis on the archive text content distribution bias comparison value for two adjacent point clusters, if the calculation result is greater than or equal to 1, the point cluster with the later order in the adjacent point clusters will be the object of continued bubble comparison analysis; if the calculation result is less than 1, the point cluster with the earlier order in the adjacent point clusters will be the object of continued bubble comparison analysis;
[0105] Based on the analysis and processing results of the archive information nodes, a tripartite association library is constructed to classify and store the archives and call and store the associated archive information;
[0106] According to the corresponding point clusters of the maximum archive text content distribution bias, the archives stored in the repository are analyzed for association; the node affinity distribution graphs corresponding to each archive in the repository are obtained respectively, and the corresponding point clusters of the maximum archive text content distribution bias corresponding to each archive are determined based on the archive text content distribution bias;
[0107] The feature processing is performed on the corresponding point clusters of the maximum archive text content distribution bias of the current archive and each archive in the storage repository, and the vectors are constructed with the central node of the corresponding point cluster and each affinity filling node respectively, and the vector synthesis processing is performed on the central node and affinity filling node vectors in each point cluster respectively to obtain the comprehensive feature vector of the corresponding point cluster; the correlation analysis is performed on the comprehensive feature vector of the corresponding point cluster of the maximum archive text content distribution bias of the current archive and each archive in the storage repository respectively, and the division threshold is set to determine the storage repository archive information associated with the current archive; its calculation formula is:
[0108] ;
[0109] By setting a division threshold, the files with the association analysis results greater than or equal to the threshold are judged as files that are associated with the stored files; otherwise, they are files that are not associated with the stored files.
[0110] Based on the node analysis data of the archives in the library and the association analysis data between the archives in the library and the archives stored in the library, a three-party association library is constructed to store information of the archives in the library;
[0111] The three-party association library includes a first storage library for entry sequence numbers, a second storage library for archive node information, and a third storage library for associated archive calls;
[0112] The first storage serial number storage repository is used to store the storage allocation serial number and file name information of the stored files and the storage location information of the files in the second file node information storage repository;
[0113] The second archive node information storage library is used to store archive text feature keyword data and archive text filling feature word data corresponding to each central node and corresponding affinity filling node corresponding to the corresponding serial number archive in the first storage serial number storage library; wherein the archive text feature keyword data and archive text filling feature word data stored in the second archive node information storage library both record the archive text position information corresponding to each feature word;
[0114] The third associated archive call storage repository is used to store the allocation serial number and archive name information of archives that are associated with the stored archives; based on the intelligent storage processing results of the stored archives, the archive storage information of the three-party associated library is displayed on the associated port in real time;
[0115] The calling method for the three-party associated library is to extract the archive serial number and name information from the first entry serial number storage library, and retrieve the archive node information in the second archive node information storage library with the extracted information, and restore the archive text content based on the retrieved archive text node information; according to the archive text content, the archive is assigned a serial number and an archive name to the associated archive in the third associated archive calling storage library.
[0116] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.
Claims
1. A method for managing archival information data based on big data, characterized in that: The method comprises the following steps: A detection node is preset based on the archive repository port, and the detection node triggers sending an inquiry event to the management port according to a preset judgment condition; The detection node determines the archive storage mode according to the query event feedback information; the archive storage mode includes a manual storage mode and a machine-selected intelligent storage mode; the machine-selected intelligent storage mode is to perform storage analysis and processing on the archive information through a computer program; When the archive storage mode is machine-selected intelligent storage, the machine-selected intelligent analysis model is enabled to analyze and store the archive information in the storage; the machine-selected intelligent analysis model performs information node analysis and processing on the archive in the storage; a tripartite association library is constructed based on the analysis and processing results of the archive information nodes in the storage, and the archive in the storage is classified and stored and the associated archive information is called and stored; Among them, when the archive storage mode is the machine-selected intelligent storage mode, the machine-selected intelligent analysis model is enabled to analyze the archive information in the storage. The specific analysis is as follows: The machine-selected intelligent analysis model retrieves the archive information to be stored, performs feature processing on the archive text, divides and summarizes the vocabulary of the archive text paragraphs, and obtains the document text feature words; performs text feature degree analysis on the archive text feature words, and determines the archive text key feature words; its calculation formula is: ; Among them, DC(m) is the text feature degree of the archival text feature word corresponding to sequence number m; n(m) is the word frequency of the archival text feature word corresponding to sequence number m in the archival text; N is the total word frequency of the archival text feature word; P is the total number of archives in the archive repository; p(m) is the number of archives in the archive repository that appear with the archival text feature word corresponding to sequence number m; x1 and x2 are weight coefficients; Perform text feature degree analysis and calculation based on archival text feature words, and classify archival text feature words by setting feature degree thresholds; if the text feature degree of an archival text feature word is greater than or equal to the feature degree threshold, it is classified as an archival text key feature word; otherwise, it is classified as an archival text filling feature word; Combine the key feature words of the archive text and the filling feature words of the archive text to perform node clustering processing, take the key feature words of the archive text as the central node, and obtain the cluster distribution of the text filling feature words as the filling node, and determine the node affinity distribution diagram of each key feature word of the archive text and the archival text filling feature word; by taking the central node as the center of the circle and setting the division radius, obtain the point cluster distribution of each central node and the corresponding affinity filling node; determine the point clusters corresponding to each central node and the corresponding filling node according to the node affinity distribution diagram, respectively, with the number of each central node and the corresponding affinity filling node, combined with the text feature degree of each central node corresponding to the key feature word of the archive text, use the bubble comparison calculation to analyze the archive text content distribution bias of adjacent point clusters, and determine the point cluster corresponding to the maximum archive text content distribution bias; its calculation formula is ; Among them, DB (g+1: g) is the comparison value of the archive text content distribution bias between the clusters with sequence numbers g+1 and g; DC (g+1) and DC (g) correspond to the text feature degree of the key feature words of the archive text of the central nodes of the clusters with sequence numbers g+1 and g, respectively; H (g+1) and H (g) correspond to the number of affinity filling nodes in the clusters with sequence numbers g+1 and g, respectively; When performing bubble comparison analysis on the archive text content distribution bias comparison value of two adjacent point clusters, if the calculation result is greater than or equal to 1, the point cluster with a later order among the adjacent point clusters will be the object of continued bubble comparison analysis; if the calculation result is less than 1, the point cluster with a forward order among the adjacent point clusters will be the object of continued bubble comparison analysis.
2. The archival information data management method based on big data according to claim 1 is characterized in that: The detection node is a file information flow detection device, which is equipped with a file information flow detection program; the file information flow detection program judges the real-time file information flow to be stored according to the preset judgment conditions; when the preset judgment conditions are met, an inquiry event is sent to the management port; The preset judgment condition is to make a threshold judgment on the number of archives to be stored in the preset period; the preset period is the average time required for manual storage of a single archive information; the threshold is manually set; When the detection node detects that the number of files to be stored is greater than or equal to a threshold value within a preset period, an inquiry event is sent to the management port; the inquiry event is to send a message to the management port asking whether to change the file storage mode.
3. The archival information data management method based on big data according to claim 2 is characterized in that: The detection node determines the change of the archive storage mode according to the feedback information of the management port for the inquiry event; If the management port's feedback information for the inquiry event is negative, the latent trigger mode is entered, and the archive storage mode is not changed; the latent trigger mode is used to send the inquiry event to the management port again by setting a latent trigger period; or within the latent trigger period, if the detection node receives new active feedback information from the management port as positive, the detection node determines the archive storage mode as the machine-selected intelligent storage mode; if within the latent trigger period, the detection node receives new active feedback information from the management port as negative, the archive storage mode is not changed, and the latent period is re-timed; If the management port's feedback information for the inquiry event is yes, the detection node determines the archive storage mode as the machine-selected intelligent storage mode; If the management port does not give feedback information to the inquiry, a timing cycle is set. If the management port still does not give feedback information when the timing cycle is met or exceeded, the node determines the archive storage mode as the machine-selected intelligent storage mode.
4. The archival information data management method based on big data according to claim 3 is characterized in that: According to the corresponding point clusters of the maximum archive text content distribution bias, the archives stored in the repository are analyzed for association; the node affinity distribution graphs corresponding to each archive in the repository are obtained respectively, and the corresponding point clusters of the maximum archive text content distribution bias corresponding to each archive are determined based on the archive text content distribution bias; Perform feature processing on the corresponding point clusters of the maximum archive text content distribution bias of the current archive and each archive in the storage repository, respectively construct vectors with the central node of the corresponding point cluster and each affinity filling node, and respectively perform vector synthesis processing on the central node and affinity filling node vectors in each point cluster to obtain the comprehensive feature vector of the corresponding point cluster; respectively perform association analysis on the comprehensive feature vectors of the corresponding point clusters of the maximum archive text content distribution bias of the current archive and each archive in the storage repository, and set a division threshold to determine the storage repository archive information associated with the current archive; The calculation formula is: ; Among them, AS (Z, Q) is the correlation between the archive corresponding to the entry sequence Z and the archive corresponding to the storage sequence Q in the storage library; Z (t) and Q (t) correspond to the maximum archive text content distribution bias corresponding to the cluster comprehensive feature vector in the archive corresponding to the entry sequence Z and the archive corresponding to the storage sequence Q in the storage library respectively; By setting a division threshold, the files with the association analysis results greater than or equal to the threshold are judged as files that are associated with the stored files; otherwise, they are files that are not associated with the stored files. Based on the node analysis data of the archives in the library and the association analysis data between the archives in the library and the archives stored in the library, a three-party association library is constructed to store information of the archives in the library; The three-party association library includes a first storage library for entry sequence numbers, a second storage library for archive node information, and a third storage library for associated archive calls; The first storage serial number storage repository is used to store the storage allocation serial number and file name information of the stored files and the storage location information of the files in the second file node information storage repository; The second archive node information storage repository is used to store archive text feature keyword data and archive text filling feature word data corresponding to each central node and corresponding affinity filling node corresponding to the corresponding serial number archive in the first storage serial number storage repository; wherein the archive text feature keyword data and archive text filling feature word data stored in the second archive node information storage repository both record the archive text position information corresponding to each feature word; The third associated archive call storage repository is used to store the allocation serial number and archive name information of the archives that are associated with the stored archives; Based on the intelligent storage processing results of the archive machine selection, the storage information of the archives in the three-party associated libraries is displayed on the associated ports in real time.
5. An archive information data management system based on big data, applying an archive information data management method based on big data as described in any one of claims 1 to 4, characterized in that: The system includes an archive detection module, a storage mode determination module, an archive information node analysis module and an associated storage module; The archive detection module triggers sending an inquiry event to the management port according to the preset judgment conditions based on the preset detection node of the archive storage port; the storage mode determination module determines the archive storage mode according to the feedback information of the inquiry event; the archive storage mode includes a manual storage mode and a machine-selected intelligent storage mode; the machine-selected intelligent storage mode is to perform storage analysis and processing on the archive information through a computer program; when the archive storage mode is the machine-selected intelligent storage, the archive information node analysis module enables the machine-selected intelligent analysis model to analyze and store the storage archive information; the machine-selected intelligent analysis model performs information node analysis and processing on the storage archive; The associated storage module constructs a three-party associated library based on the analysis and processing results of the stored archive information nodes, and classifies and stores the stored archives and calls and stores the associated archive information.
6. The archival information data management system based on big data according to claim 5 is characterized in that: The archive detection module includes an entry detection node unit and an inquiry event sending unit; The entry detection node unit is a file information flow detection device, which is equipped with a file information flow detection program; the file information flow detection program judges the real-time file information flow to be entered into the warehouse according to the preset judgment condition; the preset judgment condition is to make a threshold judgment on the number of files to be entered into the warehouse within a preset period; the preset period is the average time required for manual entry of a single file information; the threshold is manually set; The inquiry event sending unit sends an inquiry event to the management port when the preset judgment condition is met; the inquiry event is to send to the management port whether to change the file storage mode.
7. The archival information data management system based on big data according to claim 6 is characterized in that: The storage mode determination module includes a feedback information receiving unit and a storage mode judgment unit; The feedback information receiving unit determines the change of the archive storage mode according to the feedback information of the management port for the inquiry event; The storage mode determination unit determines the storage mode change based on the management port feedback information; If the management port's feedback information for the inquiry event is negative, the latent trigger mode is entered, and the archive storage mode is not changed; the latent trigger mode is used to send the inquiry event to the management port again by setting a latent trigger period; or within the latent trigger period, if the detection node receives new active feedback information from the management port as positive, the detection node determines the archive storage mode as the machine-selected intelligent storage mode; if within the latent trigger period, the detection node receives new active feedback information from the management port as negative, the archive storage mode is not changed, and the latent period is re-timed; If the management port's feedback information for the inquiry event is yes, the detection node determines the archive storage mode as the machine-selected intelligent storage mode; If the management port does not give feedback information to the inquiry, a timing cycle is set. If the management port still does not give feedback information when the timing cycle is met or exceeded, the detection node determines the archive storage mode as the machine-selected intelligent storage mode.
8. The archival information data management system based on big data according to claim 7 is characterized in that: The archive information node analysis module includes an archive text feature analysis unit and an archive information node analysis unit; When the archive text feature analysis unit determines that the archive storage mode is the machine-selected intelligent storage mode, the machine-selected intelligent analysis model is enabled to analyze and store the archive information to be stored; the archive information to be stored is retrieved, and the archive text is feature processed, and the archive text paragraph vocabulary is divided and summarized to obtain document text feature words; the archive text feature words are analyzed for text feature degree to determine the archive text key feature words; the archive text feature words are analyzed and calculated for text feature degree, and the archive text feature words are divided by setting a feature degree threshold; if the text feature degree of the archive text feature word is greater than or equal to the feature degree threshold, it is classified as the archive text key feature word; otherwise, it is the archive text filling feature word; The archive information node analysis unit performs node clustering processing in combination with archive text key feature words and archive text filling feature words, takes the archive text key feature words as the central node, obtains the cluster distribution of the text filling feature words as the filling nodes, and determines the node affinity distribution diagram of each archive text key feature word and the archive text filling feature word; by taking the central node as the center of the circle and setting the division radius, obtains the point cluster distribution of each central node and the corresponding affinity filling node; determines the point clusters corresponding to each central node and the corresponding filling node according to the node affinity distribution diagram, respectively, uses the number of each central node and the corresponding affinity filling node, combined with the text feature degree of the archive text key feature words corresponding to each central node, uses bubble comparison calculation to analyze the archive text content distribution bias of adjacent point clusters, and determines the point cluster corresponding to the maximum archive text content distribution bias.
9. The archival information data management system based on big data according to claim 8 is characterized in that: The association storage module includes an association archive analysis unit and a three-party association library; The associated archive analysis unit performs an associated analysis on the archives stored in the repository according to the corresponding point clusters of the maximum archive text content distribution bias; obtains the node affinity distribution graphs corresponding to each archive in the repository respectively, and analyzes based on the archive text content distribution bias to determine the corresponding point clusters of the maximum archive text content distribution bias corresponding to each archive; The feature processing is performed on the corresponding point clusters of the maximum archive text content distribution bias of the current archive and each archive in the storage repository, and the vectors are constructed with the central node of the corresponding point cluster and each affinity filling node respectively, and the vector synthesis processing is performed on the central node and the affinity filling node vectors in each point cluster respectively to obtain the comprehensive feature vector of the corresponding point cluster; the association analysis is performed on the comprehensive feature vector of the corresponding point cluster of the maximum archive text content distribution bias of the current archive and each archive in the storage repository respectively, and a partition threshold is set to determine the archive information of the storage repository that is associated with the current archive; by setting a partition threshold, the archive with the association analysis result greater than or equal to the threshold is judged as an archive that is associated with the archive; otherwise, it is an archive that is not associated; The three-party association library is constructed based on the node analysis data of the stored archives and the association analysis data of the stored archives and the archives stored in the storage library to store information of the stored archives; The three-party association library includes a first storage library for entry sequence numbers, a second storage library for archive node information, and a third storage library for associated archive calls; The first storage serial number storage repository is used to store the storage allocation serial number and file name information of the stored files and the storage location information of the files in the second file node information storage repository; The second archive node information storage repository is used to store archive text feature keyword data and archive text filling feature word data corresponding to each central node and corresponding affinity filling node corresponding to the corresponding serial number archive in the first storage serial number storage repository; wherein the archive text feature keyword data and archive text filling feature word data stored in the second archive node information storage repository both record the archive text position information corresponding to each feature word; The third associated archive call storage repository is used to store the allocation serial number and archive name information of the archives that are associated with the stored archives; Based on the intelligent storage processing results of the archive machine selection, the storage information of the archives in the three-party associated libraries is displayed on the associated ports in real time.
Citation Information
Patent Citations
Electronic process based filing and modeling method
CN105894142A
Data association analysis method and system for drug document
CN111353004A