Intelligent archival repository RFID archive processing system
Through the RFID archive processing system of the smart archive warehouse, multi-dimensional analysis and intelligent archiving configuration are used to solve the problem of unreasonable archive archiving in the existing technology, and the efficiency of archive retrieval and management is improved.
Patent Information
- Application Number
- CN202510367808.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-05-16
AI Technical Summary
The existing archive warehouse management system lacks intelligence and flexibility in archive archiving strategies, resulting in unreasonable archive archiving and affecting the efficiency of archive retrieval and management.
It provides a smart archive warehouse RFID archive processing system, through the status generation module, semantic clustering module, personnel tag clustering module, frequency statistics module and archive configuration module, realize multi-dimensional analysis of archive content, personnel and access frequency, and perform intelligent archive configuration.
Through multi-level intelligent archive analysis and archiving, the rationality of archive archiving is improved, ensuring that high-frequency archives are stored in the optimal position, and improving the efficiency of archive retrieval and management.
Smart Images

Figure CN120013488A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of archive data processing, and in particular to an intelligent archive warehouse RFID archive processing system. Background Art
[0002] With the rapid development of informatization, archive management is also moving towards digitalization. The current archive warehouse management has widely adopted automated systems, which can realize the basic information entry, location recording and retrieval functions of archives through technologies such as barcodes, thereby improving the efficiency and accuracy of archive management and reducing manual intervention and errors.
[0003] However, the current archive warehouse management system still has limitations in terms of archive archiving strategies. The existing archive archiving methods are usually carried out according to preset fixed rules, such as simple classification, numbering order or time order. Although this method has achieved automation, it lacks intelligence and flexibility. In particular, the lack of analysis of multi-dimensional characteristics such as archive content, users and frequency of use has led to unreasonable archive archiving. For example, related archives that are often used by the same group of people may be stored in different areas of the warehouse; archives with high frequency of use are not given priority in locations that are easy to access.
[0004] This kind of archiving method that lacks consideration of archive characteristics directly leads to irrational archiving, which in turn affects the efficiency of archive retrieval and management. Even in archive management systems with a high degree of automation, this problem still exists, especially in units with a large number of archives, which affects the overall work efficiency. Summary of the invention
[0005] The present invention aims to solve the technical problem that the existing archive filing lacks consideration of archive characteristics, resulting in unreasonable archive filing and further leading to low efficiency of archive retrieval and management, and provides an intelligent archive warehouse RFID archive processing system to solve the problem.
[0006] In order to solve the above technical problems, the present invention provides an RFID archive processing system for a smart archive warehouse, comprising: a status generation module, for generating archive access status information in response to RFID tag location record information, wherein the archive access status information includes archive content identification and archive personnel identification; a semantic clustering module, for semantically clustering the archive set according to the archive content identification, and obtaining a first-level archive clustering result; a personnel label clustering module, for traversing the first-level archive clustering result to perform personnel label clustering according to the archive personnel identification, and obtaining a second-level archive clustering result; a frequency statistics module, for traversing the second-level archive clustering result, and based on the archive access log, counting the archive access frequency set; an archiving configuration module, for traversing the second-level archive clustering result according to the archive to be stored, and performing classification configuration to obtain an attribution cluster, wherein the attribution cluster has a category access frequency in the archive access frequency set, and if the category access frequency is greater than or equal to the access frequency threshold, it is archived in a pre-identified preset warehouse area.
[0007] Optionally, the status generation module includes: a tag acquisition unit, used to obtain an RFID tag according to the RFID tag location record information; an archive data acquisition unit, used to obtain archive data uniquely associated with the RFID tag based on archive archiving records, wherein the archive data stores the archive content identifier and the archive personnel identifier when archiving.
[0008] Optionally, the semantic clustering module includes: a first content identification unit, used to obtain a first archive content identification of the archive content identification, wherein the first archive content identification includes a first content attribute set; a second content identification unit, used to obtain a second archive content identification of the archive content identification, wherein the second archive content identification includes a second content attribute set; a semantic similarity calculation unit, used to calculate the ratio of the number of intersections and unions of the first content attribute set and the second content attribute set, set it as a first semantic similarity, and add it to the semantic similarity set; a semantic clustering execution unit, used to semantically cluster the archive set based on a semantic similarity threshold and in combination with the semantic similarity set to obtain the first-level archive clustering result.
[0009] Optionally, the personnel label clustering module includes: a first personnel identification unit, used to obtain a first archive personnel identification of the archive personnel identification, wherein the first archive personnel identification includes a first personnel position type set; a second personnel identification unit, used to obtain a second archive personnel identification of the archive personnel identification, wherein the second archive personnel identification includes a second personnel position type set; a personnel similarity calculation unit, used to calculate the ratio of the number of intersections and unions of the first personnel position type set and the second personnel position type set, set it as a first personnel similarity, and add it to the personnel similarity set; a personnel label clustering execution unit, used to perform personnel label clustering on the archive set based on a personnel similarity threshold and in combination with the personnel similarity set to obtain the secondary archive clustering result.
[0010] Optionally, the archiving configuration module includes: a scoring function configuration unit, configured to configure a warehouse shelf priority scoring function: ,in, Represents the warehouse shelf priority, Represents the path distance from the shelf to the shelf's associated access position, Characterizes the path distance between the shelf-related access location and the door, and Characterize the first weight and the second weight; a priority scoring unit, used to traverse the warehouse shelves for priority scoring according to the warehouse shelf priority scoring function, and obtain a priority scoring set; a warehouse shelf clustering unit, used to cluster the warehouse shelves according to the priority scoring threshold, and obtain a warehouse shelf clustering result, wherein the warehouse shelf clustering result has a centroid priority scoring identifier; a clustering result sorting unit, used to sort the warehouse shelf clustering results from small to large according to the centroid priority scoring identifier, and obtain a warehouse shelf clustering sorting result; a pre-marking unit, used to select vacant warehouse areas for pre-marking according to the warehouse shelf clustering sorting result.
[0011] Optionally, the pre-identification unit includes: a file sorting subunit, which is used to sort the files to be stored from large to small according to the category access frequency when there are multiple files to be stored whose category access frequency is greater than or equal to the access frequency threshold, and obtain the sorting result of the files to be stored; a shelf sorting subunit, which is used to extract the category sorting result of the vacant shelves according to the cluster sorting result of the warehouse shelves; and a fitness function subunit, which is used to construct a storage fitness function: ,in, Characterize the fitness of the storage solution, Indicates the category number of the j-th pre-occupied shelf, represents the category of the kth empty shelf, H represents the total number of pre-occupied shelves, M represents the total number of empty shelves, The sequence number representing the i-th file to be stored in the sorting result of the files to be stored. The serial number of the pre-occupied shelf for the i-th file to be stored in the empty shelf category sorting result, N represents the number of files to be stored; the storage plan acquisition subunit is used to randomly place the sorting results of the files to be stored into the empty shelves to obtain several storage plans; the minimum value sorting subunit is used to perform minimum value sorting on the several storage plans according to the storage fitness function to obtain the empty warehouse area for pre-marking.
[0012] Optionally, the archiving configuration module also includes: a similarity evaluation unit, which is used to perform similarity evaluation of semantics and personnel labels with the centroid archive of each category of the secondary archive clustering result based on the archive to be stored, and perform mean calculation to obtain a similarity set; an attribution cluster acquisition unit, which is used to extract the class corresponding to the maximum similarity according to the similarity set, and set it as the attribution cluster.
[0013] The beneficial effects of the present invention are: The status generation module is used to generate archive access status information in response to the RFID tag location record information, wherein the archive access status information includes the archive content identification and the archive personnel identification; through the status generation module, the system can automatically capture and record the archive access situation and related personnel information, laying a data foundation for subsequent intelligent analysis; through the semantic clustering module, the archive collection is semantically clustered according to the archive content identification, and the first-level archive clustering result is obtained, so that the system can perform preliminary classification based on the semantic characteristics of the archive content, classify the content-related archives into the same category, and realize the content relevance analysis of the archives; through the personnel label clustering module, according to the archive personnel identification, the first-level archive clustering result is traversed to perform personnel label clustering, and the second-level archive clustering result is obtained, so that on the basis of content classification, the user factor is further considered, and the same category of archives are clustered. The files that are frequently used by the same personnel in the case are grouped in more detail, which enhances the accuracy of file classification; the frequency statistics module traverses the secondary file clustering results, and based on the file access log, the file access frequency set is counted, thereby analyzing the file usage frequency characteristics, providing a basis for subsequent optimization of the archiving location, so that the frequently used files can be placed in a location that is more convenient for access; the archiving configuration module traverses the secondary file clustering results according to the files to be stored, and classifies and configures them to obtain the belonging clusters, wherein the belonging clusters have a category access frequency in the file access frequency set. If the category access frequency is greater than or equal to the access frequency threshold, the files are archived in a pre-identified preset warehouse area, thereby realizing the intelligent configuration of the file storage location, ensuring that the frequently used file categories are stored in the optimal location, and improving the efficiency of file retrieval and management.
[0014] Through the above technical scheme, a multi-dimensional analysis of archive content, archive personnel and archive access frequency is achieved, thereby improving the rationality of archive filing, thereby achieving the technical effect of improving the efficiency of archive retrieval and management. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 A schematic diagram of the structure of an intelligent archive warehouse RFID archive processing system provided by the present invention; Figure 2 A schematic diagram of the structure of the semantic clustering module provided by the present invention; Figure 3 This is a schematic diagram of the structure of the personnel label clustering module provided by the present invention.
[0016] In the accompanying drawings, the components represented by the reference numerals are as follows: A state generation module 11, a semantic clustering module 12, a personnel tag clustering module 13, a frequency statistics module 14 and an archive configuration module 15. DETAILED DESCRIPTION
[0017] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0018] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0019] In the description of the present invention, the term "for example" is used to mean "used as an example, illustration or explanation". Any embodiment described as "for example" in the present invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is given to enable any technician in the field to implement and use the present invention. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other examples, well-known structures and processes will not be elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed in the present invention.
[0020] Embodiment 1, as Figure 1 As shown, the embodiment of the present invention provides an RFID file processing system for a smart archive warehouse, including a state generation module 11, a semantic clustering module 12, a personnel tag clustering module 13, a frequency statistics module 14 and an archiving configuration module 15. Among them, the state generation module 11 is used to generate archive access status information in response to the RFID tag location record information, wherein the archive access status information includes the archive content identification and the archive personnel identification; the semantic clustering module 12 is used to semantically cluster the archive set according to the archive content identification to obtain the first-level archive clustering result; the personnel tag clustering module 13 is used to traverse the first-level archive clustering result to perform personnel tag clustering according to the archive personnel identification to obtain the second-level archive clustering result; the frequency statistics module 14 is used to traverse the second-level archive clustering result, and based on the archive access log, count the archive access frequency set; the archiving configuration module 15 is used to traverse the second-level archive clustering result according to the archive to be stored, perform classification configuration, and obtain the belonging cluster, wherein the belonging cluster has a category access frequency in the archive access frequency set, and if the category access frequency is greater than or equal to the access frequency threshold, it is archived in a pre-identified preset warehouse area.
[0021] Specifically, the status generation module 11 is used to respond to the RFID tag location record information and generate the archive access status information. Specifically, in the actual operation of the archive warehouse, each archive is equipped with an RFID tag. When the archive is accessed, the RFID reading device will record the change of the tag position in real time and obtain the RFID tag location record information. After receiving the RFID tag location record information, the status generation module 11 quickly processes it and generates the corresponding archive access status information. Among them, the archive access status information includes the archive content identification and the archive personnel identification; the archive content identification refers to the specific content feature information of the archive, such as the archive's subject classification, keywords, file number, etc.; the archive personnel identification refers to the creator, responsible person, reviewer and other personnel information related to the archive, which has been stored in the system when the archive is archived. In this way, the status generation module 11 can accurately record the archive access information and provide basic data support for subsequent archive clustering and intelligent archiving.
[0022] The semantic clustering module 12 is used to perform a semantic level cluster analysis on the archive collection based on the archive content identification to form a first-level archive clustering result. Specifically, in practical applications, the archive warehouse will store a large number of archive materials of different types and different themes. The semantic clustering module 12 analyzes the archive content identification information provided by the state generation module 11 to perform a semantic relevance analysis on the entire archive collection. For example, the semantic clustering module 12 can use technical means such as content feature extraction and similarity calculation to identify archives with similar content themes or high semantic relevance, and aggregate them into the same category. The semantic clustering module 12 extracts the content attribute set of each archive, including elements such as subject words, keywords, and content summaries, and determines the semantic relevance strength between archives by calculating the similarity of content attributes between different archives. When the semantic similarity of two or more archives exceeds the preset semantic similarity threshold, these archives are classified into the same cluster. Through semantic clustering processing, a first-level archive clustering result based on content features can be formed, so that content-related archives can be centrally managed and the efficiency of subsequent archive retrieval and utilization can be improved. Through semantic-based archive clustering, we can effectively overcome the shortcomings of traditional archive classification methods that rely too much on manual judgment and have rough classification, and achieve accurate classification of archive content.
[0023] The personnel label clustering module 13 is used to further perform secondary clustering processing according to the personnel identification of the archives on the basis of the primary archive clustering results formed by the semantic clustering module 12, so as to obtain a more refined secondary archive clustering result. Specifically, in actual archive management work, archives are not only related to their content themes, but also closely related to specific personnel groups. The personnel label clustering module 13 makes full use of the archive personnel identification information provided by the state generation module 11, performs traversal processing for each category of the primary archive clustering, and identifies the archive subset associated with a similar personnel group under the same semantic category. For example, the personnel label clustering module 13 extracts characteristic information such as the personnel position type set of the archive, and determines the personnel association strength between the archives by calculating the similarity of the relevant personnel attributes between different archives. When the position type similarity of the relevant personnel (such as the creator, the responsible person, the reviewer, etc.) of two or more archives exceeds the preset personnel similarity threshold, these archives are classified into the same secondary cluster. Through the secondary clustering processing based on the personnel identification, it is possible to realize the personnel association grouping based on the content similarity, forming a two-dimensional clustering structure of content and personnel. This clustering structure not only reflects the content relevance of archives, but also reflects the characteristics of the archive user groups, providing a more accurate classification basis for subsequent archive configuration and access frequency analysis. Through secondary clustering, the limitation of traditional archive management that only focuses on content characteristics and ignores user characteristics is effectively solved, making archive management more humane and accurate.
[0024] The frequency statistics module 14 is used to traverse and analyze the secondary archive clustering results generated by the personnel label clustering module 13, and combine the archive access log data to statistically generate an archive access frequency set. In the actual archive management process, there are significant differences in the use frequency of different archives. This use frequency feature is an important manifestation of the value and use rules of archives. The frequency statistics module 14 traverses each cluster group in the secondary archive clustering results, extracts the historical access records of the relevant archives, and performs time series analysis to calculate the archive access frequency index of each cluster group. For example, the frequency statistics module 14 first accesses the stored archive access log database, which records the detailed information such as the access time, access personnel, and return status of the archive. Then, the frequency statistics module 14 counts the number of archive accesses within a specific time period (such as day, week, month or quarter) for each secondary cluster group, and calculates the average access frequency, peak access frequency and other indicators to form a complete archive access frequency set. Through frequency statistical analysis, the actual use and usage rules of various types of archives can be objectively reflected, providing data support for subsequent archiving configuration. This kind of archive management based on actual frequency of use breaks through the limitations of traditional archive management that only manages based on archive content or creation time, enabling the archive storage layout to better adapt to actual usage needs and improve archive retrieval and usage efficiency.
[0025] The archiving configuration module 15 is used to intelligently archive the archives to be stored, ensuring that the archive storage location matches its usage characteristics, thereby optimizing the overall archive management efficiency. In the actual operation process, when there are archives to be stored that need to be stored in the warehouse, the archiving configuration module 15 will first analyze the characteristic information of the archive, and then traverse the existing secondary archive clustering results to find the cluster group that is most similar to the archive to be stored in terms of content and personnel labels, and determine it as the belonging cluster of the archive. The archiving configuration module 15 obtains the category access frequency value of the belonging cluster through the archive access frequency set provided by the frequency statistics module 14. At the same time, the system presets a frequency access threshold as an important basis for judging the priority of the archive storage location. When the category access frequency of the belonging cluster is greater than or equal to this threshold, the archive to be stored is archived in a pre-identified preset warehouse area. Among them, the preset warehouse area is a high-quality storage space planned according to the principle of convenient access, such as an area close to the warehouse entrance and exit, with convenient transportation routes or complete access equipment. Through the frequency-oriented storage strategy, frequently used files can be arranged in a more convenient location, while infrequently used files can be stored in the depths of the warehouse or in a remote area. This intelligent archiving configuration optimizes the physical storage layout of the files, making the storage location of the files highly matched with their actual usage needs, effectively reducing the time cost and manpower consumption in the process of accessing the files, and greatly improving the work efficiency of file management.
[0026] Through the status generation module 11, the semantic clustering module 12, the personnel tag clustering module 13, the frequency statistics module 14 and the archiving configuration module 15, multi-level intelligent analysis and archiving of archives are realized, which solves the problems of inaccurate archive classification, unreasonable storage location and low retrieval efficiency existing in traditional archive management, realizes the intelligent and efficient archive management, and improves the utilization efficiency and management level of archive resources.
[0027] Furthermore, the state generation module 11 includes a tag acquisition unit and an archive data acquisition unit. The tag acquisition unit is used to obtain the RFID tag according to the RFID tag location record information; the archive data acquisition unit is used to obtain the archive data uniquely associated with the RFID tag based on the archive filing record, wherein the archive data stores the archive content identifier and the archive personnel identifier when being archived.
[0028] In an optional implementation, the status generation module 11 includes two functional units, namely a label acquisition unit and an archive data acquisition unit. The complete generation process of archive access status information is realized through the coordinated work of these two units.
[0029] The tag acquisition unit is a front-end processing unit of the state generation module 11, which is used to respond to and process the RFID tag location record information. In actual applications, when the file is taken or moved, the position of the RFID tag attached to the file will change, and these changes will be captured by the RFID reading device and form the RFID tag location record information. After receiving the RFID tag location record information, the tag acquisition unit accurately identifies and obtains the unique identification code of the RFID tag with the position change through operations such as signal processing and data analysis. The archive data acquisition unit, as a back-end processing unit of the state generation module 11, is responsible for associating the tag information with the actual archive data. The unit queries the archive data uniquely corresponding to the acquired RFID tag by accessing the archive record database of the system. Among them, each archive has stored a complete archive content identification and archive personnel identification for it when it is initially archived. The archive content identification contains the content features of the archive, such as the subject, category, and keywords, while the archive personnel identification records the personnel information related to the archive, such as the creator, responsible person, and auditor.
[0030] Through the tag acquisition unit and the archive data acquisition unit, the status generation module 11 can realize the complete conversion process from the RFID tag position change to the archive retrieval status information, ensuring that the system can accurately track the real-time status and usage of each archive, and lay a data foundation for subsequent clustering analysis and archiving configuration.
[0031] Furthermore, the semantic clustering module 12 includes a first content identification unit, a second content identification unit, a semantic similarity calculation unit and a semantic clustering execution unit. The first content identification unit is used to obtain a first archive content identification of the archive content identification, wherein the first archive content identification includes a first content attribute set; the second content identification unit is used to obtain a second archive content identification of the archive content identification, wherein the second archive content identification includes a second content attribute set; the semantic similarity calculation unit is used to calculate the number ratio of the intersection and the union of the first content attribute set and the second content attribute set, set it as the first semantic similarity, and add it to the semantic similarity set; the semantic clustering execution unit is used to perform semantic clustering on the archive set based on the semantic similarity threshold and in combination with the semantic similarity set to obtain a first-level archive clustering result.
[0032] In a preferred embodiment, Figure 2 As shown, the semantic clustering module 12 includes a first content identification unit, a second content identification unit, a semantic similarity calculation unit and a semantic clustering execution unit, which work together to complete the process from archive content feature extraction to cluster formation to achieve effective semantic clustering of archive collections.
[0033] Among them, the first content identification unit is responsible for extracting the first archive content identification from the archive content identification. In the specific implementation, when it is necessary to compare the semantic similarity of two archives, the first content identification unit first obtains the content identification information of the first archive, and extracts the various characteristic elements that constitute the first content attribute set, such as subject terms, keywords, content summaries, document types and other attributes. These attributes together constitute the first content attribute set that describes the characteristics of the archive content. The second content identification unit has similar functions to the first content identification unit, and is mainly responsible for obtaining the content identification information of the second archive, and extracting and forming the second content attribute set. Through the parallel work of these two units, the system can simultaneously obtain the content feature sets of the two archives to be compared, laying the foundation for the subsequent similarity calculation.
[0034] The semantic similarity calculation unit is used to calculate the semantic similarity between the two acquired content attribute sets. The unit adopts the intersection-and-union ratio method, that is, it calculates the ratio of the number of intersection elements of the first content attribute set and the second content attribute set to the number of union elements, and defines the ratio as the first semantic similarity, thereby reflecting the similarity of the two archives in terms of content. The calculated first semantic similarity value is then added to the semantic similarity set maintained by the system, which records the semantic similarity relationship between all archive pairs. The semantic clustering execution unit is responsible for clustering the entire archive set based on the preset semantic similarity threshold and the generated semantic similarity set. When the semantic similarity of two archives is greater than or equal to the preset threshold, the system determines that the two archives belong to the same semantic category; otherwise, they are determined to be different categories. Through this threshold-based judgment mechanism, the system can automatically aggregate archives with similar semantics to form a first-level archive clustering result.
[0035] Through the above four functional units, the semantic clustering module 12 can efficiently complete the semantic analysis and clustering processing of archives, provide a basic classification structure for subsequent personnel label clustering, and promote the intelligent operation of the entire smart archive warehouse system.
[0036] Furthermore, the personnel label clustering module 13 includes a first personnel identification unit, a second personnel identification unit, a personnel similarity calculation unit and a personnel label clustering execution unit. The first personnel identification unit is used to obtain a first file personnel identification of the file personnel identification, wherein the first file personnel identification includes a first personnel position type set; the second personnel identification unit is used to obtain a second file personnel identification of the file personnel identification, wherein the second file personnel identification includes a second personnel position type set; the personnel similarity calculation unit is used to calculate the number ratio of the intersection and the union of the first personnel position type set and the second personnel position type set, set it as the first personnel similarity, and add it to the personnel similarity set; the personnel label clustering execution unit is used to perform personnel label clustering on the file set based on the personnel similarity threshold and in combination with the personnel similarity set to obtain a secondary file clustering result.
[0037] In a preferred embodiment, Figure 3 As shown, the personnel label clustering module 13 includes a first personnel identification unit, a second personnel identification unit, a personnel similarity calculation unit and a personnel label clustering execution unit to achieve fine clustering of archives based on personnel characteristics.
[0038] The first personnel identification unit is responsible for extracting the first file personnel identification information from the file personnel identification. In the actual application process, when the system needs to compare the personnel relevance of two files, the first personnel identification unit first obtains the personnel identification information of the first file and extracts the first personnel position type set from it. This set contains the position type information of various types of personnel related to the file (such as creators, responsible persons, reviewers, etc.), such as department managers, engineers, financial directors, etc. These position type information constitute the first personnel position type set that describes the characteristics of the file personnel. The second personnel identification unit has similar functions to the first personnel identification unit and is responsible for obtaining the personnel identification information of the second file and extracting it to form the second personnel position type set. Through the parallel operation of these two units, the system can simultaneously obtain the personnel feature sets of the two files to be compared, providing necessary data for subsequent similarity calculations.
[0039] The personnel similarity calculation unit is used to calculate the similarity between the two obtained personnel position type sets. This unit also adopts the intersection-and-union ratio calculation method, that is, it calculates the ratio of the number of intersection elements of the first personnel position type set and the second personnel position type set to the number of union elements, and defines this ratio as the first personnel similarity. The first personnel similarity reflects the matching degree of the two files in terms of personnel position type. The calculated first personnel similarity is then added to the personnel similarity set maintained by the system, which records the personnel association similarity relationship between all file pairs. The personnel label clustering execution unit is responsible for further subdividing and clustering the file set after semantic clustering based on the preset personnel similarity threshold and the generated personnel similarity set. When the personnel similarity of the two files is greater than or equal to the preset threshold, the system determines that the two files belong to the same category in terms of personnel characteristics; otherwise, they are determined to be different categories. Through the judgment mechanism based on the personnel similarity threshold, the system can further aggregate files with similar personnel characteristics into more refined subcategories on the basis of semantic clustering to form a secondary file clustering result.
[0040] Through the personnel label clustering module 13, the personnel feature analysis and clustering processing of the archives can be effectively realized, so that the archive management not only focuses on the content dimension, but also can carry out refined management from the personnel association dimension, thereby improving the intelligence level and utilization efficiency of the archive management.
[0041] Furthermore, the archive configuration module 15 includes a scoring function configuration unit, a priority scoring unit, a warehouse shelf clustering unit, a clustering result sorting unit and a pre-marking unit. Among them, the scoring function configuration unit is used to configure the warehouse shelf priority scoring function: ,in, Represents the warehouse shelf priority, Represents the path distance from the shelf to the shelf's associated access position, Characterizes the path distance between the shelf-related access location and the door, and The first weight and the second weight are represented; the priority scoring unit is used to traverse the warehouse shelves for priority scoring according to the warehouse shelf priority scoring function to obtain a priority scoring set; the warehouse shelf clustering unit is used to cluster the warehouse shelves according to the priority scoring threshold to obtain the warehouse shelf clustering result, wherein the warehouse shelf clustering result has a centroid priority scoring identifier; the clustering result sorting unit is used to sort the warehouse shelf clustering results from small to large according to the centroid priority scoring identifier to obtain the warehouse shelf clustering sorting result; the pre-identification unit is used to select vacant warehouse areas for pre-identification according to the warehouse shelf clustering sorting result.
[0042] In a preferred embodiment, the archive configuration module 15 includes a scoring function configuration unit, a priority scoring unit, a warehouse shelf clustering unit, a clustering result sorting unit and a pre-identification unit, which completes the process from shelf scoring to space pre-allocation, thereby realizing intelligent archive configuration of archives.
[0043] The scoring function configuration unit is used to set a mathematical model for evaluating the warehouse shelf priority. The warehouse shelf priority scoring function configured by this unit is: In this function, Characterizes the priority of warehouse shelves and is a comprehensive indicator to measure the storage suitability of shelves; The path distance that characterizes the movement of the shelf to the shelf-associated access location reflects the convenience of moving the archive from the storage location to the regular use location; It represents the path distance between the shelf-related access location and the door, reflecting the convenience of the archive access location to the warehouse entrance and exit; and The first weight and the second weight are respectively represented to adjust the importance ratio of the two distance factors in the priority scoring. Through this exponential weighting function, the system can comprehensively consider the multi-dimensional factors of the shelf location and realize scientific and reasonable priority evaluation. The priority scoring unit traverses and scores all the shelves in the warehouse according to the above-mentioned warehouse shelf priority scoring function. The priority scoring unit collects the location information and distance data of each shelf, substitutes them into the warehouse shelf priority scoring function for calculation, and finally generates a priority score value for each shelf, and summarizes these score values to form a priority score set, providing a data basis for subsequent shelf clustering.
[0044] The warehouse shelf clustering unit is responsible for classifying and aggregating the scored warehouse shelves according to the priority score threshold. This unit adopts a threshold-based clustering algorithm to classify shelves with similar priority scores into the same category to form a warehouse shelf clustering result. Among them, each cluster group has a centroid priority score identifier, which represents the average priority level of this type of shelf and is an important indicator to measure the overall priority of this type of shelf. The clustering result sorting unit is responsible for systematically sorting the clustered shelf groups. This unit sorts the warehouse shelf clustering results in the order of the centroid priority score identifier from small to large to generate the warehouse shelf clustering sorting results. This sorting method ensures that shelf categories with higher priority (smaller score values) can be given priority, providing a more convenient storage location for frequently used archives. The pre-marking unit is responsible for selecting suitable vacant warehouse areas for pre-marking according to the sorted shelf clustering results. This unit will give priority to the vacant locations in the shelf categories with higher rankings (higher priority), pre-mark these locations as potential storage areas for frequently used archives, and prepare space for the subsequent actual archiving of archives.
[0045] Through the above five functional units, the archive configuration module 15 can realize intelligent space allocation based on multi-factor evaluation, so that the physical storage layout of the archives is highly matched with its usage frequency and convenience requirements, thereby improving the efficiency of archive management and user experience.
[0046] Furthermore, the pre-identification unit includes an archive sorting subunit, a shelf sorting subunit, a fitness function subunit, a storage scheme acquisition subunit, and a minimum value sorting subunit. Among them, the archive sorting subunit is used to sort the archives to be stored from large to small according to the category access frequency when there are multiple archives to be stored whose category access frequency is greater than or equal to the access frequency threshold, and obtain the sorting result of the archives to be stored; the shelf sorting subunit is used to extract the category sorting result of the vacant shelves according to the warehouse shelf cluster sorting result; the fitness function subunit is used to construct the storage fitness function: ,in, Characterize the fitness of the storage solution, Indicates the category number of the j-th pre-occupied shelf, represents the category of the kth empty shelf, H represents the total number of pre-occupied shelves, M represents the total number of empty shelves, The sequence number representing the i-th file to be stored in the sorting result of the files to be stored. The serial number of the pre-occupied shelf for the i-th file to be stored in the empty shelf category sorting result is represented, and N represents the number of files to be stored; the storage plan acquisition subunit is used to randomly place the sorting results of the files to be stored into the empty shelves to obtain several storage plans; the minimum value sorting subunit is used to perform minimum value sorting on several storage plans according to the storage fitness function, and obtain the empty warehouse area for pre-marking.
[0047] In a preferred embodiment, the pre-identification unit includes an archive sorting subunit, a shelf sorting subunit, a fitness function subunit, a storage scheme acquisition subunit and a minimum value sorting subunit to achieve efficient and accurate archive pre-identification and archiving.
[0048] The archive sorting subunit prioritizes the archives to be stored that are used frequently. When the system detects that the category access frequencies of multiple archives to be stored are greater than or equal to the preset access frequency threshold, the subunit will sort these archives in order from large to small according to the size of the category access frequency, and generate the sorting results of the archives to be stored. This sorting mechanism ensures that archives with higher frequency of use can obtain higher storage priority, laying the foundation for the subsequent optimal storage location allocation. Then, the shelf sorting subunit processes the warehouse space resource information, extracts the currently available empty shelf information from the aforementioned warehouse shelf cluster sorting results, and forms the empty shelf category sorting results. This result retains the sorting characteristics of the shelf priority, providing a space resource basis for subsequent archive-shelf matching.
[0049] The fitness function subunit is responsible for building a mathematical model for evaluating the quality of storage solutions. The storage fitness function defined by this subunit is: In this function, Characterizes the adaptability of the storage solution, that is, the quality of the solution; The serial number representing the category of the j-th pre-occupied shelf reflects the attributes of the allocated shelf; represents the category of the kth empty shelf, representing the available storage resources; H represents the total number of pre-occupied shelves; M represents the total number of empty shelves; The serial number representing the i-th file to be stored in the sorting result of the files to be stored; represents the number of the pre-occupied shelf of the i-th file to be stored in the result of the classification of the vacant shelves; N represents the number of files to be stored. This function comprehensively considers the matching degree between the file priority and the shelf priority, as well as factors such as resource utilization, and can comprehensively evaluate the rationality of the storage plan.
[0050] The storage scheme acquisition subunit is responsible for generating a variety of possible archive storage schemes. This subunit generates multiple different storage combination schemes by randomly assigning the archives in the sorting results of the archives to be stored to the empty shelves. This random allocation strategy can generate a variety of candidate schemes, providing a rich selection space for the subsequent optimal scheme screening. The minimum value sorting subunit is the decision-making execution part of the pre-identification unit, responsible for screening out the optimal solution from multiple candidate storage schemes. Based on the aforementioned storage fitness function, this subunit calculates the FIT value of each storage scheme and selects the scheme with the smallest FIT value as the final decision. Since the smaller the FIT value, the higher the scheme fitness, this minimum value sorting strategy can effectively identify the storage scheme that best meets the system optimization goal, and determine the vacant warehouse area to be identified accordingly.
[0051] Through the collaborative work of the above five sub-units, the pre-identification unit can realize the intelligent pre-allocation of storage locations for frequently used archives, so that the physical storage location of the archives is highly matched with their usage characteristics, thereby improving the efficiency of archive management and user experience.
[0052] Furthermore, the archive configuration module 15 also includes a similarity evaluation unit and an attribution cluster acquisition unit. The similarity evaluation unit is used to perform semantic and personnel label similarity evaluations based on the archives to be stored and the centroid archives of each category of the secondary archive clustering results, and perform mean calculations to obtain a similarity set; the attribution cluster acquisition unit is used to extract the class corresponding to the maximum similarity based on the similarity set and set it as the attribution cluster.
[0053] In a preferred embodiment, the archive configuration module 15 further includes a similarity evaluation unit and a belonging cluster acquisition unit to achieve intelligent matching of the archives to be stored with the existing archive clusters, thereby improving the accuracy of the archive configuration.
[0054] In a preferred embodiment, the archive configuration module 15 executes the steps further including, based on the archive to be stored, respectively evaluating the similarity of semantics and personnel labels with the centroid archive of each category of the secondary archive clustering result, and performing mean calculation to obtain a similarity set, which includes: Through the warehouse management terminal, configure the file content identification and the file personnel identification of the files to be stored.
[0055] The similarity evaluation unit is an analysis and evaluation component of the archive configuration module 15, which is used to perform a comprehensive similarity analysis on the archives to be stored and the existing clusters. Specifically, when the archives to be stored need to be stored, the unit compares the archives to be stored with the centroid archives of each category in the secondary archive clustering results. This comparison includes not only similarity evaluation at the content semantic level, but also similarity evaluation at the personnel label level, thereby realizing multi-dimensional matching analysis. For each cluster, the similarity evaluation unit calculates the mean of semantic similarity and personnel label similarity to form a comprehensive similarity index, and summarizes these indexes into a similarity set. This mean calculation method balances the importance of content features and personnel features, making archive classification more comprehensive and reasonable.
[0056] The attribution cluster acquisition unit, as a decision-making execution component of the archive configuration module 15, is responsible for determining the best attribution of the archive to be stored based on the similarity evaluation results. The unit identifies the cluster with the highest similarity to the archive to be stored by analyzing the similarity set. Specifically, the attribution cluster acquisition unit extracts the cluster corresponding to the maximum similarity value from the similarity set, and sets the cluster as the attribution cluster of the archive to be stored. This attribution determination method based on maximum similarity ensures that the archive to be stored can be classified into the archive category that best matches its characteristics.
[0057] Through the similarity evaluation unit and the attribution cluster acquisition unit, the archiving configuration module 15 can achieve accurate classification of the archives to be stored, providing a reliable classification basis for subsequent storage location optimization. This archive attribution determination method based on multi-dimensional similarity improves the accuracy and rationality of archive classification and enhances the intelligence level of the entire smart archive warehouse system.
[0058] The embodiment of the present invention provides a smart archive warehouse RFID file processing system, which has at least the following technical effects: Through the status generation module, in response to the RFID tag location record information, the archive access status information is generated, wherein the archive access status information includes the archive content identification and the archive personnel identification, providing raw data support for subsequent analysis. Through the semantic clustering module, according to the archive content identification, the archive collection is semantically clustered to obtain the first-level archive clustering result, and the content characteristics of the archive are analyzed, and archives with similar themes or related contents are classified into the same category to form a preliminary archive classification structure. Through the personnel label clustering module, according to the archive personnel identification, the first-level archive clustering result is traversed to perform personnel label clustering, and the second-level archive clustering result is obtained, and further subdivision is achieved based on the semantic clustering according to the archive personnel identification, and archives with similar personnel characteristics are classified into the same subclass to achieve a two-dimensional classification structure of content and personnel. Through the frequency statistics module, the second-level archive clustering result is traversed, and based on the archive access log, the archive access frequency set is statistically analyzed to provide data basis for the optimization of archive storage location. Through the archiving configuration module, according to the archives to be stored, the secondary archive clustering results are traversed for classification configuration to obtain the belonging cluster, wherein the belonging cluster has a category access frequency in the archive access frequency set. If the category access frequency is greater than or equal to the access frequency threshold, it is archived in a pre-identified preset warehouse area, thereby arranging the archives in a pre-identified warehouse area that is easy to access, realizing the priority arrangement of high-frequency used archives, thereby improving the efficiency of archive retrieval and management.
[0059] It should be noted that in the above embodiments, the description of each embodiment has its own emphasis, and for parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0060] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0061] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0062] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0063] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0064] Although preferred embodiments of the present invention have been described, additional changes and modifications may occur to these embodiments once those skilled in the art understand the basic inventive concepts.
[0065] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention belong to the scope of the present invention and its equivalent technologies, the present invention is also intended to include these changes and variations.
Claims
1. A smart archive warehouse RFID file processing system, characterized in that: include: A status generation module, used to generate archive access status information in response to the RFID tag location record information, wherein the archive access status information includes an archive content identifier and an archive personnel identifier; A semantic clustering module, used to perform semantic clustering on the archive collection according to the archive content identifier to obtain a first-level archive clustering result; A personnel label clustering module, used to perform personnel label clustering by traversing the first-level file clustering results according to the file personnel identification, and obtain a second-level file clustering result; A frequency statistics module, used to traverse the secondary archive clustering results and count the archive access frequency set based on the archive access log; The archiving configuration module is used to traverse the secondary archive clustering results according to the archives to be stored, perform classification configuration, and obtain the belonging cluster, wherein the belonging cluster has a category access frequency in the archive access frequency set. If the category access frequency is greater than or equal to the access frequency threshold, it is archived in a pre-identified preset warehouse area.
2. The system according to claim 1, characterized in that The state generation module comprises: A tag acquisition unit, configured to obtain an RFID tag according to the RFID tag location record information; The archive data acquisition unit is used to obtain the archive data uniquely associated with the RFID tag based on the archive filing record, wherein the archive data stores the archive content identifier and the archive personnel identifier when being filed.
3. The system according to claim 1, characterized in that The semantic clustering module includes: A first content identification unit, used to obtain a first archive content identification of the archive content identification, wherein the first archive content identification includes a first content attribute set; A second content identification unit, used to obtain a second archive content identification of the archive content identification, wherein the second archive content identification includes a second content attribute set; A semantic similarity calculation unit, used to calculate the ratio of the number of intersections and unions of the first content attribute set and the second content attribute set, set it as a first semantic similarity, and add it into a semantic similarity set; The semantic clustering execution unit is used to perform semantic clustering on the archive set based on a semantic similarity threshold and in combination with the semantic similarity set to obtain the first-level archive clustering result.
4. The system according to claim 1, characterized in that The personnel label clustering module includes: A first personnel identification unit, configured to obtain a first archive personnel identification of the archive personnel identification, wherein the first archive personnel identification includes a first personnel position type set; A second personnel identification unit, used to obtain a second archive personnel identification of the archive personnel identification, wherein the second archive personnel identification includes a second personnel position type set; A personnel similarity calculation unit, used for calculating the ratio of the number of intersections and unions of the first personnel position type set and the second personnel position type set, setting it as a first personnel similarity, and adding it into a personnel similarity set; The personnel label clustering execution unit is used to perform personnel label clustering on the archive set based on the personnel similarity threshold and in combination with the personnel similarity set to obtain the secondary archive clustering result.
5. The system according to claim 1, wherein: The archiving configuration module includes: Scoring function configuration unit, used to configure warehouse shelf priority scoring function: , in, Represents the warehouse shelf priority, Represents the path distance from the shelf to the shelf's associated access position, Characterizes the path distance between the shelf-related access location and the door, and characterizing a first weight and a second weight; A priority scoring unit, used to traverse the warehouse shelves to perform priority scoring according to the warehouse shelf priority scoring function, and obtain a priority scoring set; A warehouse shelf clustering unit, used for clustering the warehouse shelves according to the priority score threshold value to obtain a warehouse shelf clustering result, wherein the warehouse shelf clustering result has a centroid priority score identifier; A clustering result sorting unit, used to sort the warehouse shelf clustering results from small to large according to the centroid priority score identifier to obtain a warehouse shelf clustering sorting result; The pre-marking unit is used to select vacant warehouse areas for pre-marking according to the warehouse shelf clustering and sorting results.
6. The system according to claim 5, characterized in that The pre-marking unit comprises: The file sorting subunit is used to sort the files to be stored from large to small according to the category access frequency when there are multiple files to be stored whose category access frequency is greater than or equal to the access frequency threshold, and obtain the sorting result of the files to be stored; The shelf sorting subunit is used to extract the category sorting results of the vacant shelves according to the cluster sorting results of the warehouse shelves; The fitness function subunit is used to construct and store the fitness function: , in, Characterize the fitness of the storage solution, Indicates the category number of the j-th pre-occupied shelf, represents the category of the kth empty shelf, H represents the total number of pre-occupied shelves, M represents the total number of empty shelves, The sequence number representing the i-th file to be stored in the sorting result of the files to be stored. The number of the pre-occupied shelf representing the i-th file to be stored in the sorting result of the vacant shelves, and N represents the number of files to be stored; A storage plan acquisition subunit is used to randomly place the sorting results of the archives to be stored into empty shelves to obtain several storage plans; The minimum value sorting subunit is used to perform minimum value sorting on the plurality of storage schemes according to the storage fitness function, and obtain the vacant warehouse area for pre-marking.
7. The system according to claim 1, characterized in that The archiving configuration module also includes: A similarity evaluation unit is used to evaluate the similarity of semantics and personnel labels with the centroid file of each category of the secondary file clustering result based on the file to be stored, and perform mean calculation to obtain a similarity set; The belonging cluster acquisition unit is used to extract the class corresponding to the maximum similarity according to the similarity set, and set it as the belonging cluster.
8. The system according to claim 7, characterized in that Based on the archive to be stored, the similarity of semantics and personnel labels is evaluated with the centroid archive of each category of the secondary archive clustering result, and the mean is calculated to obtain a similarity set, which includes: Through the warehouse management terminal, configure the file content identification and the file personnel identification of the files to be stored.
Citation Information
Cited By
Intelligent archival repository management method and system based on multiple modules
CN121032397A
Intelligent archive library archive query method based on RFID technology
CN121119319A
File box optimization design printing method based on electronic file system
CN121168043A