File management method and system

By analyzing user access logs and preloading common archive data to the cache area, the problem of long wait time during archive call is solved, and the efficiency of file search and user experience is improved.

CN120067065APending Publication Date: 2025-05-30陕西省水工环地质调查中心
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510147785.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

During the process of file call, users need to spend more time waiting for the file to be loaded or saved, resulting in reduced work efficiency and productivity and poor user experience.

Method used

By analyzing the user's access log, identifying the feature set of commonly used files, and preloading the file data corresponding to these features into the cache area. When the user initiates an archive retrieval request, it first querys from the cache area. If the cache hits, the data will be returned directly, otherwise querying from the storage area.

Benefits of technology

It improves the search efficiency of commonly used files, reduces the waiting time for users during file call, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067065A_ABST
    Figure CN120067065A_ABST
Patent Text Reader

Abstract

The invention discloses an archive management method and system, belongs to the technical field of archive management, and can improve the search efficiency of common archives and improve the user experience. Comprising the following steps: in response to a received request access of a first user, obtaining a user access log of the first user in a predetermined historical time period; analyzing the user access log to obtain a common file feature set of the first user; preloading archive data corresponding to each archive feature in the common archive feature set to a cache region; when a user initiates an archive retrieval request, firstly querying from the cache region, and if the cache is hit, reading corresponding archive data and returning the corresponding archive data to the first user; and if the cache is not hit, querying and reading the corresponding archive data from the storage area and returning the archive data to the first user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of file management, and specifically to a file management method and system. Background Art

[0002] File management is a systematic process of organizing and managing documents and records, which covers multiple aspects such as document creation, classification, storage, retrieval, protection, retention, and disposal. The importance of file management is self-evident. It not only ensures the security, reliability, and availability of information, but also helps maintain the effective operation of personnel relations.

[0003] With the development of artificial intelligence technology, the application of artificial intelligence in file management is becoming increasingly widespread. Through technical means such as natural language processing, machine learning, and data analysis, artificial intelligence can achieve automated classification and archiving of files, quickly retrieve and extract information, formulate intelligent archiving strategies, and identify potential risks in management. This intelligent file management not only improves management efficiency and information availability, but also helps reduce human errors and management costs.

[0004] However, during the process of file retrieval, users need to spend more time waiting for the file to be loaded or saved, which also reduces work efficiency and productivity. For users who need to access files frequently, the cumbersome file retrieval process will also lead to a poor user experience and affect user satisfaction. Therefore, "how to cache frequently used files" is the technical problem to be solved by the present invention.

[0005] The disclosure of the above background art content is only for assisting in understanding the concept and technical solution of the present invention, and it does not necessarily belong to the prior art of this patent application. Without clear evidence indicating that the above content was publicly available on the filing date of this patent application, the above background art should not be used to evaluate the novelty and inventiveness of this application. Summary of the Invention

[0006] This application provides a file management method and system, which can improve the search efficiency of frequently used files and enhance the user experience.

[0007] To achieve the above object, the embodiments of this application disclose the following technical solutions:

[0008] In the first aspect, the embodiments of this application provide a file management method, including the following steps:

[0009] In response to receiving a request for access from a first user, obtain the user access log of the first user within a predetermined historical time period;

[0010] Analyze the user access log to obtain the set of frequently used file features of the first user;

[0011] Pre-load the file data corresponding to each file feature in the set of common file features into the cache area;

[0012] When a user initiates a file retrieval request, first query from the cache area. If the cache hits, read the corresponding file data and return it to the first user; if the cache misses, query and read the corresponding file data from the storage area and return it to the first user.

[0013] In the embodiment of the present application, when a user requests to access a file, the access logs of the user within a predetermined historical time period will be obtained. By analyzing these logs, the set of file features commonly used by the user can be identified, so as to understand the user's access habits and identify which files are frequently accessed.

[0014] Next, pre-load the data corresponding to the features of these common files into the cache area. When a user initiates a file retrieval request, first query from the cache area. If the cache hits, read the corresponding file data and return it to the first user. Therefore, when the user requests these common files again, the data can be quickly returned directly from the cache. In this way, the search efficiency of common files can be improved, the waiting time of the user during the file call process can be effectively reduced, and the user experience can be enhanced.

[0015] In some possible implementation manners of the first aspect, obtain the user access logs of the first user in the past two days, and the user access logs are independently recorded daily. In this way, the time range is more reasonable, and the recent user behavior can be analyzed more quickly and accurately.

[0016] In some possible implementation manners of the first aspect, the steps of analyzing the user access logs to obtain the set of common file features of the first user include:

[0017] Read the access records in the user access logs one by one;

[0018] Parse each access record to identify the field containing the file identifier and extract the file identifier;

[0019] Based on the extracted multiple file identifiers, determine the file features of each access request;

[0020] Perform access count statistics, merging and sorting on the multiple file features to obtain the access count statistics result;

[0021] According to the access count statistics result, select the file features with the top N access counts and store them in a preset set to obtain the set of common file features.

[0022] In this way, the set of file features in the cache can be updated in real time and dynamically adjusted according to the latest access data, ensuring that the content in the cache always conforms to the actual needs of the user. Additionally, by setting a threshold for the number of accesses (the top N file features), the amount of features in the cache can be effectively limited, preventing too many infrequently used features from being loaded into the cache.

[0023] In some possible implementation manners of the first aspect, the steps of parsing each access record to identify the field containing the file identifier and extracting the file identifier include:

[0024] Read each access record in the user access log line by line to obtain the original text data of the record;

[0025] Apply a preset deep learning model to perform word segmentation on the original text data to identify multiple fields;

[0026] Determine whether each field contains a specific identifier keyword, where the keyword is used to indicate the field where the file identifier is located;

[0027] After identifying the field containing the file identifier, extract the specific content within each field to obtain multiple file identifiers.

[0028] In some possible implementation manners of the first aspect, the preset deep learning model is a Transformer model.

[0029] In some possible implementation manners of the first aspect, the file management method further includes the following steps: Dynamically adjust the file distribution in the storage area according to the frequently used file features. In this way, it is beneficial to improve the access efficiency of the storage system.

[0030] In a second aspect, an embodiment of the present application provides a file management system, including:

[0031] A first acquisition module, configured to obtain the user access log of the first user within a predetermined historical time period in response to receiving a request for access from the first user;

[0032] A first analysis module, configured to analyze the user access log to obtain a set of frequently used file features of the first user;

[0033] A first loading module, configured to preload the file data corresponding to each file feature in the set of frequently used file features into the cache area;

[0034] A first query module, configured to first query from the cache area when the user initiates a file retrieval request. If the cache is hit, read the corresponding file data and return it to the first user; if the cache is not hit, query and read the corresponding file data from the storage area and return it to the first user.

[0035] In some possible embodiments of the second aspect, the first acquisition module is specifically configured to acquire the user access logs of the first user in the past two days, and the user access logs are independently recorded daily.

[0036] In some possible embodiments of the second aspect, the first analysis module is specifically configured to: read the access records in the user access logs one by one; parse each access record to identify the fields containing the file identifiers and extract the file identifiers; based on the extracted multiple file identifiers, determine the file features of each access request; perform access count statistics, merging, and sorting on the multiple file features to obtain the access count statistics result; according to the access count statistics result, select the file features with the top N access counts and store them in a preset set to obtain the set of common file features.

[0037] In some possible embodiments of the second aspect, the first analysis module is specifically further configured to: read each access record in the user access logs line by line to obtain the original text data of the record; apply a preset deep learning model to perform word segmentation on the original text data to identify multiple fields; determine whether each field contains a specific identifier keyword, where the keyword is used to indicate the field where the file identifier is located; after identifying the field containing the file identifier, extract the specific content within each field to obtain multiple file identifiers.

[0038] In some possible embodiments of the second aspect, the preset deep learning model is a Transformer model.

[0039] In some possible embodiments of the second aspect, the file management system further includes a first adjustment module, configured to dynamically adjust the file distribution in the storage area according to the common file features.

[0040] In a third aspect, an embodiment of the present application provides an electronic device, including one or more processors; a storage device storing one or more programs thereon; when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any one of the technical solutions of the first aspect.

[0041] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, having a computer program stored thereon, and when the computer program is executed by a processor, it implements the method described in any one of the technical solutions of the first aspect.

[0042] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method described in any one of the technical solutions of the first aspect.

[0043] Among them, for the technical effects brought by any one of the design methods in the second to fifth aspects, reference can be made to the technical effects brought by different design methods in the first aspect, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary. For those of ordinary skill in the art, without creative efforts, other implementation drawings can also be obtained according to the provided drawings.

[0045] Figure 1 Schematic flowchart of a file management method provided by some embodiments of the present application;

[0046] Figure 2 Schematic structural diagram of a file management system provided by some embodiments of the present application;

[0047] Figure 3 Schematic structural diagram of an electronic device suitable for implementing some embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] Now the specific implementation schemes of the present invention will be described in detail. Although the present invention is described in conjunction with these specific implementation schemes, it should be understood that it is not intended to limit the present invention to these specific implementation schemes. On the contrary, these implementation schemes are intended to cover alternative, changed or equivalent implementation schemes that may be included within the spirit and scope of the invention defined by the claims. In the following description, a large number of specific details are set forth in order to provide a comprehensive understanding of the present invention. The present invention can be implemented without some or all of these specific details.

[0049] When used in conjunction with the terms "comprising", "the method comprises", or similar language in this specification and the appended claims, the singular forms "a", "an", "the" include plural references unless the context clearly dictates otherwise. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the art to which this invention belongs.

[0050] Application Overview: File management is a systematic process of organizing and managing documents and records. This process covers multiple aspects such as document creation, classification, storage, retrieval, protection, retention, and disposal. The importance of file management is self-evident. It not only ensures the security, reliability, and availability of information but also helps maintain the effective operation of personnel relations.

[0051] With the development of artificial intelligence technology, the application of artificial intelligence in archive management has become increasingly widespread. Through technical means such as natural language processing, machine learning, and data analysis, artificial intelligence can achieve automated classification and archiving of archives, quickly retrieve and extract information, formulate intelligent archiving strategies, and identify potential risks in management. This intelligent archive management not only improves management efficiency and information availability but also helps reduce human errors and management costs. However, during the process of accessing archives, users need to spend more time waiting for the archives to be loaded or saved, which also reduces work efficiency and productivity. For users who need to access archives frequently, the sluggish archive access process will also lead to a poor user experience and affect user satisfaction. Therefore, "how to cache frequently used archives" is the technical problem to be solved by the present invention.

[0052] In response to the above technical problems, the general idea of the technical solution provided in this application is as follows: Provide an archive management method, including the following steps: Upon receiving a request for access from a first user, obtain the user access log of the first user within a predetermined historical time period; analyze the user access log to obtain the set of frequently used archive features of the first user; preload the archive data corresponding to each archive feature in the set of frequently used archive features into the cache area; when the user initiates an archive retrieval request, first query from the cache area. If the cache hits, read the corresponding archive data and return it to the first user; if the cache misses, query and read the corresponding archive data from the storage area and return it to the first user.

[0053] In this way, when a user requests to access an archive, the access log of the user within a predetermined historical time period will be obtained. By analyzing these logs, the set of frequently used archive features of the user can be identified, so as to understand the user's access habits and identify which archives are frequently accessed.

[0054] Next, preload the data corresponding to the features of these frequently used archives into the cache area. When the user initiates an archive retrieval request, first query from the cache area. If the cache hits, read the corresponding archive data and return it to the first user. Therefore, when the user requests these frequently used archives again, the data can be quickly returned directly from the cache. In this way, the search efficiency of frequently used archives can be improved, the waiting time of users during the archive access process can be effectively reduced, and the user experience can be enhanced.

[0055] After introducing the basic principle of this application, the various non-limiting implementation manners of this application will be specifically introduced below with reference to the accompanying drawings of the specification. Please refer to Figure 1 , an embodiment of this application provides an archive management method, including the following steps:

[0056] S101: In response to receiving a request for access from a first user, obtaining a user access log of the first user within a predetermined historical time period; wherein the user's request for access refers to the user's attempt to log in to or access the archive management system itself.

[0057] Specifically, in some embodiments, user access logs are recorded independently every day. The predetermined historical time period is the user access logs within the past two days from the current moment. In this way, the time range is more reasonable, and recent user behavior can be analyzed more quickly and accurately. Of course, the present application is not limited to this. In other embodiments, the predetermined historical time period can also be the user access logs within the past three days, four days, or five days from the current moment.

[0058] S102: Analyze the user access log to obtain a common profile feature set of the first user;

[0059] Specifically, in some embodiments, the user access log may be analyzed through the following steps to obtain a common profile feature set of the first user:

[0060] The first step is to read the access records in the user access log one by one;

[0061] In the second step, each access record is parsed to identify the field containing the archive identifier and extract the archive identifier;

[0062] Specifically, in some embodiments, each access record may be parsed through the following steps to identify a field containing an archive identifier and extract the archive identifier:

[0063] In the first sub-step, each access record in the user access log is read line by line to obtain the original text data of the record;

[0064] In the second sub-step, a preset deep learning model is applied to perform word segmentation processing on the original text data to identify multiple fields; illustratively, the preset deep learning model can be but is not limited to a Transformer model.

[0065] The third sub-step is to determine whether each field contains a specific identifier keyword, where the keyword is used to indicate the field where the archive identifier is located;

[0066] The fourth sub-step is to extract the specific content in each field after identifying the fields containing the archive identifiers to obtain multiple archive identifiers.

[0067] The third step is to determine the archive characteristics of each access request based on the extracted multiple archive identifiers;

[0068] The fourth step is to perform access count statistics and merge and sort multiple archive features to obtain access count statistics results;

[0069] The fifth step is to select the archive features with the top N access times according to the access times statistics and store them in a preset set to obtain a commonly used archive feature set.

[0070] In this way, the archive feature set in the cache can be updated and dynamically adjusted in real time according to the latest access data, ensuring that the cache content always meets the actual needs of the user. In addition, by setting the threshold of the number of accesses (the top N archive features), the amount of features in the cache can be effectively limited to avoid loading too many, infrequently used features into the cache.

[0071] S103: preloading the archive data corresponding to each archive feature in the common archive feature set into the cache area;

[0072] S104: When a user initiates an archive retrieval request, the cache area is first queried. If the cache hits, the corresponding archive data is read and returned to the first user; if the cache does not hit, the corresponding archive data is queried and read from the storage area and returned to the first user.

[0073] Preferably, in some embodiments, the archive management method further comprises the following steps: dynamically adjusting the archive distribution of the storage area according to the common archive characteristics, which is conducive to improving the access efficiency of the storage system.

[0074] See also Figure 2 Based on the same inventive concept as the archive management method in the aforementioned embodiment, the embodiment of the present application provides an archive management system, including:

[0075] A first acquisition module 201 is used to acquire a user access log of the first user within a predetermined historical time period in response to receiving an access request from the first user;

[0076] A first analysis module 202, configured to analyze the user access log to obtain a common profile feature set of the first user;

[0077] The first loading module 203 is used to preload the archive data corresponding to each archive feature in the common archive feature set into the cache area;

[0078] The first query module 204 is used to query the cache area first when the user initiates a file retrieval request. If the cache hits, the corresponding file data is read and returned to the first user; if the cache does not hit, the corresponding file data is queried and read from the storage area and returned to the first user.

[0079] In some embodiments, the first acquisition module 201 is specifically configured to acquire the user access logs of the first user in the past two days, and the user access logs are independently recorded daily.

[0080] In some embodiments, the first analysis module 202 is specifically configured to: read the access records in the user access logs one by one; parse each access record to identify the fields containing the file identifiers and extract the file identifiers; based on the extracted multiple file identifiers, determine the file features of each access request; perform access count statistics, merging, and sorting on the multiple file features to obtain the access count statistics result; according to the access count statistics result, select the file features with the top N access counts and store them in a preset set to obtain the set of common file features.

[0081] In some embodiments, the first analysis module 202 is further specifically configured to: read each access record in the user access logs line by line to obtain the original text data of the record; apply a preset deep learning model to perform word segmentation on the original text data to identify multiple fields; determine whether each field contains a specific identifier keyword, where the keyword is used to indicate the field where the file identifier is located; after identifying the field containing the file identifier, extract the specific content within each field to obtain multiple file identifiers.

[0082] In some embodiments, the preset deep learning model is a Transformer model.

[0083] In some embodiments, the file management system further includes a first adjustment module 205, which is configured to dynamically adjust the file distribution in the storage area according to the common file features.

[0084] It can be understood that the various modules described in this file management system correspond to the respective steps in the file management method described with reference to Figure 1 Therefore, the operations, features, and beneficial effects described above for the method also apply to the file management system and the modules included therein, and will not be repeated here.

[0085] Please refer to Figure 3, based on the inventive concept of an archive management method in the foregoing embodiments, an embodiment of the present application provides an electronic device. The electronic device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), etc., and fixed terminals such as digital TVs, desktop computers, etc. The electronic device includes a processing device 301 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in the ROM 302 (Read Only Memory) or a program loaded from the storage device 308 into the RAM 303 (Random Access Memory). In the RAM 303, various programs and data required for the operation of the electronic device are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output interface (i.e., the I / O interface 305) is also connected to the bus 304.

[0086] Generally, the following devices can be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 can allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data.

[0087] Specifically, according to some embodiments of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the method shown in the flowchart. In such some embodiments, the computer program can be downloaded and installed from the network through the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above functions defined in the method of some embodiments of the present application are executed.

[0088] It should be noted that the computer-readable media described in some embodiments of the present application may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present application, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0089] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed network.

[0090] The above computer-readable medium may be included in the above electronic device; or it may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device is caused to: obtain the user access log of the first user within a predetermined historical time period; analyze the user access log to obtain a set of common profile features of the first user; preload the profile data corresponding to each profile feature in the set of common profile features into a cache area; when the user initiates a profile retrieval request, first query from the cache area. If the cache hits, read the corresponding profile data and return it to the first user; if the cache misses, query and read the corresponding profile data from the storage area and return it to the first user.

[0091] Computer program code for performing the operations of some embodiments of the present application may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages - such as Java, Smalltalk, C++; and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0092] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0093] The modules described in some embodiments of the present application can be implemented in software or in hardware. The described modules can also be provided in a processor. For example, they can be described as: a first acquisition module, a first analysis module, a first loading module, a first query module, and a first adjustment module. Among them, the names of these modules do not constitute a limitation to the module itself in some cases. For example, the first acquisition module can also be described as an "access log acquisition module".

[0094] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.

[0095] Some embodiments of the present application also provide a computer program product, including a computer program which, when executed by a processor, implements any of the above file management methods.

[0096] Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it on the basis of the present invention, which will be obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection required by the present invention.

Claims

1. A file management method, characterized in that: The following steps are involved: In response to receiving a request for access from a first user, obtaining a user access log of the first user within a predetermined historical time period; Analyzing the user access log to obtain a common profile feature set of the first user; Preloading the archive data corresponding to each archive feature in the common archive feature set into a cache area; When a user initiates a file retrieval request, the cache area is first queried. If the cache hits, the corresponding file data is read and returned to the first user; If the cache does not hit, the corresponding archive data is queried and read from the storage area and returned to the first user.

2. The file management method according to claim 1, characterized in that: Obtain the user access logs of the first user in the past two days, where the user access logs are recorded independently every day.

3. The file management method according to claim 2, characterized in that: The step of analyzing the user access log to obtain the common profile feature set of the first user includes: Read the access records in the user access log one by one; Parsing each of the access records to identify a field containing a profile identifier and extracting the profile identifier; Determining profile characteristics of each access request based on the extracted plurality of profile identifiers; Perform access count statistics on the plurality of archive features, merge and sort them, and obtain access count statistics results; According to the access count statistics, the archive features with the top N access counts are selected and stored in a preset set to obtain the commonly used archive feature set.

4. The file management method according to claim 3, characterized in that: The steps of parsing each of the access records to identify a field containing an archive identifier and extracting the archive identifier include: Read each access record in the user access log line by line to obtain original text data of the record; Applying a preset deep learning model to perform word segmentation processing on the original text data to identify multiple fields; Determining whether each of the fields contains a specific identifier keyword, the keyword being used to indicate the field where the archive identifier is located; After identifying the fields containing the archive identifiers, the specific content in each field is extracted to obtain a plurality of the archive identifiers.

5. The file management method according to claim 4, characterized in that: The preset deep learning model is the Transformer model.

6. The file management method according to any one of claims 1 to 5, characterized in that: The following steps are also included: The file distribution of the storage area is dynamically adjusted according to the common file characteristics.

7. A file management system, characterized in that: include: A first acquisition module, configured to, in response to receiving an access request from a first user, acquire a user access log of the first user within a predetermined historical time period; A first analysis module, configured to analyze the user access log to obtain a common profile feature set of the first user; A first loading module, used for preloading the archive data corresponding to each archive feature in the common archive feature set into a cache area; The first query module is used to query from the cache area when the user initiates a file retrieval request. If the cache hits, the corresponding file data is read and returned to the first user; if the cache does not hit, the corresponding file data is queried and read from the storage area and returned to the first user.

8. An electronic device, characterized in that: include: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processing device, the method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processing device, the method according to any one of claims 1 to 6 is implemented.