A file management method and system based on cloud storage and artificial intelligence authentication
By analyzing the application scenarios and user groups of archival information through artificial intelligence authentication methods, the archival management in the cloud storage system is optimized, the problems of archival data redundancy and low query efficiency are solved, and efficient and convenient archival management is achieved.
Patent Information
- Application Number
- CN202411809511.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2044-12-10
AI Technical Summary
The existing cloud storage archive management has problems such as archive data redundancy, insufficient content query convenience and low index query efficiency. In particular, the utilization rate of old archive data is low and errors are prone to occur during query.
Through artificial intelligence authentication methods, the application scenarios and user groups of archival information are analyzed, the scenario utilization rate and user utilization rate are calculated, low-utilization archives are identified, and channel optimization, content adjustment and location improvement plans are generated to optimize archive management.
Reduce the redundancy of archival data, improve the convenience of content query and the efficiency of index query, and ensure the efficiency and accuracy of archival management.
Smart Images

Figure CN119739915B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a file management method and system based on cloud storage and artificial intelligence authentication, belonging to the technical field of file management. Background Art
[0002] With the popularization of cloud storage, most existing archives have been separated from the paper-based archive management method and have been stored in cloud systems. The method of archive management through cloud storage is efficient and reliable, but there are still some problems.
[0003] First, when users upload paper archive records in the warehouse to the cloud, they will not only upload the recently added archive data, but also upload old archive data from many years ago. With the development of the times, some old archive data may be unavailable because they cannot adapt to the development of the times, so they are "accumulated dust" in the cloud system. Therefore, archive management needs to consider the archive data redundancy brought by these old archive data. Secondly, the archive layout, content and order of querying archives (using the mouse to click on the A archive data on the front-end interface to find the B archive data, and then from the B archive data to find the C archive data) may not be convenient for users to read. At this time, some professionals responsible for data collection and data statistics may use their personal In the name of , the layout, content and sequence of these archives are reconstructed to attract users to watch and earn traffic. This process shows that the layout, content and reading sequence of the archives in the cloud system are not suitable for users. Therefore, the layout, content and sequence of the archives need to be optimized to increase the number of times users enter the cloud system to query archives. Finally, when managing archives, an index structure of the archives will be constructed, and archive data will be queried through multi-level indexes. When users use the index structure to query data, the names of the two indexes E and F are similar, which may cause users to want to query archives in the F index but mistakenly query archives in the E index. At this time, the index of the archives also needs to be managed.
[0004] Therefore, there is an urgent need for a solution that can reduce the redundancy of archival data, improve the convenience of content search and the efficiency of index search. Summary of the Invention
[0005] The present invention provides a file management method and system based on cloud storage and artificial intelligence authentication, the main purpose of which is to improve the intelligence of ventilator data analysis and treatment recommendations.
[0006] To achieve the above objectives, the present invention provides a cloud storage and artificial intelligence authentication file management method, comprising:
[0007] Obtaining archival information from a cloud storage system, identifying the actual time of invocation of the archival information, collecting activity scenario data at the actual time of invocation, analyzing the application scenario of the archival information from the activity scenario data, and querying the user group of the archival information in the reference scenario based on an artificial intelligence authentication method;
[0008] Calculating the scenario utilization rate of the archival information based on the application scenario, calculating the user utilization rate of the archival information according to the user group, and determining the total utilization rate of the archival information by using the scenario utilization rate and the user utilization rate;
[0009] Extracting low-utilization files from the file information using the total utilization, obtaining a regular calling channel for the low-utilization file, identifying an actual calling channel for the low-utilization file, and determining whether the low-utilization file is called by a regular channel using the actual calling channel and the regular calling channel;
[0010] When the low-utilization file is not called by a regular channel, generating a channel optimization plan for the low-utilization file in the file information;
[0011] When the low-utilization file is called by a regular channel, constructing a content adjustment plan for the low-utilization file in the file information;
[0012] Assigning an information serial number to the archival information, analyzing a serial number collision value of the archival information based on the information serial number, and generating a position improvement plan for the archival information using the serial number collision value;
[0013] The channel optimization scheme, the content adjustment scheme and the location improvement scheme are used to perform archive management on the archive information in the cloud storage system to obtain an archive management result of the archive information.
[0014] Optionally, the calculating the scenario utilization rate of the archive information based on the application scenario includes:
[0015] Setting the unit call time period of the archive information;
[0016] Extract the number of unit calls of the archive information within the unit call period;
[0017] Query the application scenario time period of the application scenario;
[0018] Obtaining a scene calling period belonging to the application scene period in the unit calling period;
[0019] Extract the scene call times belonging to the scene call period from the unit call times:
[0020] Calculating the scene density of the archive information within the application scene period based on the unit call count and the scene call count;
[0021] The scene density is used as the scene utilization rate of the archive information.
[0022] Optionally, calculating the user utilization rate of the profile information according to the user group includes:
[0023] Setting the characteristic category of the archival information;
[0024] Extract the number of category calls of the archive information within the feature category;
[0025] Recording the identity characteristics of the user group;
[0026] Obtaining user features belonging to the identity features in the feature category;
[0027] Extracting the feature call count belonging to the user feature from the category call count;
[0028] Calculating the feature density of the profile information with respect to the identity feature based on the category call count and the feature call count;
[0029] The feature density is used as the user utilization rate of the profile information.
[0030] Optionally, the determining whether the low-utilization file is called by a regular channel by using the actual calling channel and the regular calling channel includes:
[0031] Selecting a first calling channel among the actual calling channels that is consistent with the regular calling channel;
[0032] Selecting a second calling channel in the actual calling channel that is inconsistent with the regular calling channel;
[0033] Query the first call count and the second call count of the first call channel and the second call channel respectively;
[0034] Calculating a regular call rate of the low-utilization file according to the first call number and the second call number;
[0035] When the regular call rate is greater than the preset call rate threshold, it is determined that the low-utilization file is called by the regular channel; when the regular call rate is not greater than the preset call rate threshold, it is determined that the low-utilization file is not called by the regular channel.
[0036] Optionally, the generating of a channel optimization solution for low-utilization archives in the archive information includes:
[0037] Querying the informal files when the low-utilization files are not called by the formal channels;
[0038] identifying a first file format, a first index order, and a first file keyword of the informal file;
[0039] Obtaining a second file format, a second index sequence, and a second file keyword for the low-utilization file;
[0040] Calculating the layout similarity between the first file layout and the second file layout;
[0041] Calculating a sequence similarity between the first index sequence and the second index sequence;
[0042] Calculating keyword similarity between the first archive keyword and the second archive keyword;
[0043] Using the typesetting similarity, the sequence similarity and the keyword similarity, respectively, the second file typesetting, the second index sequence and the second file keywords are screened for abnormal file typesetting, abnormal index sequence and abnormal file keywords;
[0044] The first file layout, the first index sequence and the first file keyword are used to set a channel optimization solution corresponding to the abnormal file layout, the abnormal index sequence and the abnormal file keyword.
[0045] Optionally, the step of constructing a content adjustment plan for low-utilization archives in the archive information includes:
[0046] Performing information segmentation on the archival information to obtain segmented information;
[0047] Acquire non-low-utilization files in the file information;
[0048] querying whether the segmentation information exists in the non-low utilization file;
[0049] When the segmentation information exists in the non-low-utilization file, distinguishing existing information from non-existing information in the segmentation information;
[0050] Setting a first content adjustment scheme of deleting the existing information from the low-utilization archive;
[0051] Constructing a transition probability matrix of the non-existence information;
[0052] Based on the transition probability matrix, constructing an inference probability model of the non-existence information;
[0053] When the value of the inference probability model is greater than a preset probability threshold, determining whether the amount of inferred information in the inference probability model is lower than a preset amount threshold;
[0054] When the amount of inferred information in the inference probability model is lower than a preset amount threshold, deleting the inferred information from the non-existent information as a second content adjustment scheme;
[0055] The first content adjustment scheme and the second content adjustment scheme are used to determine a content adjustment scheme for low-utilization files in the file information.
[0056] Optionally, allocating an information serial number to the archival information includes:
[0057] An information node for collecting the archival information;
[0058] Sort the information nodes according to the superior-subordinate relationship and the peer relationship in the information nodes to obtain a node sequence;
[0059] Assign node values to the node sequence in ascending order to obtain information sequence numbers.
[0060] Optionally, analyzing the sequence number collision value of the archive information based on the information sequence number includes:
[0061] Collect the number of information calls corresponding to the first serial number in the information serial number;
[0062] Querying target information called after calling the information corresponding to the first sequence number within a continuous period;
[0063] Obtaining a second serial number of the target information from the information serial number;
[0064] When the second serial number is not greater than the first serial number, determining a serial number collision value of the archive information;
[0065] When the second serial number is greater than the first serial number, a serial number collision value of the archive information is determined.
[0066] Optionally, the position improvement scheme for generating the archive information using the sequence number collision value includes:
[0067] Get the sequence number collision value when the second sequence number is greater than the first sequence number;
[0068] Determining whether the second sequence number is a leaf sequence number of the first sequence number;
[0069] When the second serial number is a leaf serial number of the first serial number, generating a location improvement plan for the archive information;
[0070] When the second sequence number is not the leaf sequence number of the first sequence number, exchanging the node structure of the first sequence number with the node structure of the second sequence number is used as a position improvement solution.
[0071] In order to solve the above problems, the present invention also provides a cloud storage and artificial intelligence authentication file management system, the system comprising:
[0072] A user query module is used to obtain archival information from the cloud storage system, identify the actual call time of the archival information, collect activity scene data at the actual call time, analyze the application scenario of the archival information from the activity scene data, and query the user group of the archival information in the reference scenario based on artificial intelligence authentication methods;
[0073] a utilization rate determination module, configured to calculate a scenario utilization rate of the archival information based on the application scenario, calculate a user utilization rate of the archival information based on the user group, and determine a total utilization rate of the archival information using the scenario utilization rate and the user utilization rate;
[0074] a regularity determination module, configured to extract low-utilization files from the file information using the total utilization, obtain regular call channels for the low-utilization files, identify actual call channels for the low-utilization files, and determine whether the low-utilization files are called by regular channels using the actual call channels and the regular call channels;
[0075] a channel optimization module, configured to generate a channel optimization plan for the low-utilization file in the file information when the low-utilization file is not called by a regular channel;
[0076] A content construction module, configured to construct a content adjustment plan for the low-utilization archive in the archive information when the low-utilization archive is called by a regular channel;
[0077] a position generation module, configured to assign an information serial number to the archival information, analyze a serial number collision value of the archival information based on the information serial number, and generate a position improvement plan for the archival information using the serial number collision value;
[0078] The file management module is used to perform file management on the file information in the cloud storage system by using the channel optimization solution, the content adjustment solution and the location improvement solution to obtain the file management result of the file information.
[0079] Compared with the problem described in the background technology, the embodiment of the present invention calculates the scenario utilization rate of the archival information based on the application scenario, so as to use the scenario utilization rate to characterize whether a certain archival information is frequently mobilized in the application scenario. In non-application scenarios, the user will not need this archival information and will not call it. Therefore, the archival call rate when the user has a demand for this archival information is counted to characterize whether this archival information is useful, that is, whether the user needs this archival data. Furthermore, the embodiment of the present invention determines whether the low-utilization archive is called by the regular channel by using the actual call channel and the regular call channel, so as not to immediately determine that the archival information is invalid data that the user does not need when the archival information has a low utilization rate, but to determine whether the user is calling because the channel is blocked by identifying the channel through which the user queries this archival information. Failure to find archival information. If the user cannot find archival information because the query channel is blocked, it means that the low utilization rate of archival information is not because the user does not need it. The embodiment of the present invention constructs a content adjustment scheme for low-utilization archives in the archival information to remove redundant data that the user does not need. Furthermore, the embodiment of the present invention analyzes the sequence number collision value of the archival information based on the information sequence number to identify whether the user will exit the window after accessing the archival data in a certain front-end interface window and access other windows with similar names. If accessed, it means that the archive in the window accessed for the first time is the archive that the user mistakenly clicked in to view. In other words, if the user mistakenly clicked in to view, a conflict value and a collision value will be generated. Furthermore, the embodiment of the present invention generates a location improvement scheme for the archival information by using the sequence number collision value to improve the efficiency of index query. Therefore, the cloud storage and artificial intelligence authentication-based archive management method proposed by the present invention can reduce archival data redundancy, improve content query convenience and index query efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] Figure 1 A flowchart of a cloud storage and artificial intelligence authentication file management method according to an embodiment of the present invention;
[0081] Figure 2 A schematic diagram of modules for implementing the cloud storage and artificial intelligence authentication file management method provided in one embodiment of the present invention.
[0082] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0083] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0084] The present embodiment provides a method for managing archives based on cloud storage and artificial intelligence authentication. The method can be executed by at least one of a server, a terminal, or other electronic device capable of executing the method provided by the present embodiment. In other words, the method can be executed by software or hardware installed on a terminal or server. The server can include, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster.
[0085] Example 1:
[0086] Reference Figure 1 FIG. 1 is a flow chart of a method for managing archives based on cloud storage and artificial intelligence authentication according to an embodiment of the present invention. In this embodiment, the method for managing archives based on cloud storage and artificial intelligence authentication includes:
[0087] S1. Obtain archival information from a cloud storage system, identify the actual call time of the archival information, collect activity scene data at the actual call time, analyze the application scenario of the archival information from the activity scene data, and query the user group of the archival information in the reference scenario based on an artificial intelligence authentication method.
[0088] In an embodiment of the present invention, the archival information refers to archival data applied to a certain scenario, such as the treatment file of a patient in a medical scenario. The actual call time refers to the moment when the archival information is called and queried by the front-end user. The activity scenario data refers to the activity data that occurs at the actual call time. For some archival data that is not often called, such as personal tax files, only when personal income tax data needs to be counted every year, the user calls his own tax payment file. That is to say, the activity of needing to count personal income tax data every year is the activity scenario data. The application scenario refers to the activity related to the archival information in the activity scenario data. For example, the aforementioned personal tax statistical activity is not related to the archival data in the medical scenario, so the personal tax statistical activity cannot be used as an application scenario. Furthermore, the artificial intelligence authentication method refers to authenticating whether the user's identity is legal when the front-end user logs in to the front-end system of the archival data, and recording the time, age, gender and other data when the user calls the file.
[0089] Optionally, the process of querying the user group of the archival information in the reference scenario based on the artificial intelligence authentication method refers to querying the users who make archival calls and the time, age, gender and other data of these users who call the archives.
[0090] The user group refers to users who have called up files and include data such as time, age, and gender.
[0091] S2. Based on the application scenario, calculate the scenario utilization rate of the archive information, calculate the user utilization rate of the archive information according to the user group, and determine the total utilization rate of the archive information using the scenario utilization rate and the user utilization rate.
[0092] The embodiment of the present invention calculates the scenario utilization rate of the archival information based on the application scenario, so as to use the scenario utilization rate to characterize whether a certain archival information is frequently mobilized in the application scenario. In non-application scenarios, this archival information will not be called because there is no demand for it. Therefore, the archival call rate when the user has demand for this archival information is counted to characterize whether this archival information is useful, that is, whether the user needs this archival data.
[0093] In one embodiment of the present invention, the scenario utilization rate of the archival information is calculated based on the application scenario, including: setting a unit call time period for the archival information; extracting the number of unit calls of the archival information within the unit call time period; querying the application scenario time period of the application scenario; obtaining the scene call time period belonging to the application scenario time period in the unit call time period; extracting the number of scene calls belonging to the scene call time period from the unit call number; and calculating the scenario density of the archival information within the application scenario time period based on the unit call number and the scene call number using the following formula:
[0094]
[0095] Among them, a represents the scene density, x i represents the number of unit calls in the i-th unit call period, n represents the total number of unit call periods, represents the number of scene calls in the jth scene call period of the kth application scenario, and m represents the total number of scene call periods of the kth application scenario;
[0096] The scene density is used as the scene utilization rate of the archive information.
[0097] Among them, the unit call period refers to a series of continuous periods in a certain year or month. For example, if each week of a certain year is a unit call period, then the first week, the second week, etc. of a certain year are regarded as a series of continuous unit call periods. The number of unit calls refers to the total number of times this file information is called within a unit call period.
[0098] In one embodiment of the present invention, the calculating of the user utilization rate of the profile information based on the user group includes: setting a feature category of the profile information; extracting the number of category calls of the profile information within the feature category; recording the identity features of the user group; obtaining user features belonging to the identity features in the feature category; extracting the number of feature calls belonging to the user features from the number of category calls; and calculating the feature density of the profile information with respect to the identity features based on the number of category calls and the number of feature calls using the following formula:
[0099]
[0100] Among them, b represents the feature density, x' i' represents the number of category calls for the i'th feature category, n' represents the total number of feature categories, represents the number of scene calls of the jth user feature of the k'th user group, and m represents the total number of user features of the kth user group;
[0101] The feature density is used as the user utilization rate of the profile information.
[0102] The feature category refers to the identity feature of the user when the profile information is called by the user, such as gender, age, etc., and the category call count refers to the number of calls belonging to the same feature category.
[0103] The total utilization is the weighted sum of the scenario utilization and the user utilization.
[0104] S3. Use the total utilization to extract low-utilization files from the file information, obtain the regular calling channel of the low-utilization files, identify the actual calling channel of the low-utilization files, and use the actual calling channel and the regular calling channel to determine whether the low-utilization files are called by the regular channel.
[0105] Optionally, the process of extracting low-utilization archives from the archive information using the total utilization rate can filter the total utilization rate by setting a utilization rate threshold to obtain low utilization rate, and then obtain archive information corresponding to the low utilization rate.
[0106] Furthermore, the embodiment of the present invention uses the actual calling channel and the regular calling channel to determine whether the low-utilization file is called by the regular channel, so that when the file information has low utilization, it does not immediately determine that the file information is invalid data that the user does not need, but instead determines whether the user cannot query the file information because the channel is blocked by identifying the channel through which the user queries the file information. If the user cannot query the file information because the query channel is blocked, it means that the low utilization of the file information is not because the user does not need it.
[0107] Among them, the formal calling channel is set up to query the archival information through the software and system where the archival information is located. The actual calling channel includes channels that are consistent with the formal calling channel and channels that are inconsistent with it. For example, some professionals who mainly monetize traffic use channels to obtain archival information. Simply put, the channel refers to the channel through which users obtain archival information.
[0108] In one embodiment of the present invention, the determining whether the low-utilization file is called by a regular channel using the actual call channel and the regular call channel includes: selecting a first call channel from the actual call channel that is consistent with the regular call channel; selecting a second call channel from the actual call channel that is inconsistent with the regular call channel; querying the first call count and the second call count of the first call channel and the second call channel respectively; and calculating the regular call rate of the low-utilization file according to the first call count and the second call count using the following formula:
[0109]
[0110] Where c represents the regular call rate, u represents the number of first calls, and v represents the number of second calls;
[0111] When the regular call rate is greater than the preset call rate threshold, it is determined that the low-utilization file is called by the regular channel; when the regular call rate is not greater than the preset call rate threshold, it is determined that the low-utilization file is not called by the regular channel.
[0112] The call rate threshold is set according to the actual scenario and will not be further described here.
[0113] S4. When the low-utilization file is not called by a regular channel, a channel optimization solution for the low-utilization file in the file information is generated.
[0114] The embodiment of the present invention generates a channel optimization solution for low-utilization archives in the archive information to improve the convenience of content query.
[0115] In one embodiment of the present invention, the channel optimization scheme for generating low-utilization archives in the archive information includes: querying informal archives when the low-utilization archives are not called by formal channels; identifying the first file typesetting, first index sequence and first file keywords of the informal archives; obtaining the second file typesetting, second index sequence and second file keywords of the low-utilization archives; calculating the typesetting similarity between the first file typesetting and the second file typesetting; calculating the sequence similarity between the first index sequence and the second index sequence; calculating the keyword similarity between the first file keyword and the second file keyword; using the typesetting similarity, the sequence similarity and the keyword similarity to respectively screen out abnormal file typesetting, abnormal index sequence and abnormal file keywords in the second file typesetting, second index sequence and second file keywords; and using the first file typesetting, the first index sequence and the first archive keywords to set the channel optimization scheme corresponding to the abnormal file typesetting, the abnormal index sequence and the abnormal file keywords.
[0116] Among them, the layout similarity is the similarity between the Sth position in the first file layout and the Sth position in the second file layout, and the calculation method is similar to the calculation of image pixel similarity. The sequential similarity refers to the similarity of the front and back order of the query file data. For example, first query the A file data, and then query the B file data at the next level of the A file data. The query process is a tree node traversal process, because the multi-level structure of the index is a tree structure. The abnormal file layout, abnormal index sequence and abnormal file keywords refer to the second file layout, second index sequence and second file keywords that are lower than the similarity threshold. The channel optimization scheme refers to correcting the abnormal file layout, abnormal index sequence and abnormal file keywords to the first file layout, first index sequence and first file keywords, that is, imitating the file management method of informal files.
[0117] S5. When the low-utilization file is called by a regular channel, a content adjustment plan for the low-utilization file in the file information is constructed.
[0118] The embodiment of the present invention constructs a content adjustment scheme for low-utilization archives in the archive information to remove redundant data that is not required by users.
[0119] In one embodiment of the present invention, constructing a content adjustment scheme for low-utilization archives in the archive information includes: performing information segmentation on the archive information to obtain segmentation information; obtaining non-low-utilization archives in the archive information; querying whether the segmentation information exists in the non-low-utilization archives; when the segmentation information exists in the non-low-utilization archives, distinguishing existing information from non-existing information in the segmentation information; setting a first content adjustment scheme for deleting the existing information from the low-utilization archives; and constructing a transition probability matrix for the non-existing information using the following matrix:
[0120]
[0121]
[0122] Among them, P represents the transition probability matrix, p IJ , p(J|I) represents the state transition probability of inferring the non-existence of information J from the non-existence of information I, p(IJ) represents the probability that the non-existence of information I and the non-existence of information J exist at the same time, p(I) represents the probability that the non-existence of information I accounts for the total non-existence of information, d represents the number of categories of non-existence information, represents the probability in the transition probability matrix;
[0123] Based on the transition probability matrix, the inference probability model of the non-existence information is constructed using the following formula:
[0124]
[0125] Among them, P(X t+1 ∣…X t-2 ,X t-1 ,X t ) represents the inference probability model, ...X t-2 ,X t-1 ,X t Indicates that the non-existence information contains X t+1 Any number t of non-existence information in the remaining non-existence information, X t+1 Indicates the non-existence information of the currently calculated inference probability value, p(…X t-2 ,X t-1 ,X t ,X t+1 )、p(…X t-2 ,X t-1 ,X t ) represents the probability calculated by the transition probability matrix;
[0126] When the value of the inference probability model is greater than a preset probability threshold, determine whether the amount of inferred information in the inference probability model is lower than the preset quantity threshold; when the amount of inferred information in the inference probability model is lower than the preset quantity threshold, delete the inferred information from the non-existent information as a second content adjustment scheme; use the first content adjustment scheme and the second content adjustment scheme to determine the content adjustment scheme for low-utilization archives in the archive information.
[0127] The segmentation information can be keywords or sentences, etc. It should be noted that, by constructing the inference probability model, it can be determined whether, when a part of the segmentation information exists, other segmentation information can be inferred based on this part of the known segmentation information. If so, only this part of the known segmentation information can be retained, and the other inferred segmentation information can be deleted. The probability threshold is used to filter the larger value in the inference probability model. This value represents the value that can be obtained by...X t-2 ,X t-1 ,X t To infer X t+1 , and the quantity threshold is set to 2, that is, when the amount of inferred information in the inference probability model is lower than the preset quantity threshold, it means that in the presence of ...X t-2 ,X t-1 ,X t Under such information conditions, only one P(X t+1 ∣…X t-2 ,X t-1 ,X t ), which can make it possible to have...X t-2 ,X t-1 ,X t Under such information conditions, there is only one inference result, namely X t+1 , and the inferred information is X t+1 .
[0128] S6. Assign an information serial number to the archival information, analyze a serial number collision value of the archival information based on the information serial number, and generate a position improvement plan for the archival information using the serial number collision value.
[0129] In one embodiment of the present invention, the assigning of information serial numbers to the archival information includes: collecting information nodes of the archival information; sorting the information nodes according to the superior-subordinate relationship and the peer relationship in the information nodes to obtain a node sequence; and assigning node values to the node sequence in ascending order to obtain an information serial number.
[0130] The information node includes a file collection, for example, a file collection of a certain type. In a medical scenario, a file collection belonging to the category of patient identity information is an information node.
[0131] Optionally, the process of assigning information serial numbers to the archive information is a process of assigning 1, 2, 3, 4, 5, ... to the information nodes of the tree structure from top to bottom and from left to right.
[0132] Furthermore, the embodiment of the present invention analyzes the serial number collision value of the archive information based on the information serial number to identify whether the user will exit the window after accessing the archive data in a front-end interface window and access other windows with similar names. If so, it means that the archive in the window accessed for the first time is the archive that the user accidentally clicked in to view. In other words, if the user accidentally clicked in to view, a conflict value and a collision value will be generated.
[0133] In one embodiment of the present invention, analyzing the sequence number collision value of the archive information based on the information sequence number includes: collecting the number of information calls corresponding to the first sequence number in the information sequence number; querying the target information called after calling the information corresponding to the first sequence number within a continuous time period; obtaining the second sequence number of the target information in the information sequence number; determining the sequence number collision value of the archive information when the second sequence number is not greater than the first sequence number; and determining the sequence number collision value of the archive information when the second sequence number is greater than the first sequence number.
[0134] Among them, the first serial number refers to the serial number clicked when accessing the archive information for the first time, and the target information refers to the next information called after calling the information corresponding to the first serial number. When the second serial number is not greater than the first serial number, the serial number collision value of the archive information is determined to be 0. When the second serial number is greater than the first serial number, the serial number collision value of the archive information is determined to be the number of collisions, that is, the number of times the user accidentally clicks in to view it.
[0135] Furthermore, the embodiment of the present invention generates a location improvement scheme for the archive information by utilizing the sequence number collision value, so as to improve the efficiency of index query.
[0136] In one embodiment of the present invention, the use of the serial number collision value to generate a position improvement plan for the archive information includes: obtaining a serial number collision value when the second serial number is greater than the first serial number; determining whether the second serial number is a leaf serial number of the first serial number; generating a position improvement plan for the archive information when the second serial number is a leaf serial number of the first serial number; and exchanging the node structure of the first serial number with the node structure of the second serial number as a position improvement plan when the second serial number is not a leaf serial number of the first serial number.
[0137] Among them, the leaf serial number refers to the serial number corresponding to the leaf node of the tree structure. When the second serial number is the leaf serial number of the first serial number, the location improvement plan for generating the archive information is no improvement. The node structure of the first serial number refers to the tree structure with the first serial number as the root node, and the node structure of the second serial number is the same.
[0138] S7. Performing archive management on the archive information in the cloud storage system using the channel optimization solution, the content adjustment solution, and the location improvement solution to obtain an archive management result for the archive information.
[0139] Compared with the problem described in the background technology, the embodiment of the present invention calculates the scenario utilization rate of the archival information based on the application scenario, so as to use the scenario utilization rate to characterize whether a certain archival information is frequently mobilized in the application scenario. In non-application scenarios, the user will not need this archival information and will not call it. Therefore, the archival call rate when the user has a demand for this archival information is counted to characterize whether this archival information is useful, that is, whether the user needs this archival data. Furthermore, the embodiment of the present invention determines whether the low-utilization archive is called by the regular channel by using the actual call channel and the regular call channel, so as not to immediately determine that the archival information is invalid data that the user does not need when the archival information has a low utilization rate, but to determine whether the user is calling because the channel is blocked by identifying the channel through which the user queries this archival information. Failure to find archival information. If the user cannot find archival information because the query channel is blocked, it means that the low utilization rate of archival information is not because the user does not need it. The embodiment of the present invention constructs a content adjustment scheme for low-utilization archives in the archival information to remove redundant data that the user does not need. Furthermore, the embodiment of the present invention analyzes the sequence number collision value of the archival information based on the information sequence number to identify whether the user will exit the window after accessing the archival data in a certain front-end interface window and access other windows with similar names. If accessed, it means that the archive in the window accessed for the first time is the archive that the user mistakenly clicked in to view. In other words, if the user mistakenly clicked in to view, a conflict value and a collision value will be generated. Furthermore, the embodiment of the present invention generates a location improvement scheme for the archival information by using the sequence number collision value to improve the efficiency of index query. Therefore, the cloud storage and artificial intelligence authentication-based archive management method proposed by the present invention can reduce archival data redundancy, improve content query convenience and index query efficiency.
[0140] Example 2:
[0141] like Figure 2 The figure shows a functional module diagram of a cloud storage and artificial intelligence authentication archive management system according to the present invention.
[0142] The cloud storage and artificial intelligence authentication archive management system 200 described in the present invention can be installed in an electronic device. Depending on the functionality implemented, the cloud storage and artificial intelligence authentication archive management system may include a user query module 201, a utilization determination module 202, a formality determination module 203, a channel optimization module 204, a content construction module 205, a location generation module 206, and an archive management module 207. A module, also referred to as a unit, is a series of computer program segments that can be executed by an electronic device processor and perform a fixed function. These modules are stored in the electronic device's memory.
[0143] In the embodiment of the present invention, the functions of each module / unit are as follows:
[0144] The user query module 201 is used to obtain archival information from the cloud storage system, identify the actual call time of the archival information, collect activity scene data at the actual call time, analyze the application scenario of the archival information from the activity scene data, and query the user group of the archival information in the reference scenario based on the artificial intelligence authentication method;
[0145] The utilization rate determination module 202 is configured to calculate the scenario utilization rate of the archive information based on the application scenario, calculate the user utilization rate of the archive information based on the user group, and determine the total utilization rate of the archive information using the scenario utilization rate and the user utilization rate;
[0146] The regularity determination module 203 is configured to extract low-utilization files from the file information using the total utilization, obtain regular call channels for the low-utilization files, identify actual call channels for the low-utilization files, and determine whether the low-utilization files are called by regular channels using the actual call channels and the regular call channels;
[0147] The channel optimization module 204 is configured to generate a channel optimization solution for the low-utilization file in the file information when the low-utilization file is not called by a regular channel;
[0148] The content construction module 205 is used to construct a content adjustment plan for the low-utilization file in the file information when the low-utilization file is called by a regular channel;
[0149] The position generating module 206 is configured to assign an information serial number to the archival information, analyze a serial number collision value of the archival information based on the information serial number, and generate a position improvement solution for the archival information using the serial number collision value;
[0150] The archive management module 207 is configured to perform archive management on the archive information in the cloud storage system using the channel optimization solution, the content adjustment solution, and the location improvement solution to obtain an archive management result for the archive information.
[0151] In detail, each module in the cloud storage and artificial intelligence authentication archive management system 200 described in the embodiment of the present invention adopts the same Figure 1 The technical means are the same as the cloud storage and artificial intelligence authentication archive management method described in, and can produce the same technical effects, so I will not go into details here.
[0152] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0153] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A file management method based on cloud storage and artificial intelligence authentication, characterized in that: The method comprises: Obtaining archival information from a cloud storage system, identifying the actual time of access to the archival information, collecting activity scenario data at the actual time of access, analyzing the application scenario of the archival information from the activity scenario data, and querying the user group of the archival information in the application scenario based on an artificial intelligence authentication method; Calculating the scenario utilization rate of the archival information based on the application scenario, calculating the user utilization rate of the archival information according to the user group, and determining the total utilization rate of the archival information by using the scenario utilization rate and the user utilization rate; Extracting low-utilization files from the file information using the total utilization, obtaining a regular calling channel for the low-utilization file, identifying an actual calling channel for the low-utilization file, and determining whether the low-utilization file is called by a regular channel using the actual calling channel and the regular calling channel; When the low-utilization file is not called by a regular channel, a channel optimization plan for the low-utilization file in the file information is generated; the channel optimization plan for the low-utilization file in the file information is generated, including: Querying the informal files when the low-utilization files are not called by the formal channels; identifying a first file format, a first index order, and a first file keyword of the informal file; Obtaining a second file layout, a second index sequence, and a second file keyword for the low-utilization file; Calculating the layout similarity between the first file layout and the second file layout; Calculating a sequence similarity between the first index sequence and the second index sequence; Calculating keyword similarity between the first archive keyword and the second archive keyword; Using the typesetting similarity, the sequence similarity and the keyword similarity, respectively, the second file typesetting, the second index sequence and the second file keywords are screened for abnormal file typesetting, abnormal index sequence and abnormal file keywords; Using the first file layout, the first index order, and the first file keyword, a channel optimization solution corresponding to the abnormal file layout, the abnormal index order, and the abnormal file keyword is set; When the low-utilization file is called by a regular channel, constructing a content adjustment plan for the low-utilization file in the file information; Assigning an information serial number to the archival information, analyzing a serial number collision value of the archival information based on the information serial number, and generating a position improvement plan for the archival information using the serial number collision value; The step of allocating an information serial number to the archival information includes: An information node for collecting the archival information; Sort the information nodes according to the superior-subordinate relationship and the peer relationship in the information nodes to obtain a node sequence; Assign node values to the node sequence in ascending order to obtain information sequence numbers; The analyzing the sequence number collision value of the archive information based on the information sequence number includes: Collect the number of information calls corresponding to the first serial number in the information serial number; Querying target information called after calling the information corresponding to the first sequence number within a continuous period; Obtaining a second serial number of the target information from the information serial number; When the second serial number is not greater than the first serial number, determining a serial number collision value of the archive information; When the second serial number is greater than the first serial number, determining a serial number collision value of the archive information; The location improvement scheme for generating the archive information by using the sequence number collision value includes: Get the sequence number collision value when the second sequence number is greater than the first sequence number; Determining whether the second sequence number is a leaf sequence number of the first sequence number; When the second sequence number is a leaf sequence number of the first sequence number, generating a position improvement plan for the archive information, the position improvement plan being no improvement; When the second sequence number is not the leaf sequence number of the first sequence number, the node structure of the first sequence number and the node structure of the second sequence number are exchanged as a position improvement scheme; The first serial number refers to the serial number clicked when accessing the archive information for the first time, and the target information refers to the next information called after calling the information corresponding to the first serial number; the node structure of the first serial number refers to the tree structure with the first serial number as the root node, and the same applies to the node structure of the second serial number; The channel optimization scheme, the content adjustment scheme and the location improvement scheme are used to perform archive management on the archive information in the cloud storage system to obtain an archive management result of the archive information.
2. The cloud storage and artificial intelligence authentication file management method according to claim 1, characterized in that: The calculating the scenario utilization rate of the archive information based on the application scenario includes: Setting the unit call time period of the archive information; Extract the number of unit calls of the archive information within the unit call period; Query the application scenario time period of the application scenario; Obtaining a scene calling period belonging to the application scene period in the unit calling period; Extract the scene call times belonging to the scene call period from the unit call times: Calculating the scene density of the archive information within the application scene period based on the unit call count and the scene call count; The scene density is used as the scene utilization rate of the archive information.
3. The cloud storage and artificial intelligence authentication file management method according to claim 1, characterized in that: Calculating the user utilization rate of the archive information according to the user group includes: Setting the characteristic category of the archival information; Extract the number of category calls of the archive information within the feature category; Recording the identity characteristics of the user group; Obtaining user features belonging to the identity features in the feature category; Extracting the feature call count belonging to the user feature from the category call count; Calculating the feature density of the profile information with respect to the identity feature based on the category call count and the feature call count; The feature density is used as the user utilization rate of the profile information.
4. The cloud storage and artificial intelligence authentication file management method according to claim 1, characterized in that: The determining whether the low-utilization file is called by a regular channel by using the actual calling channel and the regular calling channel includes: Selecting a first calling channel among the actual calling channels that is consistent with the regular calling channel; Selecting a second calling channel in the actual calling channel that is inconsistent with the regular calling channel; Query the first call count and the second call count of the first call channel and the second call channel respectively; Calculating a regular call rate of the low-utilization file according to the first call number and the second call number; When the regular call rate is greater than the preset call rate threshold, it is determined that the low-utilization file is called by the regular channel; when the regular call rate is not greater than the preset call rate threshold, it is determined that the low-utilization file is not called by the regular channel.
5. The cloud storage and artificial intelligence authentication file management method according to claim 1, characterized in that: The content adjustment scheme for low-utilization archives in the archive information is constructed, including: Performing information segmentation on the archival information to obtain segmented information; Acquire non-low-utilization files in the file information; querying whether the segmentation information exists in the non-low utilization file; When the segmentation information exists in the non-low-utilization file, distinguishing existing information from non-existing information in the segmentation information; Setting a first content adjustment scheme to delete the existing information from the low-utilization archive; Constructing a transition probability matrix of the non-existence information; Based on the transition probability matrix, constructing an inference probability model of the non-existence information; When the value of the inference probability model is greater than a preset probability threshold, determining whether the amount of inferred information in the inference probability model is lower than a preset amount threshold; When the amount of inferred information in the inference probability model is lower than a preset amount threshold, deleting the inferred information from the non-existent information as a second content adjustment scheme; The first content adjustment scheme and the second content adjustment scheme are used to determine a content adjustment scheme for low-utilization files in the file information.
6. A cloud storage and artificial intelligence authentication archive management system, which implements the method according to claim 1, characterized in that: The system comprises: A user query module is used to obtain archival information from the cloud storage system, identify the actual call time of the archival information, collect activity scene data at the actual call time, analyze the application scenario of the archival information from the activity scene data, and query the user group of the archival information in the reference scenario based on artificial intelligence authentication methods; a utilization rate determination module, configured to calculate a scenario utilization rate of the archival information based on the application scenario, calculate a user utilization rate of the archival information based on the user group, and determine a total utilization rate of the archival information using the scenario utilization rate and the user utilization rate; a regularity determination module, configured to extract low-utilization files from the file information using the total utilization, obtain regular call channels for the low-utilization files, identify actual call channels for the low-utilization files, and determine whether the low-utilization files are called by regular channels using the actual call channels and the regular call channels; a channel optimization module, configured to generate a channel optimization plan for the low-utilization file in the file information when the low-utilization file is not called by a regular channel; A content construction module, configured to construct a content adjustment plan for the low-utilization archive in the archive information when the low-utilization archive is called by a regular channel; a position generation module, configured to assign an information serial number to the archival information, analyze a serial number collision value of the archival information based on the information serial number, and generate a position improvement plan for the archival information using the serial number collision value; The file management module is used to perform file management on the file information in the cloud storage system by using the channel optimization solution, the content adjustment solution and the location improvement solution to obtain the file management result of the file information.
Citation Information
Patent Citations
Multi-level file system and construction method thereof
CN117807045A
Scene-adaptive index construction method and device, equipment and storage medium
CN117909548A