Data storage method based on artificial intelligence learning
Through the data storage method based on artificial intelligence learning, data is stored dynamically in-classified, which solves the problem of low data query efficiency in the existing technology, and achieves fast query and efficient storage.
Patent Information
- Application Number
- CN202510072211.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-17
AI Technical Summary
The prior art is less efficient in data storage and querying, especially when users need to quickly query large amounts of data, traversing the database may lead to longer query times.
Using a data storage method based on artificial intelligence learning, the data storage area is determined by obtaining the data to be stored and its calling information, and the storage area of the data is divided according to the data type and attributes to realize dynamic hierarchical storage.
It improves the efficiency of data query. Through dynamic hierarchical storage, data of different importance levels and types are stored separately, which facilitates quick query and call.
Smart Images

Figure CN119988520A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data storage, and in particular to a data storage method based on artificial intelligence learning. Background Art
[0002] With the rapid progress of information technology and the development of big data technology, all industries are facing the need to store and query a large amount of data. For example, in the digital system of an enterprise, a large amount of employee data, enterprise confidential data and enterprise financial data are stored. When an enterprise makes important decisions, it may need to store a large amount of data in a short period of time. At this time, the enterprise needs to store and manage a large amount of data in a unified manner for later data application.
[0003] In the related art, data is directly stored in the database, and when the data in the database needs to be queried, all the data in the database is directly traversed; however, when the user has an emergency need for data query, that is, the user needs to query in a short period of time, it may take a long time to traverse the database for query. It can be seen that the storage of data in the related art may lead to the problem of low data query efficiency. Summary of the invention
[0004] In order to solve the above technical problems, the present invention provides a data storage method based on artificial intelligence learning.
[0005] In a first aspect, the present application provides a data storage method based on artificial intelligence learning, comprising:
[0006] Acquire multiple data to be stored and call information of each of the data to be stored;
[0007] Determine the storage area corresponding to each of the data to be stored according to the calling information;
[0008] Acquire the data type and data attribute corresponding to each of the data to be stored, and divide each of the storage areas into a plurality of sub-storage areas according to the data type;
[0009] Based on the data attributes and all the sub-storage areas, all the data to be stored are stored.
[0010] In a preferred example, the present application may be further configured as follows: the calling information includes a plurality of calling time periods and calling frequencies corresponding to the calling time periods; and determining the storage area corresponding to each of the to-be-stored data according to the calling information includes:
[0011] For each of the to-be-stored data, based on all the retrieval time periods and the retrieval frequencies corresponding to the retrieval time periods, determining an average retrieval frequency;
[0012] Obtaining a set of retrieval frequencies corresponding to each of the storage areas;
[0013] The storage area corresponding to each of the data to be stored is determined according to the average of the retrieval frequencies of each of the data to be stored and the retrieval frequency sets corresponding to each of the storage areas.
[0014] In a preferred example, the present application may be further configured as follows: based on the data attributes and all the sub-storage areas, storing all the data to be stored includes:
[0015] For each of the sub-storage areas, obtaining the storage capacity corresponding to each of the sub-storage areas;
[0016] Calculating the correlation between the data to be stored, and determining a number of associated data of each of the data to be stored according to the correlation;
[0017] Generate a data item to be stored according to all the associated data of each of the data to be stored;
[0018] Obtaining the data amount of all data items to be stored, and determining whether the data amount is less than the storable amount;
[0019] If yes, then storing the data item to be stored in the corresponding sub-storage area;
[0020] If not, then modifying the storage capacity of each of the sub-storage areas according to the storage capacity and the data attribute to obtain a modified sub-storage area;
[0021] The data item to be stored is stored in the modified sub-storage area.
[0022] In a preferred example, the present application may be further configured as follows: calculating the correlation between the data to be stored includes:
[0023] Generating a plurality of data tags corresponding to each of the data to be stored, wherein the data tags are used to identify the data to be stored;
[0024] Determining a first relevance of each of the data to be stored based on a data label corresponding to each of the data to be stored;
[0025] Acquiring the retrieval time of the data to be stored, and determining the second relevance of each data to be stored based on the retrieval time of each data to be stored and a preset corresponding relationship;
[0026] A comprehensive correlation is determined based on the first correlation and the second correlation, and the comprehensive correlation is determined as the correlation between the data to be stored.
[0027] In a preferred example, the present application may be further configured as follows: determining the first relevance of each of the data to be stored based on the data label corresponding to each of the data to be stored includes:
[0028] Acquire application information of each of the data to be stored, and determine a weight value of each of the data tags according to the application information and all of the data tags of each of the data to be stored;
[0029] Determining a number of target data tags based on the weight values of the data tags and a preset weight value threshold;
[0030] For each pair of the data to be stored, determining the same data tags from all the target data tags;
[0031] Acquire the first label quantity of the target data label of each of the data to be stored and the second label quantity corresponding to the same data label;
[0032] Determine data similarity according to the first number of tags, the second number of tags, and all the same data tags;
[0033] The data similarity is determined as a first correlation of each of the data to be stored.
[0034] In a preferred example, the present application may be further configured as follows: determining the data similarity according to the first tag quantity, the second tag quantity and all the same data tags includes:
[0035] For each of the same data tags, based on the first tag quantity and the second tag quantity, determining the occurrence frequency of the same data tag;
[0036] Acquire the data volume of the data to be stored, and determine the inverse document frequency of the same data tag according to the occurrence frequency, the data volume of the data to be stored and a preset similarity calculation formula, wherein the inverse document frequency represents the prevalence of the same data tag;
[0037] Determining a vector value based on the occurrence frequency and the inverse document frequency;
[0038] Based on the vector value, determining the data distance between the data to be stored;
[0039] The data similarity is determined according to the data distance and the target corresponding relationship, wherein the target corresponding relationship is the corresponding relationship between the data distance and the data similarity.
[0040] In summary, this application includes the following beneficial technical effects:
[0041] The higher the retrieval frequency of the data to be stored is, the higher the number of times the data to be stored is used, indicating that the data to be stored is more important, and therefore it is necessary to obtain the data to be stored and the corresponding retrieval frequency; then determine the storage area of each data to be stored according to the retrieval frequency, and store the data to be stored in different areas for different importance levels for subsequent data calls; obtain the data type and data attribute of each data to be stored, and divide the storage area into several sub-storage areas according to the data type, so as to further subdivide the data storage area on the basis of data storage of the same importance; then divide the storage area into several sub-storage areas according to the data type for classified storage; on the basis of classified storage, further refine the storage according to the data characteristics; compared with the related art, the present application divides the data to be stored into multiple levels according to the retrieval frequency, data type and data attribute of the data, and stores different levels in different areas to realize dynamic hierarchical storage of the data to be stored, which is convenient for subsequent users to perform hierarchical data query, so as to effectively improve the data query efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 A schematic diagram of a scenario of a data storage method based on artificial intelligence learning provided in an embodiment of the present application;
[0043] Figure 2 A flowchart of a data storage method based on artificial intelligence learning provided in an embodiment of the present application;
[0044] Figure 3 A schematic diagram of the structure of a data storage device based on artificial intelligence learning provided in an embodiment of the present application;
[0045] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0046] The following is combined with Figure 1 To Attachment Figure 4 This application is described in further detail.
[0047] After reading this specification, those skilled in the art may make non-creative modifications to this embodiment as needed, but such modifications are protected by patent law as long as they are within the scope of the claims of this application.
[0048] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0049] In addition, the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article, unless otherwise specified, generally means that the associated objects before and after are in an "or" relationship.
[0050] The following is a further detailed description of the embodiments of the present application in conjunction with the accompanying drawings. Figure 1 As shown, it is a schematic diagram of a data storage scenario provided in an embodiment of the present application. A user generates a storage request on a user-side device and inputs the data to be stored. After receiving the storage request and the data to be stored, the electronic device performs a comprehensive analysis of the call information, data type and data attributes of the data to be stored, and dynamically stores the data to be stored in a hierarchical manner.
[0051] The embodiment of the present application provides a data storage method based on artificial intelligence learning, which is executed by an electronic device, which can be a server or a terminal device, wherein the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services. The terminal device can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc., but is not limited to this. The terminal device and the server can be directly or indirectly connected via wired or wireless communication, and the embodiment of the present application does not limit this. Figure 2 As shown, the method includes step S101, step S102, step S103 and step S104, wherein:
[0052] Step S101: Acquire multiple data to be stored and call information of each data to be stored.
[0053] Specifically, a monitoring program is pre-integrated in the electronic device, and the monitoring program is used to monitor the triggering behavior of the acquisition request. Once it is monitored that the acquisition request is triggered, the acquisition operation is performed. Specifically, the user can generate an acquisition request by clicking or voice. The data to be stored and the call information are both pre-entered by the user into the database to be stored. The data to be stored can be enterprise data or government data, etc. The embodiment of the present application does not limit the specific data to be stored. The call information includes multiple retrieval time periods for the data to be stored and the retrieval frequency corresponding to each retrieval time period.
[0054] Step S102: Determine the storage area corresponding to each data to be stored according to the call information.
[0055] Specifically, the call information includes multiple call time periods and the call frequencies corresponding to each call time period, wherein the call time period is composed of multiple call times; the storage area for each data to be stored can be determined according to the call information with reference to the following embodiments.
[0056] Step S103: obtaining the data type and data attribute corresponding to each data to be stored, and dividing each storage area into a plurality of sub-storage areas according to the data type.
[0057] Specifically, the data type of the data to be stored can be determined by a data type identification tool. In the embodiment of the present application, the data type of the data to be stored can be continuous data, discrete data, ordered data, time series data or text data, etc., which is not limited in the embodiment of the present application. Data attributes are used to describe the characteristics of data, such as enterprise customer data and enterprise product data, etc. The embodiment of the present application does not limit the specific data attributes. There are multiple sub-storage areas; each sub-area corresponds to a data attribute.
[0058] Step S104: store all the data to be stored based on the data attributes and all the sub-storage areas.
[0059] Specifically, the specific process of storing all the data to be stored based on the data attributes and all the sub-storage areas can refer to the following embodiments. It can be understood that storing the data to be stored in the corresponding sub-storage area according to the data attributes can achieve further refined storage and dynamic storage of the data to be stored, and then the data can be quickly called in the subsequent calling process.
[0060] Based on the above embodiment, when the frequency of retrieval of the data to be stored is higher, the higher the number of times the data to be stored is applied, indicating that the data to be stored is more important, and therefore it is necessary to obtain the data to be stored and the corresponding calling information; then determine the storage area of each data to be stored according to the calling information, and store the data to be stored in different areas with different importance levels to facilitate the subsequent calling of data; obtain the data type and data attribute of each data to be stored, and divide the storage area into several sub-storage areas according to the data type, so as to further subdivide the data storage area on the basis of data storage of equal importance; then divide the storage area into several sub-storage areas according to the data type for classified storage; on the basis of classified storage, further refine the storage according to the data characteristics; compared with the related art, the present application divides the data to be stored into multiple levels according to the calling information, data type and data attribute of the data, and stores different levels in different areas to realize dynamic hierarchical storage of the data to be stored, which is convenient for subsequent users to perform hierarchical data query, so as to effectively improve the data query efficiency.
[0061] Furthermore, the call information includes multiple call time periods and the call frequencies corresponding to each call time period, and the storage area corresponding to each data to be stored is determined according to the call information, including:
[0062] For each data to be stored, based on all retrieval periods and the retrieval frequencies corresponding to each retrieval period, determine the average retrieval frequency;
[0063] Get the retrieval frequency set corresponding to each storage area;
[0064] The storage area corresponding to each data to be stored is determined according to the average value of the retrieval frequency of each data to be stored and the retrieval frequency set corresponding to each storage area.
[0065] Specifically, the number of access time periods and the total access frequency of the access time period are determined, and the average access frequency is determined according to the calculation formula.
[0066] In an embodiment of the present application, the storage area includes a first storage area, a second storage area, and a third storage area; different storage areas have different storage priorities, and the regional priorities of the storage areas are pre-set by the technician; wherein the regional priority of the first storage area is higher than the storage priority of the second storage area, and the regional priority of the second storage area is higher than the regional priority of the third storage area. Among them, the storage data corresponding to the first storage area are all high-frequency call data, the storage data corresponding to the second storage area are all medium-frequency call data, and the storage data corresponding to the third storage area are all low-frequency call data, that is, the retrieval frequency of the stored data in the first storage area is greater than the retrieval frequency of the stored data in the second storage area, and the retrieval frequency of the stored data in the second storage area is greater than the retrieval frequency of the stored data in the third storage area. The retrieval frequency set corresponding to each storage area is pre-set by the technician, and the retrieval frequency set is composed of multiple retrieval frequency values. Then match the retrieval frequency mean and the retrieval frequency set one by one to obtain the storage area corresponding to the data to be stored. It can be understood that when the frequency of retrieval of the data to be stored is high, it indicates that the data to be stored is more important. When used later, important data will be scheduled from this storage area first. Compared with traversing all storage areas, this application can effectively improve the retrieval efficiency.
[0067] Based on the above embodiment, the retrieval frequency mean is determined according to the retrieval period and the corresponding retrieval frequency, so that the retrieval frequency mean can represent the overall retrieval frequency of the data to be stored; then the retrieval frequency set of each storage area is obtained, and the storage area corresponding to each data to be stored is determined according to the retrieval frequency mean and the retrieval frequency set, so as to effectively improve the accuracy of the division of the area to be stored.
[0068] Furthermore, based on the data attributes and all sub-storage areas, all the data to be stored are stored, including:
[0069] For each sub-storage area, obtaining the storage capacity corresponding to each sub-storage area;
[0070] Calculate the correlation between the data to be stored, and determine a number of associated data of each data to be stored according to the correlation;
[0071] Generate a data item to be stored according to all associated data of each data to be stored;
[0072] Obtain the data volume of all data items to be stored, and determine whether the data volume is less than the storable volume;
[0073] If yes, the data item to be stored is stored in the corresponding sub-storage area;
[0074] If not, the storage capacity of each sub-storage area is corrected according to the storage capacity and the data attribute to obtain a corrected sub-storage area;
[0075] The data item to be stored is stored in the modified sub-storage area.
[0076] Specifically, the storable amount of each sub-storage area is pre-set by a technician, and the embodiments of the present application do not limit the storable amount. The specific process of calculating the correlation between the data to be stored can refer to the following embodiments. The correlation is compared with a preset correlation threshold. If the correlation is greater than the preset correlation threshold, the data to be stored is determined as the associated data of another data to be stored, and the data to be stored that are associated with each other are generated as data items (including at least two data to be stored) for storage. It can be understood that storing the data to be stored in the form of data items during the storage process helps to maintain the consistency and integrity of the data to be stored, and when the user needs to query and access it later, he only needs to perform one operation to access all related data, which effectively improves the efficiency of data query. The data volume of the data item to be stored is input by the user. If the amount of data is less than the storage capacity of the sub-storage area, the data to be stored can be directly stored; if not, the storage capacity of the sub-storage area is corrected. The specific process of correcting the storage capacity of each sub-storage area includes: obtaining the attribute weight value of the preset data attribute, and correcting the storage capacity of the sub-storage area according to the corresponding relationship between the attribute weight value and the corrected storage capacity to obtain the corrected storage capacity, and then storing the corrected data items to be stored in the corresponding sub-storage area. The correspondence between the attribute weight value and the corrected storage capacity is pre-set by the technician. It can be understood that the attribute weight value is used to describe the importance of the data attribute. As the attribute weight value increases, the corresponding data to be stored becomes more important, and the amount of important data to be stored will also increase accordingly. Therefore, it is necessary to appropriately expand the storage capacity to avoid data loss due to insufficient storage capacity.
[0077] Based on the above embodiment, the storable amount of each sub-storage area is obtained and the storage capacity of the sub-storage area is corrected according to the storable amount and data attributes. The importance level of data with different characteristics is also different. In order to avoid data loss due to storage failure of more important data, the storage capacity needs to be corrected so as to store the complete data to be stored; the correlation between the data to be stored is calculated and the associated data of the data to be stored is determined according to the correlation, and then the data items to be stored are generated according to all the associated data of all the data to be stored, thereby realizing the associated storage between the data; when querying the data subsequently, only one query is needed to realize the query of all related data, which effectively improves the query efficiency.
[0078] Furthermore, the correlation between the data to be stored is calculated, including:
[0079] Generate a number of data tags corresponding to each data to be stored, the data tags are used to identify the data to be stored;
[0080] Determining a first relevance of each piece of data to be stored based on a data label corresponding to each piece of data to be stored;
[0081] Acquire the retrieval time of the data to be stored, and determine the second relevance of each data to be stored based on the retrieval time of each data to be stored and a preset corresponding relationship;
[0082] A comprehensive correlation is determined based on the first correlation and the second correlation, and the comprehensive correlation is determined as the correlation between the data to be stored.
[0083] Specifically, data tags are used to identify the data to be stored. For example, when the data to be stored is enterprise product data, the corresponding data tags may be product sales status, popularity, etc. Data tags may be generated by a preset data tag annotation model. The data to be stored is input into the data tag annotation model, and the data tag annotation model outputs the data to be stored and its corresponding data tags. Each data tag for the data to be stored may be one or more. The data label annotation model is obtained by training based on multiple training data. The training process of the data label annotation model includes: inputting multiple training storage data and their corresponding sample data labels into the time series neural network model, the time series neural network model outputs the training data labels corresponding to each training storage data, and calculates the loss value between the training data label and the sample data label according to the training data, the sample data label and the loss value function, and determines whether the loss value is greater than the preset loss value threshold. If the loss value is not greater than the preset loss value threshold, continue to use the training storage data and the corresponding sample data label to train the time series neural network model until the loss value is greater than the preset loss value threshold; if the loss value is greater than the preset loss value threshold, continue to train the time series neural network model for a preset number of times, and calculate the average loss value after the loss value is greater than the preset loss value threshold, and then calculate the loss value variance based on the average loss value, and compare the loss value variance and the variance threshold. If the loss value variance is less than the variance threshold, stop training and determine the time series neural network model as the data label annotation model; otherwise, continue to train the time series neural network model until the loss value variance is less than the variance threshold. The above loss value function can be a mean square error loss function, a binary cross entropy loss function, etc., and the preset loss value threshold and variance threshold are both preset by the technician. It can be understood that when the loss value variance is less than the variance threshold, it indicates that the change amplitude of the loss value is small at this time, that is, the time series neural network model can output data labels with higher accuracy. The specific process of determining the first correlation between the data to be stored according to the data label of the data to be stored can refer to the following embodiment. The retrieval time of the data to be stored is sent by the user to the electronic device, and the retrieval time is a single time value; the preset corresponding relationship is set by the technician based on work experience. The specific process of determining the comprehensive correlation according to the first correlation and the second correlation includes: obtaining a first weight value corresponding to the first correlation and a second weight value corresponding to the second correlation, and obtaining a comprehensive correlation according to the first weight value, the first correlation, the second weight value and the second correlation, that is, the comprehensive correlation = first weight value * first correlation + second weight value * second correlation. In an embodiment of the present application, the first weight value is greater than the second weight value; it can be understood that assigning different weight values to the first correlation and the second correlation can calculate the comprehensive correlation in a targeted manner to obtain a more accurate comprehensive correlation, and determine the comprehensive correlation as the correlation between the two data to be stored.
[0084] Based on the above embodiments, data labels are helpful to further refine and distinguish the data to be stored, and thus it is necessary to generate data labels for the data to be stored so as to analyze the correlation between the data to be stored from a micro perspective, and then determine the first correlation based on the data labels of the data to be stored to effectively improve the accuracy of the first correlation; then determine the second correlation based on the retrieval time and the corresponding relationship, and determine the comprehensive correlation based on the first correlation and the second correlation, thereby achieving a comprehensive analysis of the correlation between the data to be stored from different dimensions, thereby improving the accuracy of determining the correlation between the data to be stored.
[0085] Further, based on the data labels corresponding to the data to be stored, determining the first relevance of the data to be stored includes:
[0086] Acquire application information of each data to be stored, and determine a weight value of each data tag according to the application information and all the data tags corresponding to each data to be stored;
[0087] Determine a number of target data labels based on the weight value of each data label and a preset weight value threshold;
[0088] For each pair of data to be stored, determine the same data labels from all target data labels;
[0089] Obtain the first label quantity of the target data label of each data to be stored and the second label quantity corresponding to the same data label;
[0090] Determine data similarity based on the number of first tags, the number of second tags, and all identical data tags;
[0091] The data similarity is determined as the first correlation of each data to be stored.
[0092] Specifically, the application information of the data to be stored represents the application scenario identifier of the data to be stored. For example, when the data to be stored is enterprise product data, the corresponding application scenario can be an application scenario such as enterprise annual display or product revenue accounting; the application information can be sent synchronously when the user sends the data to be stored to the electronic device. The specific process of determining the weight value of the data tag according to the application information and the data tag includes: obtaining the preset weight value corresponding to each application scenario identifier, matching the data tag with all application scenario identifiers through a semantic matching algorithm, and then determining the weight value of the application scenario identifier corresponding to the data tag as the weight value of the data tag. The weight value of each data tag is compared with the preset weight value threshold, and the data tag with a weight value greater than the preset weight value threshold is determined as the target data tag, and then the first relevance of the data to be stored is determined according to the representative data tag to improve the accuracy of the first relevance calculation. In addition, filtering the data tags by the weight value can also reduce the interference of irrelevant data tags to further improve the accuracy of the first relevance. The number of first tags and the number of second tags can be obtained from the data tag statistics library. The specific process of determining data similarity according to the first tag quantity, the second tag quantity and the same data tags may refer to the following embodiment, and the data similarity is determined as the first correlation between the data to be stored.
[0093] Based on the above embodiments, when the application scenarios are different, the degree to which the same data tag intuitively reflects the data to be stored is also different. Therefore, different data tags have different importance to the stored data. Therefore, it is necessary to determine the weight value of each data tag based on the application information; then determine the target data tag based on the weight value of the data tag and the preset weight value threshold, that is, use the more important data tag that can more intuitively reflect the data to be stored as a reference; determine the same data tag from all target data tags for every two data to be stored, and the more the number of the same data tags between the data to be stored, the higher the correlation between the data to be stored; then obtain the first tag number of the target data tag and the second tag number of the same data tag, and determine the data similarity based on the first tag number, the second tag data and the same data tag, and then determine the data similarity as the first correlation between the data to be stored.
[0094] Further, determining data similarity according to the number of first tags, the number of second tags, and all identical data tags includes:
[0095] For each identical data tag, determining the occurrence frequency of the identical data tag based on the first tag quantity and the second tag quantity;
[0096] Obtain the amount of data to be stored, and determine the inverse document frequency of the same data label based on the frequency of occurrence, the amount of data to be stored and a preset similarity calculation formula, where the inverse document frequency represents the prevalence of the same data label;
[0097] Determine the vector value based on the occurrence frequency and the inverse document frequency;
[0098] Based on the vector value, determine the data distance between the data to be stored;
[0099] The data similarity is determined according to the data distance and the target correspondence, and the target correspondence is the correspondence between the data distance and the data similarity.
[0100] Specifically, the frequency of occurrence of the same data tag can be determined according to a calculation formula. In the embodiment of the present application, the frequency of occurrence represents the frequency of occurrence of the same data tag in the target data tag; the calculation formula is: In the embodiment of the present application, the preset similarity calculation formula is: Among them, the preset value is a positive integer greater than 0. The embodiment of the present application does not limit the specific preset value, and it can be 1, 2, etc.; it can be understood that in order to avoid the denominator in the above formula being 0, which makes the result a negative value, the preset value needs to be considered.
[0101] The inverse document frequency characterizes the prevalence of data tags between two data to be stored. The higher the prevalence of data tags, the stronger the correlation between the two data to be stored. The vector value can be obtained according to the vector value calculation formula, which is: vector value = frequency of occurrence * inverse document frequency, that is, the vector value is the distance between the data to be stored determined from the same data tag dimension. For the same data tags of the data to be stored two by two, the sub-data distance corresponding to the two data to be stored can be calculated by the Euclidean algorithm. The embodiment of the present application does not limit the specific calculation process. In an achievable manner, the average sub-data distance is calculated according to the sub-data distance corresponding to all the same data tags, and the average sub-data distance is determined as the data distance between the data to be stored two by two; it can be understood that the average value can more accurately reflect the degree of change between the data, so the average sub-data distance can be selected as the data distance between the data to be stored. The data distance and the target correspondence are matched one by one, and then the data similarity corresponding to the data distance can be obtained. In the embodiment of the present application, the smaller the data distance, the greater the data similarity; the target correspondence is pre-entered into the electronic device by the technician.
[0102] Based on the above embodiment, the frequency of occurrence of the same data tag in the target data tag is determined according to the number of first tags and the number of second tags; the data volume of the data to be stored is obtained, and the prevalence of the same data tag is determined according to the frequency of occurrence, the data to be stored and the preset similarity calculation formula; then the vector value corresponding to the same data tag is determined according to the frequency of occurrence and the inverse document frequency, and the data distance between the data to be stored is determined according to the vector value. When the correlation between the two data to be stored is higher, the data distance between the data to be stored is closer, so it is necessary to determine the data distance between the data to be stored; then the data similarity is determined according to the data distance and the target correspondence; the same data tag and the data to be stored are quantified in the form of a vector, and the data distance is determined according to the vector value, and the data similarity is reflected by the data distance, so as to effectively improve the accuracy of the data similarity.
[0103] The above-mentioned embodiment introduces a data storage method based on artificial intelligence learning from the perspective of method flow. The following embodiment introduces a data storage device based on artificial intelligence learning from the perspective of a virtual module or a virtual unit. Please refer to the following embodiment for details.
[0104] The present application embodiment provides a data storage device based on artificial intelligence learning, such as Figure 3 As shown, the data storage device based on artificial intelligence learning may specifically include:
[0105] An acquisition module 201 is used to acquire a plurality of data to be stored and call information of each data to be stored;
[0106] The storage area determination module 202 is used to determine the storage area corresponding to each data to be stored according to the call information;
[0107] The sub-storage area determination module 203 is used to obtain the data type and data attribute corresponding to each data to be stored, and divide each storage area into a plurality of sub-storage areas according to the data type;
[0108] The storage module 204 is used to store all the data to be stored based on the data attributes and all the sub-storage areas.
[0109] The higher the retrieval frequency of the data to be stored is, the higher the number of times the data to be stored is used, indicating that the data to be stored is more important, and therefore it is necessary to obtain the data to be stored and the corresponding retrieval frequency; then determine the storage area of each data to be stored according to the retrieval frequency, and store the data to be stored in different areas for different importance levels for subsequent data calls; obtain the data type and data attribute of each data to be stored, and divide the storage area into several sub-storage areas according to the data type, so as to further subdivide the data storage area on the basis of data storage of the same importance; then divide the storage area into several sub-storage areas according to the data type for classified storage; on the basis of classified storage, further refine the storage according to the data characteristics; compared with the related art, the present application divides the data to be stored into multiple levels according to the retrieval frequency, data type and data attribute of the data, and stores different levels in different areas to realize dynamic hierarchical storage of the data to be stored, which is convenient for subsequent users to perform hierarchical data query, so as to effectively improve the data query efficiency.
[0110] In a possible implementation of the embodiment of the present application, the call information includes multiple call time periods and the call frequencies corresponding to each call time period. When the storage area determination module 202 determines the storage area corresponding to each to-be-stored data according to the call information, it is specifically used to:
[0111] For each data to be stored, based on all retrieval periods and the retrieval frequencies corresponding to each retrieval period, determine the average retrieval frequency;
[0112] Get the retrieval frequency set corresponding to each storage area;
[0113] The storage area corresponding to each data to be stored is determined according to the average value of the retrieval frequency of each data to be stored and the retrieval frequency set corresponding to each storage area.
[0114] In a possible implementation of the embodiment of the present application, when the storage module 204 performs storage of all data to be stored based on data attributes and all sub-storage areas, it is specifically used to:
[0115] For each sub-storage area, obtaining the storage capacity corresponding to each sub-storage area;
[0116] Calculate the correlation between the data to be stored, and determine a number of associated data of each data to be stored according to the correlation;
[0117] Generate a data item to be stored according to all associated data of each data to be stored;
[0118] Obtain the data volume of all data items to be stored, and determine whether the data volume is less than the storable volume;
[0119] If yes, the data item to be stored is stored in the corresponding sub-storage area;
[0120] If not, the storage capacity of each sub-storage area is corrected according to the storage capacity and the data attribute to obtain a corrected sub-storage area;
[0121] The data item to be stored is stored in the modified sub-storage area.
[0122] In a possible implementation of the embodiment of the present application, when the storage module 204 calculates the correlation between the data to be stored, it is specifically used to:
[0123] Generate a number of data tags corresponding to each data to be stored, the data tags are used to identify the data to be stored;
[0124] Determining a first relevance of each piece of data to be stored based on a data label corresponding to each piece of data to be stored;
[0125] Acquire the retrieval time of the data to be stored, and determine the second relevance of each data to be stored based on the retrieval time of each data to be stored and a preset corresponding relationship;
[0126] A comprehensive correlation is determined based on the first correlation and the second correlation, and the comprehensive correlation is determined as the correlation between the data to be stored.
[0127] In a possible implementation of the embodiment of the present application, when the storage module 204 determines the first relevance of each data to be stored based on the data label corresponding to each data to be stored, it is specifically used to:
[0128] Obtain application information of each data to be stored, and determine the weight value of each data tag according to the application information and all data tags of each data to be stored;
[0129] Determine a number of target data labels based on the weight value of each data label and a preset weight value threshold;
[0130] For each pair of data to be stored, determine the same data labels from all target data labels;
[0131] Obtain the first label quantity of the target data label of each data to be stored and the second label quantity corresponding to the same data label;
[0132] Determine data similarity based on the number of first tags, the number of second tags, and all identical data tags;
[0133] The data similarity is determined as the first correlation of each data to be stored.
[0134] In a possible implementation of the embodiment of the present application, when the storage module 204 determines the data similarity according to the first tag number, the second tag number and all the same data tags, it is used to:
[0135] For each identical data tag, determining the occurrence frequency of the identical data tag based on the first tag quantity and the second tag quantity;
[0136] Obtain the amount of data to be stored, and determine the inverse document frequency of the same data label based on the frequency of occurrence, the amount of data to be stored and a preset similarity calculation formula, where the inverse document frequency represents the prevalence of the same data label;
[0137] Determine the vector value based on the occurrence frequency and the inverse document frequency;
[0138] Based on the vector value, determine the data distance between the data to be stored;
[0139] The data similarity is determined according to the data distance and the target correspondence, and the target correspondence is the correspondence between the data distance and the data similarity.
[0140] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the data storage device based on artificial intelligence learning described above can refer to the corresponding process in the aforementioned method embodiment and will not be repeated here.
[0141] An electronic device is provided in an embodiment of the present application, such as Figure 4 As shown, Figure 4 The electronic device shown includes: a processor 301 and a memory 303. The processor 301 and the memory 303 are connected, such as through a bus 302. Optionally, the electronic device may further include a transceiver 304. It should be noted that in actual applications, the transceiver 304 is not limited to one, and the structure of the electronic device does not constitute a limitation on the embodiments of the present application.
[0142] Processor 301 may be a CPU (Central Processing Unit), a general purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of this application. Processor 301 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0143] The bus 302 may include a path to transmit information between the above components. The bus 302 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus 302 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 Only one thick line is used in the diagram, but it does not mean that there is only one bus or only one type of bus.
[0144] The memory 303 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compressed optical disk, laser disk, optical disk, digital versatile disk, Blu-ray disk, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to these.
[0145] The memory 303 is used to store the application code for executing the solution of the present application, and the execution is controlled by the processor 301. The processor 301 is used to execute the application code stored in the memory 303 to implement the contents shown in the above method embodiment.
[0146] The electronic devices include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and fixed terminals such as digital TVs, desktop computers, etc. It can also be a server, etc. Figure 4 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0147] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer-readable storage medium is run on a computer, the computer can execute the corresponding content in the aforementioned method embodiment. Compared with the related art, in the embodiment of the present application, when the frequency of retrieval of the data to be stored is higher, the number of times the data to be stored is applied is higher, indicating that the data to be stored is more important, and therefore it is necessary to obtain the data to be stored and the corresponding retrieval frequency; then determine the storage area of each data to be stored according to the retrieval frequency, and store the data to be stored in different areas for different importance levels for subsequent data calls; obtain the data type and data attribute of each data to be stored, and divide the storage area into several sub-storage areas according to the data type, so as to further subdivide the data storage area on the basis of data storage of equal importance; then divide the storage area into several sub-storage areas according to the data type for classified storage; on the basis of classified storage, further refine the storage according to the data characteristics; compared with the related art, the present application divides the data to be stored into multiple levels according to the frequency of retrieval of the data, the data type and the data attribute, and stores different levels in different areas to realize dynamic hierarchical storage of the data to be stored, which is convenient for subsequent users to perform hierarchical data query, so as to effectively improve the data query efficiency.
[0148] It should be understood that, although the steps in the flowchart of the accompanying drawings are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.
[0149] The above description is only a partial implementation method of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A data storage method based on artificial intelligence learning, characterized in that: include: Acquire multiple data to be stored and call information of each of the data to be stored; Determine the storage area corresponding to each of the data to be stored according to the calling information; Acquire the data type and data attribute corresponding to each of the data to be stored, and divide each of the storage areas into a plurality of sub-storage areas according to the data type; Based on the data attributes and all the sub-storage areas, all the data to be stored are stored.
2. The data storage method based on artificial intelligence learning according to claim 1 is characterized in that: The calling information includes a plurality of calling time periods and the calling frequencies corresponding to the calling time periods. The determining, according to the call information, a storage area corresponding to each of the data to be stored comprises: For each of the to-be-stored data, based on all the retrieval time periods and the retrieval frequencies corresponding to the retrieval time periods, determining an average retrieval frequency; Obtaining a set of retrieval frequencies corresponding to each of the storage areas; The storage area corresponding to each of the data to be stored is determined according to the average of the retrieval frequencies of each of the data to be stored and the retrieval frequency sets corresponding to each of the storage areas.
3. The data storage method based on artificial intelligence learning according to claim 1 is characterized in that: The storing of all the data to be stored based on the data attributes and all the sub-storage areas includes: For each of the sub-storage areas, obtaining the storage capacity corresponding to each of the sub-storage areas; Calculating the correlation between the data to be stored, and determining a number of associated data of each of the data to be stored according to the correlation; Generate a data item to be stored according to all the associated data of each of the data to be stored; Obtaining the data amount of all data items to be stored, and determining whether the data amount is less than the storable amount; If yes, then storing the data item to be stored in the corresponding sub-storage area; If not, then modifying the storage capacity of each of the sub-storage areas according to the storage capacity and the data attribute to obtain a modified sub-storage area; The data item to be stored is stored in the modified sub-storage area.
4. The data storage method based on artificial intelligence learning according to claim 3 is characterized in that: The calculating the correlation between the data to be stored includes: Generating a plurality of data tags corresponding to each of the data to be stored, wherein the data tags are used to identify the data to be stored; Determining a first relevance of each of the data to be stored based on a data label corresponding to each of the data to be stored; Acquiring the retrieval time of the data to be stored, and determining the second relevance of each data to be stored based on the retrieval time of each data to be stored and a preset corresponding relationship; A comprehensive correlation is determined based on the first correlation and the second correlation, and the comprehensive correlation is determined as the correlation between the data to be stored.
5. The data storage method based on artificial intelligence learning according to claim 4 is characterized in that: The determining, based on the data labels corresponding to the data to be stored, the first relevance of each of the data to be stored, comprises: Acquire application information of each of the data to be stored, and determine a weight value of each of the data tags according to the application information and all of the data tags of each of the data to be stored; Determining a number of target data tags based on the weight values of the data tags and a preset weight value threshold; For each pair of the data to be stored, determining the same data tags from all the target data tags; Acquire the first label quantity of the target data label of each of the data to be stored and the second label quantity corresponding to the same data label; Determine data similarity according to the first number of tags, the second number of tags, and all the same data tags; The data similarity is determined as a first correlation of each of the data to be stored.
6. The data storage method based on artificial intelligence learning according to claim 5 is characterized in that: The determining of data similarity according to the first tag quantity, the second tag quantity and all the same data tags includes: For each of the same data tags, based on the first tag quantity and the second tag quantity, determining the occurrence frequency of the same data tag; Acquire the data volume of the data to be stored, and determine the inverse document frequency of the same data tag according to the occurrence frequency, the data volume of the data to be stored and a preset similarity calculation formula, wherein the inverse document frequency represents the prevalence of the same data tag; Determining a vector value based on the occurrence frequency and the inverse document frequency; Based on the vector value, determining the data distance between the data to be stored; The data similarity is determined according to the data distance and the target corresponding relationship, wherein the target corresponding relationship is the corresponding relationship between the data distance and the data similarity.
Citation Information
Patent Citations
Data fragmentation method and device based on artificial intelligence, equipment and medium
CN114860722A
Working condition data management method, device, equipment and medium
CN116842223A
Data storage method of computer database
CN118193643A
Document editing method and device, server, terminal, and storage medium
WO2022095520A1