Data storage method and apparatus
By acquiring and associating target hotspot information with current hotspot data and user historical query data, the problems of storage media lifespan and user experience are solved, achieving efficient data access and extending storage media lifespan.
Patent Information
- Application Number
- CN202111509697.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-07-13
- Filing Date
- 2021-12-10
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2041-12-10
AI Technical Summary
Existing technologies for storing knowledge graph data result in a poor user experience, reduce user stickiness, and affect the life of the storage medium.
By acquiring current hot data and user historical query data, target hot information is identified, and its features are extracted for one-to-one association and storage. The data content in the storage medium is dynamically adjusted to improve access efficiency and reduce the frequency of data transfer between different storage media.
It improves data access efficiency, extends the lifespan of storage media, and enhances the user experience.
Smart Images

Figure CN114186099B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence of computer technology, in particular to a data storage method. The present application also relates to a data storage device, a computing device, and a computer readable storage medium. BACKGROUND
[0002] Artificial intelligence (AI) refers to the ability of an engineered (i.e., designed and manufactured) system to perceive the environment, and the ability to acquire, process, apply and represent knowledge. The development status of key technologies in the field of artificial intelligence includes machine learning, knowledge graph, natural language processing, computer vision, human-computer interaction, biometric identification, virtual reality / augmented reality, and other key technologies. Knowledge graph describes concepts, entities and their relationships in a structured form, expresses the information on the Internet in a form closer to human cognition of the world, and provides the ability to better organize, manage and understand the vast amount of information on the Internet. With the development of computer technology, various access technologies have emerged. For large-scale knowledge graph access technology, a graph database or a more stable distributed big data platform is usually used to store large-scale knowledge graph. That is, in large-scale data storage, a small part of data is stored in the memory for efficient reading, and most of the data is stored in the hard disk. However, during user retrieval and access, high-frequency data reading will be accompanied, which greatly affects the data access efficiency and the service life of the storage medium.
[0003] In the prior art, the size of the cache and the memory is reasonably utilized, and a garbage collector is used to move high-frequency access data in and out during the data active time, or a memory management system is designed to complete data management, keep high-frequency access data in the cache, and realize efficient data access. However, for the data stored in the knowledge graph, the above method will result in poor user experience, thereby reducing user stickiness. Therefore, an effective solution is needed to solve the above problems. SUMMARY
[0004] Therefore, the embodiments of the present application provide a data storage method to solve the technical defects in the prior art. The embodiments of the present application also provide a data storage device, a computing device, and a computer readable storage medium.
[0005] According to a first aspect of the embodiments of the present application, a data storage method is provided, comprising:
[0006] obtaining current hot data and user historical query data;
[0007] determine at least one target hot information according to the current hot information and the user historical query data;
[0008] extract target features of each target hot information, and store the target hot information and the target features of the target hot information in a one-to-one corresponding storage manner.
[0009] According to a second aspect of the embodiment of the present application, a data storage device is provided, comprising:
[0010] a obtaining module configured to obtain current hot information and user historical query data;
[0011] a determining module configured to determine at least one target hot information according to the current hot information and the user historical query data;
[0012] a storing module configured to extract target features of each target hot information, and store the target hot information and the target features of the target hot information in a one-to-one corresponding storage manner.
[0013] According to a third aspect of the embodiment of the present application, a computing device is provided, comprising:
[0014] a memory and a processor;
[0015] the memory is configured to store computer executable instructions, and the processor is configured to execute the computer executable instructions to implement the steps of the data storage method.
[0016] According to a fourth aspect of the embodiment of the present application, a computer readable storage medium is provided, which stores computer executable instructions, and the instructions are executed by a processor to implement the steps of the data storage method.
[0017] According to a fifth aspect of the embodiment of the present application, a chip is provided, which stores computer instructions, and the computer instructions are executed by the chip to implement the steps of the data storage method.
[0018] The data storage method provided by the present application can store high-frequency access information, dynamically adjust the data content in the storage medium, improve the access efficiency, reduce the frequency of data transfer in different storage media, solve the data caching from the data contact level, and improve the service life of the storage medium. BRIEF DESCRIPTION OF DRAWINGS
[0019] Figure 1 is a flowchart of a data storage method according to an embodiment of the present application;
[0020] Figure 2 is a flowchart of a data storage method according to an embodiment of the present application;
[0021] Figure 3 is a flowchart of a data storage method according to an embodiment of the present application;
[0022] Figure 4 is a structural diagram of a data storage device according to an embodiment of the present application;
[0023] Figure 5 is a structural diagram of a computing device according to an embodiment of the present application. DETAILED DESCRIPTION
[0024] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be practiced without the specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to obscure the present application. Unless otherwise noted, the description of a particular embodiment of the application herein specifically includes all combinations of elements including that particular embodiment and all combinations of elements including any other
[0025] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of one or more embodiments of the present application. As used in this disclosure and the appended claims, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0026] It will be understood that, although the terms first, second, etc. can be used herein to describe various information, these terms are not intended to denote a temporal or chronological order. Rather, these terms are used solely to distinguish one from another only. For example, a first item could be termed a second item, and, similarly, a second item could be termed a first item without departing from the scope of one or more embodiments of the present application.
[0027] First, the nomenclature related to one or more embodiments of the present application is explained.
[0028] Graph Database: Graph Database is a non-relational database that applies graph theory to store entity and relationship information between entities. Common graph databases include Neo4j (a network-oriented graph database), OrientDB (a flexible document-graph database that combines the flexibility of a document database and the link management capabilities of a graph database), titan (a client library that depends on a storage engine), and the like. Graph database is an important carrier of knowledge graph data.
[0029] Graph Embedding: is a process of mapping graph data into low-dimensional dense vectors; is to convert an attributed graph into a vector or a set of vectors. The embedding should capture the topology of the graph, the relationships between vertices, and other relevant information about the graph, subgraphs, and vertices. Graph embedding represents an entire graph with a single vector. Such embedding is used to make predictions at the level of the graph or to compare, visualize the entire graph.
[0030] Network media dynamic monitoring: is to monitor the information such as information reports, forwarding, comments and the like of media, is an important decision reference basis for enterprises and governments at all levels to master media dynamics, understand the latest public opinion information hotspots, thereby deeply understand the network reputation of themselves, maintain and improve the image of themselves, and prevent the loss of interests.
[0031] In the present application, a data storage method is provided. The present application also relates to a data storage device, a computing device, and a computer readable storage medium, which are described in detail in the following embodiments.
[0032] Figure 1 A flowchart of a data storage method according to an embodiment of the present application is shown, which specifically includes the following steps:
[0033] Step 102: obtaining current hot data and user historical query data.
[0034] Specifically, the current hot spot refers to a thing or matter that currently attracts widespread attention and discussion in the media and is often controversial. The current hot spot data can be at least one of text, pictures, sounds, videos, etc. representing a current hot topic, focus, or other field with high current discussion. The current hot spot data is obtained in the form of a set, for example, the text of a flash sale, price reduction promotion, popular movie ticket, star releasing a new album, and related descriptions corresponding to the text are collected one by one as current hot spot data. User historical query data refers to data formed by user queries and browsing through multimedia carriers, which can be at least one of text, pictures, sounds, videos, etc. User historical query data is also obtained in the form of a set, for example, if user U1 queries "2021 college entrance examination" through a webpage, the content related to "2021 college entrance examination" browsed by user U1 on the webpage is user historical query data.
[0035] In practical applications, when obtaining current hot spot data, network media dynamic monitoring technology can be used to obtain real-time network hot spot data, i.e. current hot spot data. For example, through network crawler technology, according to certain rules, the current hot spot data on the network can be automatically captured. The current hot spot data can also be obtained by searching for "hot spot", "hot search", "hot discussion" and the like on a search website. In addition, user query behavior needs to be recorded to further obtain user historical query data. For example, through network crawler technology, real-time network hot spot data, i.e. current hot spot data, can be obtained. User historical query data can be obtained according to recorded user query behavior.
[0036] It should be noted that in the present application, current hot spot data can be obtained first, and then user historical query data can be obtained. Alternatively, user historical query data can be obtained first, and then current hot spot data can be obtained. Alternatively, current hot spot data and user historical query data can be obtained simultaneously. The present application does not limit the order in which current hot spot data and user historical query data are obtained.
[0037] The present application obtains current hot spot data and user historical query data to prepare data for determining target hot spot information, ensuring the accuracy and comprehensiveness of the target hot spot information, further improving the efficiency of data storage, and to some extent, improving the service life of the storage medium.
[0038] Step 104: determining at least one target hot spot information according to the current hot spot data and the user historical query data.
[0039] On the basis of obtaining current hot spot data and user historical query data, further, target hot spot information is determined according to the obtained current hot spot data and user historical query data.
[0040] Specifically, the target hot information can be current hot information, such as "today's hot news", or predicted hot information according to user historical query data, such as "lunar eclipse" which is searched by most users on the Internet.
[0041] In actual application, in the case of obtaining current hot data, the associated data within N-degree relationship of the current hot data in the knowledge graph corresponding to the current hot data can be determined according to the current hot data, wherein N is a positive integer, and the size of N can be set according to requirements; the N-degree relationship refers to the relationship between two nodes indirectly connected through (N-1) nodes in the knowledge graph, for example, node A and node B are connected through node C, the relationship between node A and node B is two-degree relationship, and the relationship between node A and node B and node C is one-degree relationship, that is, the data corresponding to node A is the two-degree relationship associated data of the data corresponding to node B, and the data corresponding to node A and the data corresponding to node B are respectively two-degree relationship associated data of the data corresponding to node C; the current hot data and each associated data of the current hot data correspond to a node in the knowledge graph. At the same time, a user query behavior structure graph can be constructed according to user historical query data, wherein the user query behavior structure graph refers to connecting each user historical query data as a node according to the order of being queried to obtain a graph representing the connection relationship of each data. Then, the embedding representation of each user historical query data is calculated according to the connection relationship of each node in the user query behavior structure graph, and on this basis, the embedding representation of each user historical query data is input into a pre-set prediction model for prediction to obtain predicted hot data. Further, the information corresponding to the current hot data, the associated data within N-degree relationship of the current hot data, and the predicted hot data is determined respectively, and at least one target hot information is determined according to these information.
[0042] For example, in the case of N being 2, the current hot data is D1, and the second-degree relationship of the current hot data D1 needs to be determined according to the current hot data D1 to determine the associated data of the current hot data D1 within the second-degree relationship, that is, data D12 and data D13; meanwhile, the user query behavior structure graph is constructed according to the user historical query data D2, D3 and D4, and then the embedding representation of D2, D3 and D4 is determined respectively, so that the predicted hot data is D3' and D4', wherein the predicted hot data D3' can be the same as or different from the user historical query data D3. For example, the user historical query data D3 is an article about Mars, and the predicted hot data D3' can be another article about Mars or all articles about Mars, or the article about Mars in the user historical query data D3. Similarly, the predicted hot data D4' can be the same as or different from the user historical query data D4. Further, the information corresponding to the current hot data D1, the data D12, the data D13, the predicted hot data D3' and D4' is determined respectively, and the target hot information is further determined according to the determined information.
[0043] It should be noted that, in order to further ensure the integrity of the target hot information and avoid the repetition of the target hot information, the target hot information set can be determined first, and then the target hot information can be determined from the target hot information set. In an optional embodiment of the present embodiment, the specific implementation process of determining at least one target hot information according to the current hot data and the user historical query data can be as follows:
[0044] Determining a target hot information set according to the current hot data and the user historical query data;
[0045] Determining the information in the target hot information set as the target hot information to obtain at least one target hot information.
[0046] Specifically, the target hot information set refers to a set containing one or more target hot information.
[0047] In actual application, the associated data of the current hot data can be matched according to the current hot data; then the user query behavior structure graph is constructed according to the user historical query data, and the embedding representation of the user historical query data is further determined according to the user query behavior structure graph, the hot data is predicted, and the predicted hot data is obtained. On this basis, the information corresponding to the current hot data, the information corresponding to the associated data of the current hot data and the information corresponding to the predicted hot data are combined to generate a target hot information set. Then the same information in the target hot information set is removed, and at this time the target hot information set contains at least one target hot information, that is, each information in the target hot information set is the target hot information.
[0048] According to the current hot data D1, the associated data of the current hot data D1 includes data D12 and data D13; the user query behavior structure graph is constructed according to the user historical query data D2, D3 and D4, and then the embedding representation of D2, D3 and D4 is determined respectively, and the predicted hot data is D3' and D4'. Further, the information corresponding to the current hot data D1, the information corresponding to the data D12, the information corresponding to the data D13, the information corresponding to the predicted hot data D3' and the information corresponding to the predicted hot data D4' are merged, so as to obtain the target hot information set. If the information corresponding to the current hot data D1, the information corresponding to the data D12, the information corresponding to the data D13, the information corresponding to the predicted hot data D3' and the information corresponding to the predicted hot data D4' are different from each other, the information corresponding to the current hot data D1, the information corresponding to the data D12, the information corresponding to the data D13, the information corresponding to the predicted hot data D3' and the information corresponding to the predicted hot data D4' are all target hot information. If the information corresponding to the current hot data D1 is the same as the predicted hot data D3', in order to avoid the repetition of information, the information corresponding to the current hot data D1 (or the information corresponding to the predicted hot data D3') in the target hot information set can be deleted, at this time, the information corresponding to the data D12, the information corresponding to the data D13, the information corresponding to the predicted hot data D3' (or the information corresponding to the current hot data D1) and the information corresponding to the predicted hot data D4' remaining in the target hot information set are all target hot information.
[0049] For example, the information corresponding to the current hot data includes: the postponement of the exam, the last snow in 2021; the information corresponding to the predicted hot data includes: Xiaoming winning the championship, the postponement of the exam; and the target hot information obtained by merging and removing the duplicate information is: the postponement of the exam, the last snow in 2021 and Xiaoming winning the championship.
[0050] In an optional implementation of the embodiment, the information corresponding to the current hot data and the information corresponding to the associated data of the current hot data are both current hot information, the information corresponding to the predicted hot data predicted according to the user historical query data is predicted hot information, and the target hot information set is obtained by merging the current hot information and the predicted hot information. That is, the target hot information set includes the current hot information and the predicted hot information, at this time, the target hot information set is determined according to the current hot data and the user historical query data, and the specific implementation process can be as follows:
[0051] According to the current hot data, the current hot information is determined in the knowledge graph;
[0052] According to the user historical query data, the predicted hot information is determined;
[0053] The current hot information and the predicted hot information are combined to generate a target hot information set.
[0054] Specifically, the knowledge graph (Knowledge Graph) is also called knowledge domain visualization or knowledge field mapping map in the library and information field, which is a series of various different graphs showing the development process and structural relationship of knowledge. The knowledge graph uses visualization technology to describe knowledge resources and their carriers, and mines, analyzes, constructs, draws, and displays knowledge and their mutual relationships. The knowledge graph can combine theories and methods of applied mathematics, graphics, information visualization technology, information science, and other disciplines with methods such as citation analysis and co-occurrence analysis, and use visualized graphs to display the core structure, development history, frontier field, and overall knowledge architecture of a discipline to achieve the purpose of multi-disciplinary integration. The current hot information refers to information determined according to current hot data and / or associated data of the current hot data. The predicted hot information is information determined according to user historical query data.
[0055] In practical applications, the obtained current hot data can be used as a current node in the knowledge graph, and the first-degree relationship nodes and the second-degree relationship nodes of the current node are obtained. Further, the information corresponding to the current node, the information corresponding to the first-degree relationship nodes, and the information corresponding to the second-degree relationship nodes are determined. These determined information is the current hot information, that is, the current hot information is determined in the knowledge graph according to the current hot data. In addition, user historical query data is further obtained by recording user query behavior, and a user behavior structure graph is constructed based on the user historical query data. The embedding representation of each user historical query data is determined, the hot data is predicted, and the information corresponding to the predicted hot data, that is, the predicted hot information, is obtained. The determined current hot information and the predicted hot information are combined, that is, the union of the current hot information and the predicted hot information is taken to obtain the target hot information set.
[0056] For example, according to the current hot data, and in combination with the knowledge graph, the current hot information {N1, N2, N3, N4} is determined. According to the user historical query data, the predicted hot information is {N3, N5}. The union of {N1, N2, N3, N4} and {N3, N5} is taken to obtain the target hot information set {N1, N2, N3, N4, N5}.
[0057] After the target hotspot information is determined, since different target hotspot information contents are different, the importance is also different, and the weight of each target hotspot information is also different. In order to further reflect the importance of different target hotspot information, the target hotspot information can be sorted according to the weight of each target hotspot information. That is, after at least one target hotspot information is determined, the weight of each target hotspot information needs to be determined; according to the weight of each target hotspot information, each target hotspot information is sorted. In this way, it is beneficial to store according to the weight of each target hotspot information subsequently.
[0058] The target hotspot information is determined according to the current hotspot information and the predicted hotspot information, so it can be seen that the weight of each target hotspot information is inevitably affected by the weight of the current hotspot information and the weight of the predicted hotspot information, therefore, determining the weight of each target hotspot information can be realized by the following process:
[0059] Obtaining the first weight and the second weight of each target hotspot information, the first weight is the weight of each target hotspot information in the current hotspot information, and the second weight is the weight of each target hotspot information in the predicted hotspot information;
[0060] According to the first weight and the second weight of the first target hotspot information, the weight of the first target hotspot information is determined, and the first target hotspot information is any one target hotspot information.
[0061] Specifically, the first weight refers to the weight of the target hotspot information relative to the current hotspot information, for example, the target hotspot information X1 corresponds to the hotspot information Y1 in the current hotspot information, and the weight of the hotspot information Y1 is 0.1, then the first weight of the target hotspot information X1 is 0.1, and for example, the target hotspot information X2 does not correspond to each hotspot information in the current hotspot information, then the first weight of the target hotspot information X2 is 0; the second weight refers to the weight of the target hotspot information relative to the predicted hotspot information, for example, the target hotspot information X3 corresponds to the hotspot information Y2 in the predicted hotspot information, and the weight of the hotspot information Y2 is 0.7, then the second weight of the target hotspot information X3 is 0.7, and for example, the target hotspot information X4 does not correspond to each hotspot information in the predicted hotspot information, then the second weight of the target hotspot information X4 is 0.
[0062] In actual application, in order to make the weight of each target hot information more accurate, and the target hot information is derived from the set of current hot information and predicted hot information, therefore, the weight of each target hot information can be determined from two aspects of current hot information and predicted hot information respectively. Firstly, the weight of each target hot information in current hot information is determined, that is, the weight of the corresponding current hot information of each target hot information is determined, thereby obtaining the first weight of each target hot information; meanwhile, the weight of each target hot information in predicted hot information is determined, that is, the weight of the corresponding predicted hot information of each target hot information is determined, thereby obtaining the second weight of each target hot information. In the present application, the first weight of each target hot information can be obtained firstly, and then the second weight of each target hot information is obtained; or the second weight of each target hot information can be obtained firstly, and then the first weight of each target hot information is obtained; or the first weight and the second weight of each target hot information can be obtained simultaneously, which is not limited in the present application. On the basis of obtaining the first weight and the second weight of each target hot information, further, the weight of each target hot information is determined according to the first weight and the second weight of each target hot information respectively.
[0063] For example, the target hot information has N1, N2 and N3, the current hot information and the corresponding weight, the predicted hot information and the corresponding weight are shown in Table 1. Among them: the target hot information N1 corresponds to the current hot information N1, the weight of the current hot information N1 is 0.6, then the first weight of the target hot information N1 is 0.6; the target hot information N2 corresponds to the current hot information N2, the weight of the current hot information N2 is 0.4, then the first weight of the target hot information N2 is 0.4; the target hot information N3 does not correspond to the current hot information N1 and N2, then the first weight of the target hot information N3 is 0; the target hot information N1 corresponds to the predicted hot information N1, the weight of the predicted hot information N1 is 0.55, then the second weight of the target hot information N1 is 0.55; the target hot information N2 does not correspond to the predicted hot information N1 and N3, then the second weight of the target hot information N1 is 0; the target hot information N3 corresponds to the predicted hot information N3, the weight of the predicted hot information N3 is 0.45, then the second weight of the target hot information N3 is 0.45. Further, the weight of the target hot information N1 is determined according to the first weight 0.6 and the second weight 0.55 of the target hot information N1; the weight of the target hot information N2 is determined according to the first weight 0.4 and the second weight 0 of the target hot information N2; the weight of the target hot information N3 is determined according to the first weight 0 and the second weight 0.45 of the target hot information N3.
[0064] Table 1 weight of current hot information and predicted hot information
[0065]
[0066] In addition, when the same hotspot information is contained in the current hotspot information and the predicted hotspot information, since the hotspot information is determined according to the current hotspot data in the current hotspot information, and the hotspot information is determined according to the user historical query data in the predicted hotspot information, the determination manners of the hotspot information are different, and thus the weight of the hotspot information in the current hotspot information and the weight of the hotspot information in the predicted hotspot information can be different.
[0067] It should be noted that, since each current hotspot information and each predicted hotspot information have certain differences in content and importance, the weight of each current hotspot information and each predicted hotspot information can be set in advance. In the current hotspot information, the weight of the current hotspot information corresponding to the current hotspot data is the highest, the weight of the current hotspot information corresponding to the first-degree associated data of the current hotspot data is the second highest, the weight of the current hotspot information corresponding to the second-degree associated data of the current hotspot data is the third highest, and so on. In the predicted hotspot information, the weight of the predicted hotspot information is determined according to the search and browse times of the predicted hotspot information by the user, and the higher the search and browse times, the higher the weight of the corresponding predicted hotspot information; the lower the search and browse times, the lower the weight of the corresponding predicted hotspot information.
[0068] For example, the current hotspot information contains first current hotspot information corresponding to the current hotspot data, second current hotspot information corresponding to the first-degree associated data of the current hotspot data, and third current hotspot information corresponding to the second-degree associated data of the current hotspot data, and thus the weight of the first current hotspot information is greater than the weight of the second current hotspot information, and the weight of the second current hotspot information is greater than the weight of the third current hotspot information. The weight of the first current hotspot information can be set to 0.5, the weight of the second current hotspot information can be set to 0.3, and the weight of the third current hotspot information can be set to 0.2. The predicted hotspot information contains first predicted hotspot information and second predicted hotspot information, the search and browse times of the first predicted hotspot information are 600, and the search and browse times of the second predicted hotspot information are 400, and thus the weight of the first predicted hotspot information is set to 0.6, and the weight of the second predicted hotspot information is set to 0.4.
[0069] In one or more embodiments of the present embodiment, in order to improve the efficiency of determining the weight of each target hotspot information, and to reflect that the influence degrees of the current hotspot information and the predicted hotspot information on the target hotspot information are different, the specific implementation process of determining the weight of the first target hotspot information according to the first weight and the second weight of the first target hotspot information can be as follows:
[0070] The weighted sum of the first weight and the second weight of the first target hotspot information is calculated, and the weighted sum is determined as the weight of the first target hotspot information.
[0071] Specifically, the weighted sum refers to assigning a weight to the first weight and the second weight respectively, multiplying the first weight by the weight corresponding to the first weight, multiplying the second weight by the weight corresponding to the second weight, and then adding the two products to obtain the sum.
[0072] In actual application, when calculating the weight of each target hotspot information, the product of the first weight of the target hotspot information and the weight of the first weight can be calculated first to obtain a first product; the product of the second weight of the target hotspot information and the weight of the second weight is calculated first to obtain a second product, and then the first product and the second product are added, and the sum obtained is the weight of the target hotspot information. The specific calculation process is shown in formula 1.
[0073] y = a * x1 + b * x2 (formula 1)
[0074] Wherein, y represents the weight of the target hotspot information, a represents the weight of the first weight, x1 represents the first weight of the target hotspot information, b represents the weight of the second weight, and x2 represents the second weight of the target hotspot information.
[0075] For example, the weight a of the first weight is 0.6, the weight b of the second weight is 0.4, and the target hotspot information has N1, N2 and N3. In the case that the first weight of the target hotspot information N1 is 0.6 and the second weight is 0.55, the weight of the target hotspot information N1 is 0.6*0.6+0.55*0.4 = 0.58. In the case that the first weight of the target hotspot information N2 is 0.4 and the second weight is 0, the weight of the target hotspot information N2 is 0.6*0.4+0.4*0 = 0.24; in the case that the first weight of the target hotspot information N3 is 0 and the second weight is 0.45, the weight of the target hotspot information N3 is 0.6*0+0.4*0.45 = 0.18.
[0076] It should be noted that due to different determination methods of the first weight and the second weight, in order to improve the accuracy of the weight of the target hotspot information, the first weight and the second weight can be normalized first, and then the weighted sum of the first weight and the second weight of the first target hotspot information after normalization is calculated, and the weighted sum is determined as the weight of the first target hotspot information.
[0077] In this application, after obtaining the current hotspot data and the user historical query data, at least one target hotspot information is further determined according to the current hotspot data and the user historical query data, which lays a foundation for subsequent storage of target hotspot data and improves the speed of data storage.
[0078] Step 106: extracting target features of each target hotspot information, and storing each target hotspot information and the target features of each target hotspot information in a one-to-one corresponding manner.
[0079] On the basis of determining the at least one target hotspot information according to the current hotspot data and the user historical query data, further, target features of each target hotspot information need to be extracted, and then each target hotspot information and the target feature corresponding thereto are in one-to-one correspondence and are stored in association.
[0080] Specifically, the target feature refers to a feature of the target hotspot information; the one-to-one correspondence refers to a one-to-one relationship between the target hotspot information and the target feature, for example, the target hotspot information M1 corresponds to the target feature m1, the target hotspot information M2 corresponds to the target feature m2, and the target hotspot information M3 corresponds to the target feature m3; the associated storage refers to storing the target hotspot information and the target feature of the target hotspot information in a structure pair, for example, the target hotspot information is M1, the target feature of the target hotspot information M1 is m1, and the structure of "m1-M1" or "M1-m1" is stored.
[0081] In actual application, before storing the target hotspot information, the target feature of each target hotspot information needs to be extracted, and there are many ways to extract the target feature, for example, a word library representation method, which is not limited in the present application. Further, each target hotspot information and the target feature of each target hotspot information are stored in association in a one-to-one manner. The place for associated storage can be a memory or a cache with a higher data access frequency. Since the target hotspot information belongs to high-frequency access data, storing it in the memory or the cache is beneficial to improve the speed of reading the target hotspot information by the user, thereby improving the user stickiness.
[0082] In an optional embodiment of the present embodiment, after determining the weight of each target hotspot information, each target hotspot information needs to be sorted according to the weight of each target hotspot information. On this basis, each target hotspot information and the target feature of each target hotspot information are stored in association, which can be, according to the sorting result, each target hotspot information and the target feature of each target hotspot information are stored in association in sequence. The higher the weight of the target hotspot information, the more likely each target hotspot information is accessed at a high frequency, and vice versa. Therefore, storing the target hotspot information with high weight in front of the memory or the cache can make it more convenient for the user to read, that is, according to the sorting result, the target hotspot information with high weight and the target feature of the target hotspot information are stored in association, which is more effective.
[0083] In addition, long-term unused data is also stored in the memory or cache, which is low-frequency access data, i.e. historical hot information. Since the historical hot information is not accessed most of the time, storing it in the memory or cache will occupy a large amount of storage space, thereby causing the memory or cache data reading to be slow. To avoid these phenomena, the stored historical hot information can be compressed and the original historical hot information can be overwritten. The specific implementation process can be as follows:
[0084] determining the stored historical hot information;
[0085] compressing the stored historical hot information to obtain compressed historical hot information;
[0086] replacing the stored historical hot information with the compressed historical hot information.
[0087] Specifically, compression is a process of reducing the size of historical hot information by a specific algorithm, such as compressing a certain file. The original file size is 8G, and the compressed file size is only 1G.
[0088] In practical applications, it is necessary to determine which historical hot information has been stored, and then compress the stored historical hot information to make the historical hot information smaller, replace the stored historical hot information with the compressed historical hot information, and reduce the occupied storage space. For example, the stored historical hot information is H, the stored historical hot information H is compressed to obtain historical hot information h, and then the historical hot information h is replaced with the historical hot information H. At this time, the memory or cache contains historical hot information h and does not contain historical hot information H.
[0089] Since the compressed historical hot information is used to replace the stored historical hot information, if the user reads the historical hot data, the problem of not finding in the memory or cache may occur. Based on this, the stored historical hot information can be compressed, the identification information of the stored historical hot information can be extracted, and the identification information and the compressed historical hot information can be stored in association, so as to facilitate the user to find the compressed historical hot information according to the identification information. The specific implementation process can be as follows:
[0090] extracting identification information of the historical hot information, the identification information including an abstract and / or a label;
[0091] compressing the stored historical hot information to obtain compressed historical hot information;
[0092] storing the identification information of the historical hot information in association with the compressed historical hot information, and replacing the stored historical hot information.
[0093] In actual application, the identification information (abstract and / or label) of the historical hot information needs to be extracted before the stored historical hot information is compressed, and then the stored historical hot information is compressed. On this basis, the compressed historical hot information and the corresponding identification information (abstract and / or label) are stored in association, and the stored historical hot information is replaced. In this way, when the user accesses the historical hot information, the compressed historical hot information is triggered and decompressed in the form of label or abstract information.
[0094] To further illustrate the data storage method provided in the present application, see Figure 2 . Figure 2 A flowchart of data storage provided by an embodiment of the present application is shown. First, the current hot information is determined: the current hot data is obtained from the network by using the crawler technology, and the current hot information is determined in the knowledge graph according to the current hot data, wherein the current hot information includes D1, F, and C; then, the predicted hot information is determined: the user historical query data is obtained, including two user historical query data A→B→C→D2 and A→C→E→F, and the user query behavior structure graph is constructed according to the two user historical query data (for details, see Figure 2 ), the graph embedding of A, B, C, D2, E, and F is determined based on the user query behavior graph structure, and the graph embedding of A, B, C, D2, E, and F is shown in Figure 2 , according to the prediction of the hot information, the predicted hot information D2 is obtained, wherein the current hot information D1 is the same as the predicted hot information D2; then, the current hot information D1, F, and C are merged with the predicted hot information D2, and three target hot data D3, F, and C are determined, wherein the target hot data D3 can be any one of the current hot information D1 and the predicted hot information D2. Further, the weight is determined and sorted, that is, the weight of the target hot data D3, F, and C is determined and sorted according to the weight, according to the weight sorting result, the target hot information and the target feature of each target hot information are stored in association; after the new target hot information is determined, the previously determined target hot information is determined as the historical hot information, that is, the stored historical hot information is determined, the identification information of the stored historical data is extracted, wherein the identification information includes label and abstract, the stored historical hot information is compressed to obtain the compressed historical hot information, the compressed historical hot information and the identification information are stored in association and replace the stored historical hot information.
[0095] The data storage method provided in the application comprises the following steps: obtaining current hot data and user historical query data; determining at least one target hot information according to the current hot data and the user historical query data; extracting target features of each target hot information, and storing each target hot information and the target features of each target hot information in an associated manner.
[0096] Figure 3 A flowchart diagram of another data storage method provided by an embodiment of the application is shown, and the method comprises the following steps:
[0097] In step 302, current hot data and user historical query data are obtained.
[0098] In step 304, current hot information is determined in a knowledge graph according to the current hot data.
[0099] In step 306, predicted hot information is determined according to the user historical query data.
[0100] It should be noted that step 304 and step 306 can be performed simultaneously, or step 304 can be performed first and then step 306 can be performed, or step 306 can be performed first and then step 304 can be performed, and the application does not limit this. In this embodiment, step 304 and step 306 are taken as an example and are described simultaneously.
[0101] In step 308, the current hot information and the predicted hot information are merged to generate a target hot information set.
[0102] In step 310, information in the target hot information set is determined as target hot information, and at least one target hot information is obtained.
[0103] In step 312, a first weight and a second weight of each target hot information are obtained.
[0104] The first weight is the weight of each target hot information in the current hot information, and the second weight is the weight of each target hot information in the predicted hot information.
[0105] In step 314, a weighted sum of the first weight and the second weight of the first target hot information is calculated, and the weighted sum is determined as the weight of the first target hot information, and the first target hot information is any one target hot information.
[0106] In step 316, each target hot information is sorted according to the weight of each target hot information.
[0107] Step 318: according to the sorting result, the target hotspot information and the target feature of each target hotspot information are stored in association.
[0108] Step 320: determine the stored historical hotspot information.
[0109] Step 322: extract the identification information of the historical hotspot information, and the identification information includes an abstract and / or a label.
[0110] Step 324: compress the stored historical hotspot information to obtain compressed historical hotspot information.
[0111] Step 326: store the identification information of the historical hotspot information in association with the compressed historical hotspot information, and replace the stored historical hotspot information.
[0112] The data storage method provided by the application, by obtaining current hotspot data and user historical query data; according to the current hotspot data and the user historical query data, at least one target hotspot information is determined; the target feature of each target hotspot information is extracted, and each target hotspot information and the target feature of each target hotspot information are stored in association in a one-to-one correspondence storage mode. In this way, high-frequency access information can be stored, the data content in the storage medium can be dynamically adjusted, the access efficiency can be improved, and the frequency of data transfer between different storage media can be reduced, so that data caching is solved from the data contact level, and the service life of the storage medium is improved.
[0113] Corresponding to the above method embodiment, the application also provides a data storage device embodiment, Figure 4 The structure of a data storage device provided by an embodiment of the application is shown. As shown in the figure, Figure 4 The device comprises:
[0114] The acquisition module 402 is configured to acquire current hotspot data and user historical query data;
[0115] The determination module 404 is configured to determine at least one target hotspot information according to the current hotspot data and the user historical query data;
[0116] The storage module 406 is configured to extract the target feature of each target hotspot information, and store each target hotspot information and the target feature of each target hotspot information in association in a one-to-one correspondence storage mode.
[0117] In one or more embodiments of the present embodiment, the determination module 404 is further configured to:
[0118] determine the weight of each target hotspot information;
[0119] According to the weight of each target hotspot information, the target hotspot information is sorted.
[0120] Further, the storage module 406 is further configured to:
[0121] According to the sorting result, the target hotspot information and the target feature of each target hotspot information are stored in association.
[0122] In one or more embodiments of the present embodiment, the determination module 404 is further configured to:
[0123] Obtain the first weight and the second weight of each target hotspot information, the first weight being the weight of each target hotspot information in the current hotspot information, and the second weight being the weight of each target hotspot information in the predicted hotspot information.
[0124] According to the first weight and the second weight of the first target hotspot information, the weight of the first target hotspot information is determined, the first target hotspot information being any one target hotspot information.
[0125] In one or more embodiments of the present embodiment, the determination module 404 is further configured to:
[0126] Calculate the weighted sum of the first weight and the second weight of the first target hotspot information, and determine the weighted sum as the weight of the first target hotspot information.
[0127] In one or more embodiments of the present embodiment, the determination module 404 is further configured to:
[0128] According to the current hotspot data and the user historical query data, a target hotspot information set is determined.
[0129] The information in the target hotspot information set is determined as target hotspot information, and at least one target hotspot information is obtained.
[0130] In one or more embodiments of the present embodiment, the target hotspot information set includes current hotspot information and predicted hotspot information.
[0131] Further, the determination module 404 is further configured to:
[0132] According to the current hotspot data, the current hotspot information is determined in the knowledge graph.
[0133] According to the user historical query data, the predicted hotspot information is determined.
[0134] The current hotspot information and the predicted hotspot information are merged to generate a target hotspot information set.
[0135] In one or more embodiments of the present embodiment, the apparatus further comprises a historical hotspot information processing module configured to:
[0136] determine stored historical hotspot information;
[0137] compress the stored historical hotspot information to obtain compressed historical hotspot information;
[0138] replace the stored historical hotspot information with the compressed historical hotspot information.
[0139] In one or more embodiments of the present embodiment, the historical hotspot information processing module is configured to:
[0140] extract identification information of historical hotspot information, the identification information comprising an abstract and / or a label;
[0141] the replacing the stored historical hotspot information with the compressed historical hotspot information comprises:
[0142] storing the identification information of historical hotspot information in association with the compressed historical hotspot information to replace the stored historical hotspot information.
[0143] The data storage apparatus provided in the present application obtains current hotspot data and user historical query data through an obtaining module; determines at least one target hotspot information according to the current hotspot data and the user historical query data through a determining module; further, a storage module extracts target features of each target hotspot information, and stores each target hotspot information and the target features of each target hotspot information in association in a one-to-one correspondence storage manner. In this way, high-frequency access information can be stored, the data content in the storage medium can be dynamically adjusted, the access efficiency can be improved, and the frequency of data transfer between different storage media can be reduced, thereby solving data caching from the data contact level and improving the service life of the storage medium.
[0144] The above is a schematic scheme of a data storage apparatus of the present embodiment. It should be noted that the technical scheme of the data storage apparatus belongs to the same concept as the technical scheme of the data storage method described above, and the details of the technical scheme of the data storage apparatus that are not described in detail can be referred to the description of the technical scheme of the data storage method. In addition, each component in the apparatus embodiment should be understood as a functional module necessary to establish the steps of the program flow or the steps of the method. Each functional module is not limited by actual functional division or separation. The apparatus claim defined by such a group of functional modules should be understood as a functional module architecture for realizing the solution of the computer program mainly by the description of the specification, and should not be understood as an entity apparatus for realizing the solution mainly by hardware.
[0145] Figure 5 A structural block diagram of a computing device 500 is shown according to an embodiment of the present application. The components of the computing device 500 include, but are not limited to, a memory 510 and a processor 520. The processor 520 is connected with the memory 510 through a bus 530, and a database 550 is used to save data.
[0146] The computing device 500 also includes an access device 540, which enables the computing device 500 to communicate via one or more networks 560. Examples of these networks include the public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 540 can include one or more of any type of network interface (e.g., network interface card (NIC)) such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like, either wired or wireless.
[0147] In an embodiment of the present application, the above-mentioned components of the computing device 500 and other components not shown in the figure can be connected with each other, for example, through a bus. It should be understood that, Figure 5 the computing device structure block diagram shown is only for the purpose of example, and is not a limitation on the scope of the present application. Other components can be added or replaced as needed by those skilled in the art. Figure 5 the computing device structure block diagram shown is only for the purpose of example, and is not a limitation on the scope of the present application. Other components can be added or replaced as needed by those skilled in the art.
[0148] The computing device 500 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other type of mobile device, or a stationary computing device such as a desktop computer or PC. The computing device 500 can also be a mobile or stationary server.
[0149] The processor 520 is configured to execute computer-executable instructions of the data storage method.
[0150] The above is a schematic scheme of a computing device according to an embodiment of the present application. It should be noted that the technical scheme of the computing device belongs to the same concept as the technical scheme of the data storage method described above, and the details of the technical scheme of the computing device that are not described in detail can be referred to the description of the technical scheme of the data storage method.
[0151] An embodiment of the present application further provides a computer readable storage medium storing computer instructions, which, when executed by a processor, implement the steps of the data storage method.
[0152] The above is a schematic scheme of the computer readable storage medium of the embodiment of the present application. It should be noted that the technical scheme of the storage medium and the technical scheme of the data storage method described above belong to the same concept, and the details of the technical scheme of the storage medium that are not described in detail can be referred to the description of the technical scheme of the data storage method.
[0153] An embodiment of the present application discloses a chip storing computer instructions, which, when executed by a processor, implement the steps of the data storage method.
[0154] The above describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims can be performed in a different order than those described in the embodiments and still achieve desirable results. In addition, the processes depicted in the figures do not necessarily require the particular order shown or sequential order in order to achieve the desired results. In some implementations, multitasking and parallel processing can be advantageous.
[0155] The computer instructions include computer program code, which can be in the form of source code, object code, executable code, or some intermediate form. The computer readable medium can include any entity or apparatus capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content contained in the computer readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, the computer readable medium does not include electrical carrier signals and telecommunication signals.
[0156] It should be noted that for each method embodiment described above, in order to facilitate description, each is described as a combination of a series of acts, but those skilled in the art should know that the present application is not limited by the order of the acts described, because according to the present application, certain steps can be performed in other order or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the acts and modules involved are not necessarily essential to the present application.
[0157] In the above-described embodiments, the description of each embodiment focuses on different aspects, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.
[0158] The preferred embodiments of the application disclosed above are only used to illustrate the application. Alternative embodiments do not describe all the details and limit the application to the specific embodiments described. Obviously, many modifications and changes can be made according to the content of the application. The application selects and describes these embodiments in order to better explain the principles and practical applications of the application, so that those skilled in the art can well understand and utilize the application. The application is limited by the claims and their full scope and equivalents.
Claims
1. A data storage method, characterized by, The method comprises the following steps: obtaining current hot data and user historical query data; determining at least one target hot information according to the current hot data and the user historical query data; wherein the at least one target hot information is determined according to the combination of the current hot data, the associated data of the current hot data and the predicted hot data; the associated data of the current hot data is determined in the knowledge graph corresponding to the current hot data, and the predicted hot data is obtained by inputting the embedding representation of the user historical query data into a prediction model; the embedding representation is calculated according to the connection relationship of each node in the user query behavior structure graph; the user query behavior structure graph is obtained by connecting each user historical query data as a node according to the sequence of each user historical query data to represent the connection relationship of each data; extracting target features of each target hot information, and storing each target hot information and the target features of each target hot information in a one-to-one corresponding manner.
2. The method of claim 1, wherein, After determining the at least one target hot information, the method further comprises the following steps: determining the weight of each target hot information; sorting each target hot information according to the weight of each target hot information; the method further comprises the following steps: according to the sorting result, the target hot information and the target features of each target hot information are stored in a one-to-one corresponding manner.
3. The method of claim 2, wherein, The method further comprises the following steps: obtaining the first weight and the second weight of each target hot information, the first weight being the weight of each target hot information in the current hot information, and the second weight being the weight of each target hot information in the predicted hot information; determining the weight of the first target hot information according to the first weight and the second weight of the first target hot information, wherein the first target hot information is any one of the target hot information.
4. The method of claim 3, wherein, The method further comprises the following steps: calculating the weighted sum of the first weight and the second weight of the first target hot information, and determining the weighted sum as the weight of the first target hot information.
5. The method according to claim 1 or 2, characterized in that, The method further comprises the following steps: determining a target hot information set according to the current hot data and the user historical query data; determining the target hot information in the target hot information set as the target hot information to obtain at least one target hot information.
6. The method of claim 5, wherein, The target hot information set comprises current hot information and predicted hot information. The method further comprises the following steps: determining the current hot information in the knowledge graph according to the current hot data; determining the predicted hot information according to the user historical query data; merging the current hot information and the predicted hot information to generate a target hot information set.
7. The method of claim 1, wherein, The method further comprises the following steps: determining the stored historical hot information; Compress the stored historical hotspot information to obtain compressed historical hotspot information; Replace the stored historical hotspot information with the compressed historical hotspot information.
8. The method of claim 7, wherein, Before the compression of the stored historical hotspot information, the method further comprises: Extracting identification information of the historical hotspot information, the identification information comprising an abstract and / or a label; The replacement of the stored historical hotspot information with the compressed historical hotspot information comprises: Storing the identification information of the historical hotspot information in association with the compressed historical hotspot information to replace the stored historical hotspot information.
9. A data storage device, characterized by The method comprises: An acquisition module configured to acquire current hotspot data and user historical query data; A determination module configured to determine at least one target hotspot information according to the current hotspot data and the user historical query data; wherein the at least one target hotspot information is determined according to a combination of information of the current hotspot data, associated data of the current hotspot data and predicted hotspot data; the associated data of the current hotspot data is determined according to the current hotspot data in a knowledge graph corresponding to the current hotspot data; the predicted hotspot data is obtained by inputting an embedded representation of the user historical query data into a prediction model for prediction; the embedded representation is calculated according to a connection relationship of each node in a user query behavior structure graph; the user query behavior structure graph is obtained by connecting each user historical query data as a node according to an order in which the user historical query data is queried to represent a graph of data connection relationship; A storage module configured to extract target features of each target hotspot information, and store each target hotspot information and the target features of each target hotspot information in a one-to-one corresponding storage manner.
10. A computing device, comprising: The method comprises: A memory and a processor; The memory is configured to store computer executable instructions, and the processor is configured to execute the computer executable instructions to implement the steps of the data storage method according to any one of claims 1 to 8.
11. A computer-readable storage medium storing computer instructions, wherein, The instructions, when executed by the processor, implement the steps of the data storage method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Hotspot data identification method and device
CN106709068A
Information push method, apparatus and device, and storage medium
CN106982256A
Data storage method and apparatus
CN107463514A
User-oriented graph database query optimization method and system
CN112765411A
Big data and user demand data management method and cloud computing server
CN112765463A