Data acquisition method, device, medium and electronic device based on data structure

By calculating and weighting the inverse document frequency in the data structure, determining the correlation score between the search keywords and each node, the problem of cumbersome and inaccurate data acquisition process in the hierarchical structure data storage is solved, and efficient and accurate data retrieval is achieved.

CN118747179BActive Publication Date: 2025-06-17GUANGDONG GONGDI IOT SERVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410782869.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-17
Publication Date
2025-06-17
Estimated Expiration
2044-06-17

AI Technical Summary

Technical Problem

In the prior art, the data acquisition process of hierarchical structure data storage is cumbersome, inefficient, and data retrieval based on the data structure is inaccurate, resulting in inaccurate user search results.

Method used

By responding to the search keywords input by the user, the data text of each child node in the data structure is compared with the search keywords, the first entry frequency information of the search keywords in each child node is determined, and the second entry frequency information of the root node is determined based on the information. Then, the inverse document frequency is calculated and weighted, and the target inverse document frequency is generated, and the correlation score between the search keyword and each node is determined, and the data text of the node with the largest correlation score is finally obtained.

Benefits of technology

It improves the acquisition efficiency of data retrieval in the data structure and the accuracy of data acquisition, simplifies the data acquisition process, and improves the accuracy of user search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118747179B_ABST
    Figure CN118747179B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a data acquisition method, apparatus, medium and electronic device based on a data structure. The method includes: determining first term frequency information of a retrieval keyword in a child node, determining second term frequency information of a root node corresponding to the child node in the data structure according to the first term frequency information of the child node, determining a plurality of inverse document frequency information of the retrieval keyword in the data structure according to the first term frequency information and the second term frequency information, weighting the plurality of inverse document frequencies to generate a target inverse document frequency, determining a plurality of relevance scores of the child node and the root node according to the target inverse document frequency, the first term frequency information and the second term frequency information, determining the node with the largest relevance score among the plurality of relevance scores as the target node, and acquiring the data text of the target node from the data structure. Thereby, the acquisition efficiency of data retrieval in the data structure and the accuracy of data acquisition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a data acquisition method, apparatus, medium, and electronic device based on a data structure. Background Art

[0002] Nowadays, the structured data storage method has gradually replaced the traditional data storage method and become the mainstream of data storage. Most of them adopt a hierarchical structure. In order to achieve data reuse, inheritance is usually used between each data layer, that is, the child layer has all the data of the parent layer. In the related art, in the hierarchical storage data structure, the data acquisition process is as follows: According to the inheritance relationship between the top layer and the child layer and the reference relationship between nodes in the hierarchical structure, obtain the level where the data is located, and obtain the type and name of the referenced node. Then, according to the determined level, the type and name of the referenced node, obtain the data from the hierarchical structure. However, since the above process needs to obtain the level where the referenced node is located according to the inheritance relationship between the top layer and the child layer and the reference relationship between nodes, the data acquisition process is relatively cumbersome and the efficiency is relatively low. Moreover, during the data retrieval process based on the data structure, due to the limitations of the data structure, the retrieved data obtained by the user is inaccurate. In summary, how to improve the data acquisition efficiency and the accuracy of data acquisition are technical problems that need to be solved by those skilled in the art at present. Summary of the Invention

[0003] To overcome the problems existing in the related art, the present disclosure provides a data acquisition method based on a data structure.

[0004] According to the first aspect of the embodiments of the present disclosure, a data acquisition method based on a data structure is provided. The method includes:

[0005] In response to receiving a retrieval keyword input by a user, compare the data text of each child node in the data structure with the retrieval keyword to determine the first term frequency information of the retrieval keyword in each child node;

[0006] According to the first term frequency information of each child node, determine the second term frequency information of the corresponding root node of each child node in the data structure;

[0007] According to the first term frequency information and the second term frequency information, determine multiple inverse document frequency information of the retrieval keyword in the data structure, and weight the multiple inverse document frequencies according to the weight relationship between each child node and each root node in the data structure to generate a target inverse document frequency;

[0008] Determine the relevance scores between the retrieval keyword and each child node and each root node according to the target inverse document frequency, the first term frequency information, and the second term frequency information, and generate multiple relevance scores;

[0009] Determine the node with the maximum relevance score among the multiple relevance scores as the target node, and obtain the data text of the target node from the data structure.

[0010] Optionally, the comparing the data text of each child node in the data structure with the retrieval keyword to determine the first term frequency information of the retrieval keyword in each child node includes:

[0011] Segment the data text in each child node into multiple term data according to a preset word segmentation rule;

[0012] Filter the multiple term data in each child node according to preset common terms to generate target term data; determine the multiple frequency information of each term in each child node by setting a hash table according to the target term data;

[0013] Determine the first term frequency information of the retrieval keyword in each child node according to the multiple frequency information.

[0014] Optionally, the determining the second term frequency information of the root node corresponding to each child node in the data structure according to the first term frequency information of each child node includes:

[0015] Determine one or more child nodes corresponding to each root node according to the entity relationship between the nodes in the data structure;

[0016] Determine the first term frequency feature vector of the first root node according to the first term frequency information of the multiple child nodes corresponding to the first root node, where the first root node is any root node in the data structure;

[0017] Determine the second term frequency information of the first root node according to the first term frequency feature vector.

[0018] Optionally, the determining the second term frequency information of the first root node according to the first term frequency feature vector includes:

[0019] Calculate the term frequency weight information of the multiple terms corresponding to the retrieval keyword in the first term frequency feature vector through the TF-IDF algorithm;

[0020] Determine the second term frequency information of the first root node according to the term frequency weight information.

[0021] Optionally, determining a relevance score between the retrieval keyword and each child node and each root node according to the target inverse document frequency, the first term frequency information, and the second term frequency information, and generating a plurality of relevance scores, including:

[0022] Determining a first term frequency vector corresponding to the first term frequency information and a second term frequency vector corresponding to the second term frequency information;

[0023] Determining a keyword vector corresponding to the retrieval keyword according to the target inverse document frequency;

[0024] Multiplying the first term frequency vector by the keyword vector to obtain a first relevance score for the root node;

[0025] Multiplying the second term frequency vector by the keyword vector to obtain a second relevance score for the child node.

[0026] Optionally, weighting the plurality of inverse document frequencies according to the weight relationship between each child node and each root node in the data structure to generate a target inverse document frequency, including:

[0027] Determining the connection relationship between each child node and each root node in the data structure;

[0028] Determining the weight relationship between each child node and each root node according to the connection relationship;

[0029] Weighting the plurality of inverse document frequencies according to the weight relationship to generate the target inverse document frequency.

[0030] According to a second aspect of the embodiments of the present disclosure, there is provided a data acquisition device based on a data structure, including:

[0031] A first determination module, configured to compare the data text of each child node in the data structure with the retrieval keyword in response to receiving the retrieval keyword input by the user, and determine the first term frequency information of the retrieval keyword in each child node;

[0032] A second determination module, configured to determine the second term frequency information of the root node corresponding to each child node in the data structure according to the first term frequency information of each child node;

[0033] A third determination module, configured to determine a plurality of inverse document frequency information of the retrieval keyword in the data structure according to the first term frequency information and the second term frequency information, and weight the plurality of inverse document frequencies according to the weight relationship between each child node and each root node in the data structure to generate a target inverse document frequency;

[0034] A fourth determination module, configured to determine a relevance score between the retrieval keyword and each sub-node and each root node according to the target inverse document frequency, the first term frequency information, and the second term frequency information, and generate a plurality of relevance scores;

[0035] An execution module, configured to determine, from the plurality of relevance scores, a node with the largest relevance score as the target node, and obtain the data text of the target node from the data structure.

[0036] Optionally, the first determination module is configured to:

[0037] Segment the data text in each sub-node into a plurality of term data according to a preset word segmentation rule;

[0038] Screen the plurality of term data in each sub-node according to a preset common term to generate target term data; and determine a plurality of frequency information of each term in each sub-node by setting a hash table according to the target term data;

[0039] Determine the first term frequency information of the retrieval keyword in each sub-node according to the plurality of frequency information.

[0040] According to a third aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method according to any one of the first aspects are implemented.

[0041] According to a fourth aspect of the embodiments of the present disclosure, there is provided an electronic device, including:

[0042] A processor;

[0043] A memory for storing executable instructions of the processor;

[0044] Wherein, the processor is configured to execute the executable instructions in the memory to implement the method according to any one of the first aspects.

[0045] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:

[0046] In the above manner, in response to receiving a retrieval keyword input by a user, the data text of each child node in the data structure is compared with the retrieval keyword to determine the first term frequency information of the retrieval keyword in each child node. According to the first term frequency information of each child node, the second term frequency information of the corresponding root node of each child node in the data structure is determined. According to the first term frequency information and the second term frequency information, a plurality of inverse document frequency information of the retrieval keyword in the data structure is determined, and according to the weight relationship between each child node and each root node in the data structure, the plurality of inverse document frequencies are weighted to generate a target inverse document frequency. According to the target inverse document frequency, the first term frequency information, and the second term frequency information, the correlation score between the retrieval keyword and each child node and each root node is determined, and a plurality of correlation scores are generated. The node with the largest correlation score is determined as the target node from the plurality of correlation scores, and the data text of the target node is obtained from the data structure. Thereby, the acquisition efficiency of data retrieval in the data structure and the accuracy of data acquisition are improved.

[0047] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.

[0049] Figure 1 is a flowchart of a data acquisition method based on a data structure shown according to an exemplary embodiment.

[0050] Figure 2 is a block diagram of a data acquisition device based on a data structure shown according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0052] It should be noted that all actions of obtaining signals, information, or data in this application are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where the location is located and obtaining authorization from the owner of the corresponding device.

[0053] Figure 1 is a flowchart of a data acquisition method based on a data structure shown according to an exemplary embodiment. Refer to Figure 1 As shown, the method includes:

[0054] In step S11, in response to receiving a retrieval keyword input by a user, compare the data texts of each sub-node in the data structure with the retrieval keyword to determine the first term frequency information of the retrieval keyword in each sub-node.

[0055] Exemplarily, this embodiment is applied to a terminal. In this terminal, data is stored in the form of a data structure, and this data structure is a bifurcated tree data structure. In this data structure, there are multiple root nodes, and each root node includes at least one sub-node. When the terminal receives a retrieval keyword input by a user, compare the data texts corresponding to each sub-node in the data structure with the retrieval keyword, so as to determine the first term frequency information of the retrieval keyword in each sub-node. This first term frequency information is used to indicate the frequency of the term corresponding to the retrieval keyword appearing in each node.

[0056] In step S12, according to the first term frequency information of each sub-node, determine the second term frequency information of the root node corresponding to each sub-node in the data structure.

[0057] Exemplarily, in this embodiment, according to the first term frequency information of each sub-node, determine the second term frequency of the root node corresponding to each sub-node in the data structure. For example, according to the relationship between each sub-node and each root node in the data structure, determine the sub-nodes corresponding to the root node, and calculate the average value of the first term frequency information of each sub-node, so as to obtain the second term frequency information of the root node corresponding to each sub-node in the data structure.

[0058] In step S13, according to the first term frequency information and the second term frequency information, determine multiple inverse document frequency information of the retrieval keyword in the data structure, and weight the multiple inverse document frequencies according to the weight relationship between each sub-node and each root node in the data structure to generate a target inverse document frequency.

[0059] Exemplarily, in this embodiment, the TF-IDF weighted statistics method is adopted to determine the relevance scores of the retrieval keywords in each node of the data structure, and then the data text corresponding to the node with the largest relevance score is determined from multiple relevance scores as the data text retrieved by the user through the retrieval keywords. After determining the first term frequency information of each child node and the second term frequency information of each root node in the above manner, multiple inverse document frequency information of the retrieval keywords in the data structure is determined. In this embodiment, the IDF calculation formula in the TF-IDF weighted statistics algorithm is adopted to determine the multiple inverse document frequency information in the data structure. Then, according to the weight relationship between each child node and each root node in the data structure, the multiple inverse document frequencies are weighted to generate the target inverse document frequency.

[0060] In step S14, according to the target inverse document frequency, the first term frequency information, and the second term frequency information, determine the relevance scores between the retrieval keyword and each child node and each root node, and generate multiple relevance scores.

[0061] Exemplarily, in this embodiment, the TF-IDF weighted statistics algorithm is adopted to calculate and determine the relevance scores between the retrieval keyword and the child nodes and root nodes according to the target inverse document frequency, the first term frequency, and the second term frequency, and generate multiple relevance scores.

[0062] In step S15, determine the node with the largest relevance score among the multiple relevance scores as the target node, and obtain the data text of the target node from the data structure.

[0063] Exemplarily, determine the node with the largest relevance score among the multiple relevance scores as the target node, and obtain the data text of the target node from the data structure as the data text corresponding to the retrieval keyword input by the user. In this embodiment, the efficiency and accuracy of data acquisition are improved in the above manner.

[0064] In the above manner, in response to receiving the retrieval keyword input by the user, the data text of each child node in the data structure is compared with the retrieval keyword to determine the first term frequency information of the retrieval keyword in each child node. According to the first term frequency information of each child node, the second term frequency information of the root node corresponding to each child node in the data structure is determined. According to the first term frequency information and the second term frequency information, multiple inverse document frequency information of the retrieval keyword in the data structure is determined, and according to the weight relationship between each child node and each root node in the data structure, the multiple inverse document frequencies are weighted to generate the target inverse document frequency. According to the target inverse document frequency, the first term frequency information and the second term frequency information, the correlation score between the retrieval keyword and each child node and each root node is determined to generate multiple correlation scores. The node with the largest correlation score is determined as the target node from the multiple correlation scores, and the data text of the target node is obtained from the data structure. Thereby improving the acquisition efficiency of data retrieval and the accuracy of data acquisition in the data structure.

[0065] Optionally, the comparing the data text of each child node in the data structure with the retrieval keyword to determine the first term frequency information of the retrieval keyword in each child node includes:

[0066] Dividing the data text in each child node into multiple term data according to a preset word segmentation rule;

[0067] Filtering the multiple term data in each child node according to a preset common term to generate target term data; according to the target term data, determining multiple frequency information of each term in each child node by setting a hash table;

[0068] According to the multiple frequency information, determining the first term frequency information of the retrieval keyword in each child node.

[0069] Optionally, the determining the second term frequency information of the root node corresponding to each child node in the data structure according to the first term frequency information of each child node includes:

[0070] According to the entity relationship between the nodes in the data structure, determining one or more child nodes corresponding to each root node;

[0071] According to the first term frequency information of multiple child nodes corresponding to the first root node, determining the first term frequency feature vector of the first root node, where the first root node is any root node in the data structure;

[0072] According to the first term frequency feature vector, determining the second term frequency information of the first root node.

[0073] Optionally, determining the second term frequency information of the first root node according to the first term frequency feature vector includes:

[0074] Calculating the term frequency weight information of multiple terms corresponding to the retrieval keyword in the first term frequency feature vector through the TF-IDF algorithm;

[0075] Determining the second term frequency information of the first root node according to the term frequency weight information.

[0076] Optionally, determining the correlation scores between the retrieval keyword and each child node and each root node according to the target inverse document frequency, the first term frequency information, and the second term frequency information, and generating multiple correlation scores, includes:

[0077] Determining a first term frequency vector corresponding to the first term frequency information and a second term frequency vector corresponding to the second term frequency information;

[0078] Determining a keyword vector corresponding to the retrieval keyword according to the target inverse document frequency;

[0079] Multiplying the first term frequency vector by the keyword vector to obtain a first correlation score of the root node;

[0080] Multiplying the second term frequency vector by the keyword vector to obtain a second correlation score of the child node.

[0081] Optionally, weighting the multiple inverse document frequencies according to the weight relationship between each child node and each root node in the data structure to generate a target inverse document frequency, includes:

[0082] Determining the connection relationship between each child node and each root node in the data structure;

[0083] Determining the weight relationship between each child node and each root node according to the connection relationship;

[0084] Weighting the multiple inverse document frequencies according to the weight relationship to generate the target inverse document frequency.

[0085] Figure 2 It is a block diagram of a data acquisition device based on a data structure shown according to an exemplary embodiment. As Figure 2 shown, the device 100 includes:

[0086] The first determination module 110 is configured to compare the data text of each child node in the data structure with the retrieved keyword in response to receiving the retrieved keyword input by the user, and determine the first term frequency information of the retrieved keyword in each child node;

[0087] The second determination module 120 is configured to determine the second term frequency information of the root node corresponding to each child node in the data structure according to the first term frequency information of each child node;

[0088] The third determination module 130 is configured to determine multiple inverse document frequency information of the retrieved keyword in the data structure according to the first term frequency information and the second term frequency information, and weight the multiple inverse document frequencies according to the weight relationship between each child node and each root node in the data structure to generate a target inverse document frequency;

[0089] The fourth determination module 140 is configured to determine the correlation score between the retrieved keyword and each child node and each root node according to the target inverse document frequency, the first term frequency information, and the second term frequency information, and generate multiple correlation scores;

[0090] The execution module 150 is configured to determine the node with the largest correlation score among the multiple correlation scores as the target node, and obtain the data text of the target node from the data structure.

[0091] Optionally, the first determination module is configured to:

[0092] Segment the data text in each child node into multiple term data according to a preset word segmentation rule;

[0093] Screen the multiple term data in each child node according to a preset common term to generate target term data; and determine multiple frequency information of each term in each child node by setting a hash table according to the target term data;

[0094] Determine the first term frequency information of the retrieved keyword in each child node according to the multiple frequency information.

[0095] Optionally, the second determination module 120 includes:

[0096] The first determination sub-module is configured to determine one or more child nodes corresponding to each root node according to the entity relationship between the nodes in the data structure;

[0097] The second determination sub-module is configured to determine the first term frequency feature vector of the first root node according to the first term frequency information of multiple child nodes corresponding to the first root node, where the first root node is any root node in the data structure;

[0098] A third determination sub-module, configured to determine second term frequency information of the first root node according to the first term frequency feature vector.

[0099] Optionally, the third determination sub-module is configured to

[0100] calculate term frequency weight information of multiple terms corresponding to the retrieval keyword in the first term frequency feature vector through the TF-IDF algorithm;

[0101] determine the second term frequency information of the first root node according to the term frequency weight information.

[0102] Optionally, a fourth determination module is configured to:

[0103] determine a first term frequency vector corresponding to the first term frequency information and a second term frequency vector corresponding to the second term frequency information;

[0104] determine a keyword vector corresponding to the retrieval keyword according to the target inverse document frequency;

[0105] multiply the first term frequency vector by the keyword vector to obtain a first relevance score of the root node;

[0106] multiply the second term frequency vector by the keyword vector to obtain a second relevance score of the child node.

[0107] Optionally, the third determination module is configured to:

[0108] determine the connection relationship between each child node and each root node in the data structure;

[0109] determine the weight relationship between each child node and each root node according to the connection relationship;

[0110] weight the multiple inverse document frequencies according to the weight relationship to generate the target inverse document frequency.

[0111] In the above manner, in response to receiving a search keyword input by a user, the data texts of each child node in the data structure are compared with the search keyword to determine the first term frequency information of the search keyword in each child node. According to the first term frequency information of each child node, the second term frequency information of the corresponding root node of each child node in the data structure is determined. According to the first term frequency information and the second term frequency information, a plurality of inverse document frequency information of the search keyword in the data structure is determined, and according to the weight relationship between each child node and each root node in the data structure, the plurality of inverse document frequencies are weighted to generate a target inverse document frequency. According to the target inverse document frequency, the first term frequency information, and the second term frequency information, the correlation scores between the search keyword and each child node and each root node are determined to generate a plurality of correlation scores. The node with the largest correlation score is determined as the target node from the plurality of correlation scores, and the data text of the target node is obtained from the data structure. Thereby, the acquisition efficiency of data retrieval and the accuracy of data acquisition in the data structure are improved.

[0112] An embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the method described in any one of the foregoing embodiments are implemented.

[0113] An embodiment of the present disclosure further provides an electronic device, including:

[0114] a processor;

[0115] a memory for storing executable instructions of the processor;

[0116] wherein the processor is configured to execute the executable instructions in the memory to implement the method described in any one of the foregoing embodiments.

[0117] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0118] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A data acquisition method based on data structure, characterized in that: The method comprises: In response to receiving a search keyword input by a user, comparing the data text of each sub-node in the data structure with the search keyword to determine first term frequency information of the search keyword in each sub-node; Determine, according to the first term frequency information of each child node, the second term frequency information of the root node corresponding to each child node in the data structure; The determining, based on the first term frequency information of each child node, the second term frequency information of the root node corresponding to each child node in the data structure includes: Determine one or more child nodes corresponding to each root node according to the entity relationship between the nodes in the data structure; Determine a first term frequency feature vector of the first root node according to first term frequency information of a plurality of child nodes corresponding to the first root node, wherein the first root node is any root node in the data structure; Determining second term frequency information of the first root node according to the first term frequency feature vector; Determine multiple inverse document frequency information of the search keyword in the data structure according to the first term frequency information and the second term frequency information, and weight the multiple inverse document frequencies according to the weight relationship between each child node and each root node in the data structure to generate a target inverse document frequency; Determine the relevance score between the search keyword and each child node and each root node according to the target inverse document frequency, the first term frequency information and the second term frequency information, and generate a plurality of relevance scores; The step of determining the relevance score between the search keyword and each child node and each root node according to the target inverse document frequency, the first term frequency information, and the second term frequency information, and generating a plurality of relevance scores includes: Determine a first term frequency vector corresponding to the first term frequency information, and a second term frequency vector corresponding to the second term frequency information; Determining a keyword vector corresponding to the search keyword according to the target inverse document frequency; Multiplying the first word frequency vector by the keyword vector to obtain a first relevance score of the root node; Multiplying the second word frequency vector by the keyword vector to obtain a second relevance score of the child node; A node with the largest correlation score is determined as a target node from among the multiple correlation scores, and data text of the target node is acquired from the data structure.

2. The data acquisition method based on data structure according to claim 1, characterized in that: The step of comparing the data text of each sub-node in the data structure with the search keyword to determine the first term frequency information of the search keyword in each sub-node includes: The data text in each child node is divided into multiple word data according to the preset word segmentation rules; Filtering the plurality of term data in each sub-node according to preset common terms to generate target term data; determining a plurality of frequency information of each term in each sub-node by setting a hash table according to the target term data; The first term frequency information of the search keyword in each sub-node is determined according to the multiple frequency information.

3. The data acquisition method based on data structure according to claim 1, characterized in that: The determining, according to the first word frequency feature vector, second word frequency information of the first root node includes: Calculate the word frequency weight information of the multiple entries corresponding to the search keyword in the first word frequency feature vector by using the TF-IDF algorithm; The second term frequency information of the first root node is determined according to the term frequency weight information.

4. The data acquisition method based on data structure according to claim 1, characterized in that: The step of weighting the multiple inverse document frequencies according to the weight relationship between each child node and each root node in the data structure to generate a target inverse document frequency includes: Determine the connection relationship between each child node and each root node in the data structure; Determine the weight relationship between each child node and each root node according to the connection relationship; The multiple inverse document frequencies are weighted according to the weight relationship to generate the target inverse document frequency.

5. A data acquisition device based on data structure, characterized in that: The device comprises: A first determination module is used for, in response to receiving a search keyword input by a user, comparing the data text of each sub-node in the data structure with the search keyword, and determining first term frequency information of the search keyword in each sub-node; A second determination module is used to determine the second term frequency information of the root node corresponding to each child node in the data structure according to the first term frequency information of each child node; the determination of the second term frequency information of the root node corresponding to each child node in the data structure according to the first term frequency information of each child node includes: determining one or more child nodes corresponding to each root node according to the entity relationship between each node in the data structure; determining the first term frequency feature vector of the first root node according to the first term frequency information of multiple child nodes corresponding to the first root node, the first root node being any root node in the data structure; determining the second term frequency information of the first root node according to the first term frequency feature vector; A third determination module is used to determine a plurality of inverse document frequency information of the search keyword in the data structure according to the first term frequency information and the second term frequency information, and weight the plurality of inverse document frequencies according to the weight relationship between each child node and each root node in the data structure to generate a target inverse document frequency; A fourth determination module, for determining the correlation scores between the search keyword and each child node and each root node according to the target inverse document frequency, the first term frequency information and the second term frequency information, and generating multiple correlation scores; the determining the correlation scores between the search keyword and each child node and each root node according to the target inverse document frequency, the first term frequency information and the second term frequency information, and generating multiple correlation scores, comprises: determining a first term frequency vector corresponding to the first term frequency information, and a second term frequency vector corresponding to the second term frequency information; determining a keyword vector corresponding to the search keyword according to the target inverse document frequency; multiplying the first term frequency vector by the keyword vector to obtain a first correlation score for the root node; multiplying the second term frequency vector by the keyword vector to obtain a second correlation score for the child node; An execution module is used to determine a node with the largest correlation score from the multiple correlation scores as a target node, and obtain data text of the target node from the data structure.

6. The data acquisition device based on data structure according to claim 5, characterized in that: The first determining module is used to: The data text in each child node is divided into multiple word data according to the preset word segmentation rules; Filter the multiple entry data in each sub-node according to preset common entries to generate target entry data; According to the target term data, determining multiple frequency information of each term in each child node by setting a hash table; The first term frequency information of the search keyword in each sub-node is determined according to the multiple frequency information.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 4 are implemented.

8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to execute executable instructions in the memory to implement the method described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Text matching method and system based on text correlation and electronic equipment

    CN113569009A

  • Method for searching data

    US20090171945A1