Data query method, device, equipment and storage medium based on text vector

Through the data query method based on text vectors, index items are generated and data matching model query is queried, which solves the problem of low data query efficiency in the construction cost list, and achieves fast and accurate data acquisition.

CN118520124BActive Publication Date: 2025-05-06RAFTER & PURLIN TECH (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410395665.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-02
Publication Date
2025-05-06
Estimated Expiration
2044-04-02

AI Technical Summary

Technical Problem

In the construction cost list, relevant personnel are inefficient when searching for the required data in a large amount of data, and the query words are inconsistent with the data description words, which leads to difficulty in querying.

Method used

The data query method based on text vector is adopted, and the search terms are obtained by obtaining the search terms provided by the user, synonyms and conjunctions are determined, index terms are generated, and text feature vectors are generated through the preset data matching model, and matching results are queried from the vector data table.

Benefits of technology

Improve the efficiency and accuracy of data query in the construction cost list, reduce the dependence on experienced personnel, and ensure the rapid and accurate acquisition of required data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118520124B_ABST
    Figure CN118520124B_ABST
Patent Text Reader

Abstract

The present application discloses a data query method, device, equipment and storage medium based on text vectors, and belongs to the field of data management technology. The present application obtains the search terms provided by the user, and determines the synonyms of the search terms from the pre-equipment word library, and determines the associated words that have a mapping relationship with the search terms from the pre-equipment word library; generates index items according to the search terms, synonyms and associated words; generates text feature vectors corresponding to the index items through a preset data matching model, and obtains matching results from a vector data table based on the text feature vectors; the vector data table is a feature vector data table obtained after the data matching model pre-performs text semantic analysis on the corresponding data in the construction cost list, that is, generates index items through refinement, and searches for matching results corresponding to the search terms through a preset data matching model, so as to ensure the effect of data query is achieved accurately and quickly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data management technology, and in particular to a data query method, device, equipment and storage medium based on text vectors. Background Art

[0002] In the construction industry, the construction cost list, as an important part of construction project management, involves most of the data in construction project management, such as the work content of the construction project, the amount of materials consumed, the working hours of employees, and the amount-related data.

[0003] Among them, most of the data in the construction cost list will be stored in Excel tables. In the actual construction project management, experienced senior cost engineers will manually search in the historical data according to project requirements, or perform some simple text matching to search for historical data, and formulate subsequent regulations based on the historical data.

[0004] However, due to the large amount of data corresponding to the construction cost list, the search for historical data is highly dependent on the memory of data by experienced personnel. At the same time, when they query the data, the words used for the query are not necessarily consistent with the corresponding data description words in the table, resulting in low efficiency for relevant personnel when searching for required data in a large amount of construction cost list data. Summary of the invention

[0005] The main purpose of this application is to provide a data query method, device, equipment and storage medium based on text vectors, aiming to solve the technical problem of low efficiency when relevant personnel are looking for required data in a large amount of construction cost list.

[0006] To achieve the above object, the present application provides a data query method based on text vectors, and the data query method based on text vectors comprises the following steps:

[0007] Acquire a search term provided by a user, and determine a synonym of the search term from a pre-equipped word library, and determine an associated term that has a mapping relationship with the search term from the pre-equipped word library;

[0008] Generate index items according to the search terms, the synonyms and the associated terms;

[0009] Generate a text feature vector corresponding to the index item through a preset data matching model, and query the vector data table to obtain a matching result based on the text feature vector;

[0010] Among them, the preset data matching model is obtained by iteratively training the language model to be trained based on text feature training samples and sample labels corresponding to the text feature training samples, and the vector data table is a feature vector data table obtained after the data matching model pre-performs text semantic analysis on the corresponding data in the construction cost list.

[0011] Optionally, the step of generating index items according to the search terms, the synonyms and the associated terms includes:

[0012] Combining the search term, the synonym and the associated word to obtain an initial data group, wherein the initial data is a data group pointed to by the search term or the synonym to the associated word;

[0013] According to the mapping relationship between the search term / the synonym and the associated term, the initial data group is modified to obtain a triple;

[0014] The text semantics of the triples are analyzed, and index items of different text contents are generated according to the text semantics and a preset template.

[0015] Optionally, the step of modifying the initial data set according to the mapping relationship between the search term / the synonym and the associated term to obtain a triple includes:

[0016] When the mapping relationship between the search term / the synonym and the associated term is a text format unified relationship, determining whether there is a query record corresponding to the search term in the associated term;

[0017] If it exists, modify the initial data group according to the definition formula corresponding to the query record to obtain a triple;

[0018] or,

[0019] When the mapping relationship between the search term / the synonym and the associated term is a content-associated relationship, a correlation degree analysis is performed on the initial data group to obtain an analysis result;

[0020] According to the analysis result, a data group with a correlation degree greater than a preset threshold is screened out from the initial data group, and the screened data group is modified according to the content correlation relationship to obtain a triple.

[0021] Optionally, the step of generating the text feature vector corresponding to the index item through a preset data matching model includes:

[0022] Performing semantic recognition on the data content corresponding to the index item;

[0023] Extracting multi-level feature information corresponding to the index item according to the result of semantic recognition, wherein the multi-level feature information is feature information corresponding to words in the index item corresponding to different index ranges, and the level to which each feature information belongs in the multi-level feature information is determined according to the size of the index range;

[0024] The text feature vector corresponding to the index item is generated according to the multi-level feature information through the data matching model, wherein the text feature vector includes a sub-vector corresponding to each level feature in the multi-level feature information.

[0025] Optionally, the step of generating the text feature vector corresponding to the index item according to the multi-level feature information through the data matching model includes:

[0026] By using the data matching model, according to the multi-level feature information, a feature pointing vector between feature information of each level is calculated;

[0027] A text feature vector corresponding to the index item and composed of the multi-level feature information is generated according to the feature pointing vector.

[0028] Optionally, the step of querying and obtaining a matching result from a vector data table according to the text feature vector includes:

[0029] Calculate the cosine similarity between the text feature vector and each data in the vector data table;

[0030] The corresponding data whose cosine similarity is greater than the preset similarity is taken as the matching result.

[0031] Optionally, before the steps of obtaining the search term provided by the user, determining the synonyms of the search term from the pre-equipped word library, and determining the associated words that have a mapping relationship with the search term from the pre-equipped word library, the method further includes:

[0032] Obtaining list data corresponding to the building cost list, and classifying the list data to obtain a first data table with a unified text format and a second data table with related text content;

[0033] Extracting characteristic words that meet preset conditions according to the text content of the second data table, and determining replacement words with the same semantics as the characteristic words;

[0034] The second data table is updated according to the replacement words and the characteristic words, and a pre-device selected word library is constructed according to the updated second data table and the first data table.

[0035] In addition, to achieve the above-mentioned purpose, the present application also provides a data query device based on text vectors, and the data query device based on text vectors includes:

[0036] An acquisition module, used to acquire a search term provided by a user, and determine a synonym of the search term from a pre-equipped word library, and determine an associated term that has a mapping relationship with the search term from the pre-equipped word library;

[0037] A generating module, used for generating index items according to the search terms, the synonyms and the associated terms;

[0038] The query module is used to generate a text feature vector corresponding to the index item through a preset data matching model, and to obtain a matching result from a vector data table based on the text feature vector; wherein the preset data matching model is obtained by iteratively training a language model to be trained based on text feature training samples and sample labels corresponding to the text feature training samples, and the vector data table is a feature vector data table obtained after the data matching model pre-performs text semantic analysis on the corresponding data in the construction cost list.

[0039] In addition, to achieve the above-mentioned purpose, the present application also provides a text vector-based data query device, which includes: a memory, a processor, and a text vector-based data query program stored in the memory and executable on the processor, wherein the text vector-based data query program is configured to implement the steps of the text vector-based data query method as described above.

[0040] In addition, to achieve the above-mentioned purpose, the present application also provides a computer-readable storage medium, on which a text vector-based data query program is stored. When the text vector-based data query program is executed by a processor, the steps of the text vector-based data query method as described above are implemented.

[0041] The present application obtains a search term provided by a user, and determines a synonym of the search term from a pre-selected word library, and determines an associated word that has a mapping relationship with the search term from the pre-selected word library; generates an index item according to the search term, the synonym and the associated word; generates a text feature vector corresponding to the index item through a preset data matching model, and obtains a matching result by querying from a vector data table according to the text feature vector; wherein the preset data matching model is obtained by iteratively training a language model to be trained based on text feature training samples and sample labels corresponding to the text feature training samples, and the vector data table is a feature vector data table obtained after the data matching model pre-performs text semantic analysis on the corresponding data in the construction cost list, that is, index items are generated by refinement, and matching results corresponding to the search term are found through the preset data matching model, so as to ensure accurate and fast data query effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 This is a flowchart of the first embodiment of the text vector-based data query method of the present application;

[0043] Figure 2 This is a flowchart of step S20 in the second embodiment of the data query method based on text vectors of the present application;

[0044] Figure 3 This is a structural block diagram of an embodiment of a data query device based on text vectors of the present application;

[0045] Figure 4 This is a schematic diagram of the device structure of the hardware operating environment involved in the embodiment of the present application.

[0046] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0047] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0048] Reference Figure 1 , Figure 1 This is a flowchart of the first embodiment of the data query method based on text vectors of the present application.

[0049] In a first embodiment, the data query method based on text vectors comprises the following steps:

[0050] S10, obtaining a search term provided by a user, and determining a synonym of the search term from a pre-equipped word library, and determining an associated term that has a mapping relationship with the search term from the pre-equipped word library;

[0051] It should be noted that the user in this embodiment refers to a person who needs to query corresponding data in the construction cost list. The user can retrieve the corresponding data in the list through his own query method. The query method includes the process in which the user enters a search term or searches step by step according to the form content. In this embodiment, the main problems that may be encountered in the process of quick query with the search terms provided by the user are improved. The problems that may be encountered include but are not limited to the fact that the words in the list and the search terms have the same meaning but the expression effects are very different, the meaning expressed by the search terms is related to the content originally searched by the user but cannot be directly retrieved through the search terms, or the content represented by the corresponding data in the list is consistent with the corresponding expression content of the search terms, but the search terms do not appear, etc.

[0052] In this embodiment, in order to solve the above technical problems, corresponding solutions are proposed, which specifically include two parts. First, based on the search terms provided by the user, index items with a wider search range are generated to ensure accurate search results. Second, text vectors are generated through the index items and the preset model, and the effect of fast retrieval is achieved through the corresponding features of the vectors, that is, the effects of fast retrieval and high-precision retrieval are achieved at the same time.

[0053] It is understandable that the search terms are characters or words provided by the user for general indexing situations. For example, the user enters the two words "bridge suspension" and "cost" as index conditions through the terminal to query the corresponding data from the construction cost list (hereinafter referred to as the list).

[0054] It should be noted that the search term may be inaccurate. For example, the search term is too colloquial, the text expression in the search term is different from that in the list, etc. Therefore, after obtaining the user's search term, it is necessary to determine the index items involving a larger query scope based on the search term.

[0055] It is understandable that after determining the search term, synonyms similar to the search term are found from the pre-equipped word library, and associated words that have a mapping relationship with the search term are found from the pre-equipped word library, and through the synonyms and associated words, the index scope is expanded based on the search term.

[0056] The synonyms may be understood as words or foreign words that have the same or highly similar content as the search term. For example, concrete and cement may be defined as synonyms.

[0057] Among them, the associated term is a phrase or foreign word that has a mapping relationship with the search term, and the mapping relationship includes the relationship that the semantic content corresponding to the search term and the semantic content corresponding to the data in the pre-equipped word library (semantic association) are the same (or associated), or the character format of the search term and the corresponding format of the data in the pre-equipped word library (text format association) are consistent.

[0058] Among them, the pre-equipped word library includes alternative words such as synonyms and related words for expanding the search coverage. The pre-equipped word library can be continuously iterated and updated during the actual application process. The purpose of the pre-equipped word library is mainly to provide alternative words for retrieval, and secondly, to increase the speed of increasing the index scope. At the same time, the words in the pre-equipped word library can be used to avoid overly colloquial words used in search terms (that is, it is equivalent to standardizing the search terms to text terms in the list through alternative terms, and the text terms will formulate corresponding data formats according to the text format of the corresponding author of the list data).

[0059] In this embodiment, before the steps of obtaining the search term provided by the user and determining the synonyms of the search term from the pre-equipment word library, and determining the associated words that have a mapping relationship with the search term from the pre-equipment word library, the method also includes: obtaining the list data corresponding to the construction cost list, and classifying the list data to obtain a first data table with a unified text format and a second data table related to the text content; based on the text content of the second data table, extracting feature words that meet preset conditions, and determining replacement words with the same semantics as the feature words; updating the second data table based on the replacement words and the feature words, and constructing a pre-equipment word library based on the updated second data table and the first data table.

[0060] It is understandable that synonyms and related words can be determined in the pre-equipped word library, wherein the related words include semantic associations and text format associations. Therefore, before using the pre-equipped word library, it is necessary to construct a corresponding database through semantic associations, text format associations, and synonyms and other feature situations, so as to construct the pre-equipped word library while satisfying the effect of determining synonyms and related words from it at the same time.

[0061] Among them, when constructing the pre-equipment word library, the data in the construction cost list is mainly used as the basis. First, the data in the list is identified, and the content and text format corresponding to the data are analyzed. Based on the results of semantic recognition, characteristic words are extracted, such as cost, price, bridge, concrete, component, steel bar, working hours and other characteristic words. According to the meaning represented by the characteristic word in the corresponding data content, the replacement word with similar semantics is determined, thereby obtaining a data table for determining synonyms and determining semantic content associations. At the same time, according to the recognition results of the text format, a data table of associated words for determining text format associations can be obtained.

[0062] Specifically, in this embodiment, a first data table and a second data table are divided. The first data table contains data in a unified text format, and the second data table contains data with certain correlation in text content. This is equivalent to classifying the data in the construction cost list after obtaining the list, and dividing it according to text format and text content. Therefore, there should be some overlapping data in the first data table and the second data table.

[0063] Among them, when classifying to obtain the second data table, the data are mainly divided according to the text content corresponding to each data, and the related and identical data are classified into one category. For example, the data on road construction cost, the number of personnel required for road paving, the working hours of workers and the materials corresponding to road paving are uniformly classified into one category. After classification, characteristic words are extracted from the classified data. According to the results of the above classification, characteristic words that can be extracted include: cost, construction price, etc. After the characteristic words are extracted, the replacement words corresponding to the characteristic words are determined, thereby achieving the effect of synonym expansion, and the replacement words are updated to the second data table.

[0064] It should be noted that the first data table and the second data table may be multiple tables of different types. For example, the text format may exist in multiple formats, and the text content may have multiple aspects (cost content, or different types of building content, etc.), that is, multiple data tables with different contents are stored in the pre-equipment vocabulary library.

[0065] S20, generating index items according to the search terms, the synonyms and the associated terms;

[0066] It is understandable that synonyms and related words combined with search terms can improve the search scope corresponding to the above three types of words, and generate index items by combining the above three types of words.

[0067] S30, generating a text feature vector corresponding to the index item through a preset data matching model, and obtaining a matching result from a vector data table based on the text feature vector; wherein the preset data matching model is obtained by iteratively training a language model to be trained based on text feature training samples and sample labels corresponding to the text feature training samples, and the vector data table is a feature vector data table obtained after the data matching model pre-performs text semantic analysis on the corresponding data in the construction cost list.

[0068] It can be understood that by inputting the index items into a preset data matching model, and through the data matching model, the index items are vectorized to obtain text feature vectors, and at the same time, the data in the construction cost list are vectorized to obtain a vector data table, in which the vectors corresponding to the data in the construction cost list are stored.

[0069] Among them, when the corresponding data in the construction cost list is converted into a vector through the data matching model, the text semantic analysis of the corresponding data in the list is mainly carried out, and after the analysis, it is converted into a high-dimensional feature vector, which is equivalent to semantic recognition of each paragraph of text content, and according to the results of semantic recognition, the content corresponding to each data in the current list is determined, and the vector data table can be summarized by vector description.

[0070] Among them, by using the data matching model, the text feature vector and the vector data corresponding to the vector data table are matched (similarity calculation) to obtain the matching result.

[0071] It should be noted that the preset data matching model is obtained by iterative training based on the language model to be trained. The main process in the model is to train the model by training the text feature training samples and the sample labels corresponding to the samples.

[0072] Among them, the language models to be trained include gpt4-turbo, chatGLM, LLaMa and other models. Among them, the model needs to be fed with basic knowledge related to architecture first.

[0073] Among them, when applying the data matching model, it mainly relies on AI big model technology, such as GPT, BERT, etc., to perform deep semantic understanding of historical data and user queries. These models obtain rich language understanding capabilities on large-scale text data through pre-training, and can more accurately capture the true intention behind the user's search terms and the delicate semantics in the historical data. In this embodiment, the AI ​​big model technology is used to mainly achieve two effects. The AI ​​big model technology is used to provide the user's historical search terms and the query records corresponding to the historical search terms, and the historical search terms and query records can be used to establish a mapping relationship between the data corresponding to the search terms and the query records. This mapping relationship mainly reflects the search terms and query records that are semantically associated (equivalent to semantic relationships), and the relationship is further embodied in a quantitative form. For example, the semantic relationship is classified and processed in a certain order to form a historical sample, and the current model is continuously iterated and optimized based on the historical sample.

[0074] Among them, the context perception ability of the AI ​​big model is used to conduct an in-depth analysis of the relationship between the query terms and historical data. Not only can synonyms and related words be identified, but the specific meaning and usage of words can also be judged according to the context, improving the accuracy of semantic matching, and achieving the level of understanding comparable to that of professional cost engineers. In this embodiment, corresponding positioning feature words (such as: important feature words such as cost, cost, and working hours) are preset, and context perception is performed on each paragraph of content. In the perception process, the positioning feature words in the context are pre-positioned, and the context linkage of the positioning feature words is combined to analyze the meaning represented by the data corresponding to the positioning feature words.

[0075] In addition, industry expertise and a large number of data sets related to construction cost are introduced during the model training phase to further enhance the model's ability to understand construction cost data through supervised learning or semi-supervised learning. Considering the particularity of the construction cost field, the AI ​​model can be customized to better understand industry terms, specifications, and common calculation methods.

[0076] Among them, through the enhanced learning mechanism, the model can continuously optimize and adjust the matching strategy in actual applications. The model parameters are adjusted by using user feedback and the correctness evaluation of query results, so that the semantic matching ability of the model can dynamically adapt to changes in user needs and continuously improve query accuracy.

[0077] In this embodiment, the step of querying and obtaining a matching result from a vector data table based on the text feature vector includes: calculating the cosine similarity between the text feature vector and each data in the vector data table; and taking the corresponding data whose cosine similarity is greater than a preset similarity as a matching result.

[0078] It is understandable that when calculating the similarity between the text feature vector and the data in the vector data table, the cosine similarity between the vectors is usually used for calculation, the cosine similarity between the two vectors is calculated, and the cosine similarity is compared with the preset similarity. When the cosine similarity is greater than the preset similarity, the data corresponding to the cosine similarity can be used as the matching result (high similarity).

[0079] It is assumed that the text feature vector and the corresponding vectors of the data in the vector data table are A = [a1, a2, ..., a n ] and B = [b1, b2, ..., b n ], the cosine similarity calculation formula of vector A and vector B is as follows:

[0080]

[0081] Where A·B represents the dot product (inner product) of vector A and vector B, and the calculation formula is: A·B=a1b1+a2b2+...+a n b n ;

[0082] Among them, |A| and |B| represent the modulus (or length) of vector A and vector B respectively, and the calculation formulas are as follows:

[0083]

[0084] This embodiment obtains the search term provided by the user, and determines the synonyms of the search term from the pre-equipment word library, and determines the associated words that have a mapping relationship with the search term from the pre-equipment word library; generates index items according to the search term, the synonyms and the associated words; generates a text feature vector corresponding to the index item through a preset data matching model, and obtains a matching result by querying from a vector data table according to the text feature vector; wherein the preset data matching model is obtained by iteratively training a language model to be trained based on text feature training samples and sample labels corresponding to the text feature training samples, and the vector data table is a feature vector data table obtained after the data matching model pre-performs text semantic analysis on the corresponding data in the construction cost list, that is, index items are generated by refinement, and matching results corresponding to the search term are found through the preset data matching model, so as to ensure the effect of accurate and rapid data query.

[0085] like Figure 2 As shown, based on the first embodiment, a second embodiment of the data query method based on text vectors of the present application is proposed. In this embodiment, step S20 specifically includes:

[0086] S21, combining the search term, the synonym and the associated word to obtain an initial data group, wherein the initial data is a data group pointed to by the search term or the synonym to the associated word;

[0087] It is understandable that when generating index items, corresponding index items can be obtained by combining search terms, synonyms and related terms. For example, search terms and related terms can be combined, or synonyms and related terms can be combined to obtain index formulas with a smaller index range, thereby obtaining an initial data group.

[0088] Among them, the corresponding index formula in the initial data group involves different synonyms and related words, and considering the particularity of the related words, the related words correspond to different text format associations and semantic associations, thereby determining that the index formula with related words has index content with less relevance to the content corresponding to the index words.

[0089] In order to reduce the subsequent conversion vectors and the matching time required for subsequent queries of corresponding data, the initial data group needs to be further screened to reduce the data groups in the initial data group that are irrelevant or have a low degree of relevance to the actual search terms.

[0090] S22, modifying the initial data group according to the mapping relationship between the search term / the synonym and the associated term to obtain a triple;

[0091] In this embodiment, the initial data set is screened and modified according to the mapping relationship between the search terms / synonyms and the associated terms to obtain triples.

[0092] When modifying the initial data group, the pointing information between the synonyms / search terms and the associated terms is mainly established through the mapping relationship, so that the triples are constructed through the triples to ensure the integrity of the triple data.

[0093] When the initial data is screened, the similarity between the search terms / synonyms and the associated terms in the initial data group is mainly determined based on the mapping relationship, and a suitable data group is screened out from the initial data group based on the similarity.

[0094] In this embodiment, the step of modifying the initial data set according to the mapping relationship between the search term / the synonym and the associated term to obtain a triple includes two methods, which are as follows:

[0095] One of them is: when the mapping relationship between the search term / the synonym and the associated term is a text format unified relationship, determining whether there is a query record corresponding to the search term in the associated term; if so, modifying the initial data group according to the definition formula corresponding to the query record to obtain a triple;

[0096] It is understandable that when the mapping relationship is a unified text format relationship (that is, the text format relationship in the above content), it can be judged that the number of data in the current associated words that are the same as the search word content is small, and most of the associated words are data with consistent text formats but inconsistent actual text contents. Therefore, such associated words need to be screened.

[0097] Specifically, in this embodiment, the process of screening the associated words is mainly through querying the historical index records, and through the records, determining the corresponding definition formula, thereby modifying the initial data group according to the definition formula, screening out some data groups that are irrelevant to the definition formula or have a low degree of association, and then obtaining triples similar to the search terms entered by the user.

[0098] The query record includes the corresponding index item, index formula or definition formula, that is, the queried data results corresponding to the above three styles.

[0099] Another one is: when the mapping relationship between the search term / the synonym and the associated term is a content-associated relationship, the initial data group is analyzed for a degree of association to obtain an analysis result; based on the analysis result, a data group whose degree of association is greater than a preset threshold is screened out from the initial data group, and based on the content-associated relationship, the screened data group is modified to obtain a triple.

[0100] It is understandable that the initial data group contains content that has a low correlation with the search terms. For example, the user wants to search for monetary cost, but the associated word corresponding to the initial data group may be time cost. There is a certain correlation between time cost and monetary cost. Both are cost-related data, but they are still very different in nature. Therefore, in order to avoid the influence of subsequent retrieval of data with low relevance to the search terms on the queried matching results, it is necessary to filter the associated words in the initial data, where the filtering action is mainly analyzed by judging the degree of correlation between the associated words and the search terms.

[0101] Among them, the content-related association relationship mainly refers to the semantic relationship involved in the above content.

[0102] Among them, the correlation degree relationship corresponding to the initial data group is analyzed, mainly analyzing the correlation degree between the associated words and the search words / synonyms. If the correlation degree is high, it will be retained, and if the correlation degree is low, it will be removed, that is, the data group with a correlation degree greater than a preset threshold is screened out from the initial data group, and the data group is retained. Through the content correlation relationship, the initial data group is modified to obtain a triple.

[0103] S23, analyzing the text semantics of the triples, and generating index items of different text contents according to the text semantics and a preset template.

[0104] In this embodiment, after obtaining the triples, it is necessary to generate corresponding index items according to the triples. The index items mainly refer to different index types formed by different qualifiers and / or triples.

[0105] When generating the index formula, the text semantics of the triples are mainly judged, and the index items of different text contents are determined by combining the text semantics and the preset template.

[0106] The preset template mainly refers to a template in an indexed format.

[0107] Among them, the index items of different text contents refer to the fact that the index scope corresponding to the index items has a certain tendency. Taking the index word "bridge construction cost" as an example, its index item tendency can be divided into: labor cost, material cost, time planning cost, etc. However, the main content of the index item needs to be consistent with the corresponding content of the text semantics corresponding to the triple.

[0108] This embodiment obtains an initial data group by combining the search term, the synonym and the associated word, wherein the initial data is a data group pointed to by the search term or the synonym to the associated word, and the initial data group is modified according to the mapping relationship between the search term / the synonym and the associated word to obtain a triple, so that the text semantics of the triple can be analyzed, and index items of different text contents can be generated according to the text semantics and a preset template, that is, index items of different text content tendencies are generated by the effect of the combination between the search term, the synonym and the associated word, thereby ensuring a wide range of indexing.

[0109] like Figure 3 As shown, based on the first embodiment, a third embodiment of the data query method based on text vectors of the present application is proposed. In this embodiment, step S30 specifically includes:

[0110] S31, performing semantic recognition on the data content corresponding to the index item;

[0111] S32, extracting multi-level feature information corresponding to the index item according to the result of semantic recognition, wherein the multi-level feature information is feature information corresponding to words in the index item corresponding to different index ranges, and the level to which each feature information belongs in the multi-level feature information is determined according to the size of the index range;

[0112] S33, generating a text feature vector corresponding to the index item according to the multi-level feature information through the data matching model, wherein the text feature vector includes a sub-vector corresponding to each level feature in the multi-level feature information.

[0113] It is understandable that as index conditions for querying data in the list, the corresponding high accuracy and wide coverage of the index items can ensure that the final query results are closer to the user's search intention, which is equivalent to ensuring the search accuracy and efficiency.

[0114] In this embodiment, in order to ensure that the index item can query more data, it is necessary to construct a specific text feature vector based on the content corresponding to the index item, and through the text feature vector, with a wider range and strong logical characteristics, query the appropriate results from the large amount of data corresponding to the list.

[0115] Specifically, the concept of multi-level feature information is proposed in this embodiment, that is, when converting index items into text feature vectors, the content of semantic recognition corresponding to the index items is preferentially converted into multi-level preferential total energy information, and the multi-level feature information is passed through a data matching model to generate a corresponding multi-level feature text feature vector.

[0116] It should be noted that the multi-level feature information corresponds to feature information corresponding to words in different index ranges, wherein the index range corresponding to each level is different in size, and different feature information is clustered to obtain multi-level feature information.

[0117] Among them, taking the feature information in which the multi-level feature information is divided into three levels of A, B and C as an example, the A-level feature information corresponds to the level with the largest index range, and B and C gradually reduce the index range. The index range refers to the amount of data that can be indexed by feature information of a certain level. The larger the index range, the larger the amount of data that can be indexed.

[0118] Specifically, the number of corresponding index elements in level A should be less than the number of index elements in level B.

[0119] Specifically, the words corresponding to the corresponding index elements in level A have a higher degree of superiority than those in other levels. For example, the index item in level A may be "metal", and the index items in level B may be "metal or steel or copper" and "cost", etc.

[0120] In addition, after obtaining the multi-level feature information, it is necessary to perform vector conversion on the multi-level feature information through a data matching model. During the conversion process, it is considered that there are certain differences in the meanings represented by the feature information of each level. Therefore, when converting the feature vector, a corresponding vector group set will be converted. The vector group set includes the vectors corresponding to the index items that have not divided the multi-level feature information, and the sub-vectors corresponding to each level of feature information after the multi-level feature information is divided.

[0121] It should be noted that when converting to obtain a sub-vector, the relationship information between the sub-vector and the vector corresponding to the index item is also converted, as well as the relationship information between each sub-vector, thereby further improving the dimension and richness of the converted vector, thereby ensuring the accuracy of the matching result obtained by the final query.

[0122] Specifically, in this embodiment, the step of generating a text feature vector corresponding to the index item according to the multi-level feature information through the data matching model includes: calculating a feature pointing vector between feature information of each level according to the multi-level feature information through the data matching model; and generating a text feature vector corresponding to the index item composed of the multi-level feature information according to the feature pointing vector.

[0123] In this embodiment, before converting the multi-level feature information into a sub-vector corresponding to the multi-level feature information through the data matching model, the feature pointing vector between the feature information of each level will be calculated, which is equivalent to constructing a sub-triplet corresponding to the feature information from the previous level to the next level through the feature pointing vector. Each level of feature information can also be connected through the feature pointing vector to construct another sub-triplet consisting of the first-level feature information, the last-level feature information and the feature pointing vector. Specifically, for example, the feature pointing vectors corresponding to level A and level B and between level A and level B constitute a sub-triplet, and the feature pointing vectors corresponding to level A and level C and between level A and level C constitute a sub-triplet.

[0124] In summary, in this embodiment, the text feature vector converted by the data matching model includes the overall feature vector directly converted from the index item, as well as the sub-vectors corresponding to the feature information of each level, and the sub-vectors corresponding to the sub-triplets composed of the corresponding features of each level, and a corresponding relationship between the overall feature vector and each sub-vector is established, so as to ensure that the subsequent vector similarity matching can accurately obtain the matching result closest to the index term.

[0125] This embodiment performs semantic recognition on the data content corresponding to the index item, and extracts multi-level feature information corresponding to the index item based on the result of the semantic recognition, wherein the multi-level feature information is feature information corresponding to words in the index item corresponding to different index ranges, and the level to which each feature information belongs is determined according to the size of the index range in the multi-level feature information, so that a text feature vector corresponding to the index item can be generated according to the multi-level feature information through the data matching model, wherein the text feature vector includes a sub-vector corresponding to each level feature in the multi-level feature information, that is, by converting the index item into a vector and converting it according to the multi-level features, vector enrichment is achieved, thereby ensuring that a wider index range can be obtained in the subsequent vector similarity comparison, and the matching result can be accurately locked according to the multi-level feature content.

[0126] In addition, the present application also proposes a data query device based on text vectors, referring to Figure 3 , the data query device based on text vector includes:

[0127] The acquisition module 10 is used to acquire the search term provided by the user, and determine the synonyms of the search term from the pre-equipped word library, and determine the associated words that have a mapping relationship with the search term from the pre-equipped word library;

[0128] A generating module 20, for generating index items according to the search terms, the synonyms and the associated terms;

[0129] The query module 30 is used to generate a text feature vector corresponding to the index item through a preset data matching model, and obtain a matching result from a vector data table based on the text feature vector; wherein the preset data matching model is obtained by iteratively training a language model to be trained based on text feature training samples and sample labels corresponding to the text feature training samples, and the vector data table is a feature vector data table obtained after the data matching model pre-performs text semantic analysis on the corresponding data in the construction cost list.

[0130] This embodiment obtains the search term provided by the user, and determines the synonyms of the search term from the pre-equipment word library, and determines the associated words that have a mapping relationship with the search term from the pre-equipment word library; generates index items according to the search term, the synonyms and the associated words; generates a text feature vector corresponding to the index item through a preset data matching model, and obtains a matching result by querying from a vector data table according to the text feature vector; wherein the preset data matching model is obtained by iteratively training a language model to be trained based on text feature training samples and sample labels corresponding to the text feature training samples, and the vector data table is a feature vector data table obtained after the data matching model pre-performs text semantic analysis on the corresponding data in the construction cost list, that is, index items are generated by refinement, and matching results corresponding to the search term are found through the preset data matching model, so as to ensure the effect of accurate and rapid data query.

[0131] It should be noted that each module in the above-mentioned device can be used to implement each step in the above-mentioned method and achieve corresponding technical effects, which will not be described in detail in this embodiment.

[0132] Reference Figure 4 , Figure 4 A schematic diagram of the structure of the device of the hardware operating environment involved in the embodiment of the present application.

[0133] like Figure 4As shown, the device may include: a processor 1001, such as a CPU, a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard), and the optional user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory, or it may be a stable memory (non-volatile memory), such as a disk memory. The memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0134] Those skilled in the art will understand that Figure 4 The structure shown in the figure does not constitute a limitation of the device, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.

[0135] like Figure 4 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module, and a data query program based on text vectors.

[0136] exist Figure 4 In the device shown, the network interface 1004 is mainly used for data communication with an external network; the user interface 1003 is mainly used for receiving user input instructions; the device calls the text vector-based data query program stored in the memory 1005 through the processor 1001, and performs the following operations:

[0137] Acquire a search term provided by a user, and determine a synonym of the search term from a pre-equipped word library, and determine an associated term that has a mapping relationship with the search term from the pre-equipped word library;

[0138] Generate index items according to the search terms, the synonyms and the associated terms;

[0139] Generate a text feature vector corresponding to the index item through a preset data matching model, and query the vector data table to obtain a matching result based on the text feature vector;

[0140] Among them, the preset data matching model is obtained by iteratively training the language model to be trained based on text feature training samples and sample labels corresponding to the text feature training samples, and the vector data table is a feature vector data table obtained after the data matching model pre-performs text semantic analysis on the corresponding data in the construction cost list.

[0141] Further, the processor 1001 may call a text vector-based data query program stored in the memory 1005, and further perform the following operations:

[0142] Combining the search term, the synonym and the associated word to obtain an initial data group, wherein the initial data is a data group pointed to by the search term or the synonym to the associated word;

[0143] According to the mapping relationship between the search term / the synonym and the associated term, the initial data group is modified to obtain a triple;

[0144] The text semantics of the triples are analyzed, and index items of different text contents are generated according to the text semantics and a preset template.

[0145] Further, the processor 1001 may call a text vector-based data query program stored in the memory 1005, and further perform the following operations:

[0146] When the mapping relationship between the search term / the synonym and the associated term is a text format unified relationship, determining whether there is a query record corresponding to the search term in the associated term;

[0147] If it exists, modify the initial data group according to the definition formula corresponding to the query record to obtain a triple;

[0148] or,

[0149] When the mapping relationship between the search term / the synonym and the associated term is a content-associated relationship, a correlation degree analysis is performed on the initial data group to obtain an analysis result;

[0150] According to the analysis result, a data group with a correlation degree greater than a preset threshold is screened out from the initial data group, and the screened data group is modified according to the content correlation relationship to obtain a triple.

[0151] Further, the processor 1001 may call a text vector-based data query program stored in the memory 1005, and further perform the following operations:

[0152] Performing semantic recognition on the data content corresponding to the index item;

[0153] Extracting multi-level feature information corresponding to the index item according to the result of semantic recognition, wherein the multi-level feature information is feature information corresponding to words in the index item corresponding to different index ranges, and the level to which each feature information belongs in the multi-level feature information is determined according to the size of the index range;

[0154] The text feature vector corresponding to the index item is generated according to the multi-level feature information through the data matching model, wherein the text feature vector includes a sub-vector corresponding to each level feature in the multi-level feature information.

[0155] Further, the processor 1001 may call a text vector-based data query program stored in the memory 1005, and further perform the following operations:

[0156] By using the data matching model, according to the multi-level feature information, a feature pointing vector between feature information of each level is calculated;

[0157] A text feature vector corresponding to the index item and composed of the multi-level feature information is generated according to the feature pointing vector.

[0158] Further, the processor 1001 may call a text vector-based data query program stored in the memory 1005, and further perform the following operations:

[0159] Calculate the cosine similarity between the text feature vector and each data in the vector data table;

[0160] The corresponding data whose cosine similarity is greater than the preset similarity is taken as the matching result.

[0161] Further, the processor 1001 may call a text vector-based data query program stored in the memory 1005, and further perform the following operations:

[0162] Obtaining list data corresponding to the building cost list, and classifying the list data to obtain a first data table with a unified text format and a second data table with related text content;

[0163] Extracting characteristic words that meet preset conditions according to the text content of the second data table, and determining replacement words with the same semantics as the characteristic words;

[0164] The second data table is updated according to the replacement words and the characteristic words, and a pre-device selected word library is constructed according to the updated second data table and the first data table.

[0165] This embodiment obtains the search term provided by the user, and determines the synonyms of the search term from the pre-equipment word library, and determines the associated words that have a mapping relationship with the search term from the pre-equipment word library; generates index items according to the search term, the synonyms and the associated words; generates a text feature vector corresponding to the index item through a preset data matching model, and obtains a matching result by querying from a vector data table according to the text feature vector; wherein the preset data matching model is obtained by iteratively training a language model to be trained based on text feature training samples and sample labels corresponding to the text feature training samples, and the vector data table is a feature vector data table obtained after the data matching model pre-performs text semantic analysis on the corresponding data in the construction cost list, that is, index items are generated by refinement, and matching results corresponding to the search term are found through the preset data matching model, so as to ensure the effect of accurate and rapid data query.

[0166] In addition, an embodiment of the present application further proposes a computer-readable storage medium, on which a text vector-based data query program is stored. When the text vector-based data query program is executed by a processor, the following operations are implemented:

[0167] Acquire a search term provided by a user, and determine a synonym of the search term from a pre-equipped word library, and determine an associated term that has a mapping relationship with the search term from the pre-equipped word library;

[0168] Generate index items according to the search terms, the synonyms and the associated terms;

[0169] Generate a text feature vector corresponding to the index item through a preset data matching model, and query the vector data table to obtain a matching result based on the text feature vector;

[0170] Among them, the preset data matching model is obtained by iteratively training the language model to be trained based on text feature training samples and sample labels corresponding to the text feature training samples, and the vector data table is a feature vector data table obtained after the data matching model pre-performs text semantic analysis on the corresponding data in the construction cost list.

[0171] This embodiment obtains the search term provided by the user, and determines the synonyms of the search term from the pre-equipment word library, and determines the associated words that have a mapping relationship with the search term from the pre-equipment word library; generates index items according to the search term, the synonyms and the associated words; generates a text feature vector corresponding to the index item through a preset data matching model, and obtains a matching result by querying from a vector data table according to the text feature vector; wherein the preset data matching model is obtained by iteratively training a language model to be trained based on text feature training samples and sample labels corresponding to the text feature training samples, and the vector data table is a feature vector data table obtained after the data matching model pre-performs text semantic analysis on the corresponding data in the construction cost list, that is, index items are generated by refinement, and matching results corresponding to the search term are found through the preset data matching model, so as to ensure the effect of accurate and rapid data query.

[0172] It should be noted that when the above-mentioned computer-readable storage medium is executed by a processor, it can also implement the various steps in the above-mentioned method and achieve the corresponding technical effects, which will not be described in detail in this embodiment.

[0173] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.

[0174] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0175] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0176] The above are only preferred embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A data query method based on text vectors, characterized in that: The data query method based on text vectors comprises the following steps: Acquire a search term provided by a user, and determine a synonym of the search term from a pre-equipped word library, and determine an associated term that has a mapping relationship with the search term from the pre-equipped word library; Generate index items according to the search terms, the synonyms and the associated terms; Generate a text feature vector corresponding to the index item through a preset data matching model, and query a matching result from a vector data table based on the text feature vector, wherein the preset data matching model is obtained by iteratively training a language model to be trained based on text feature training samples and sample labels corresponding to the text feature training samples, and the vector data table is a feature vector data table obtained after the data matching model pre-performs text semantic analysis on the corresponding data in the building cost list; The step of generating the text feature vector corresponding to the index item through a preset data matching model includes: Performing semantic recognition on the data content corresponding to the index item; According to the result of semantic recognition, the multi-level feature information corresponding to the index item is extracted, wherein the multi-level feature information is feature information corresponding to words in the index item corresponding to different index ranges, and the level to which each feature information belongs is determined according to the size of the index range in the multi-level feature information, and the index range is divided according to the amount of data indexed by the hierarchical feature information; Generate a text feature vector corresponding to the index item according to the multi-level feature information through the data matching model, wherein the text feature vector includes a sub-vector corresponding to each level feature in the multi-level feature information; The step of generating the text feature vector corresponding to the index item according to the multi-level feature information through the data matching model includes: By using the data matching model, according to the multi-level feature information, a feature pointing vector between feature information of each level is calculated; A text feature vector corresponding to the index item and composed of the multi-level feature information is generated according to the feature pointing vector.

2. The data query method based on text vector according to claim 1, characterized in that: The step of generating index items according to the search terms, the synonyms and the associated terms comprises: Combining the search term, the synonym and the associated word to obtain an initial data group, wherein the initial data group is a data group pointed to by the search term or the synonym to the associated word; Modifying the initial data set according to the mapping relationship between the search term or the synonym and the associated term to obtain a triple; The text semantics of the triples are analyzed, and index items of different text contents are generated according to the text semantics and a preset template.

3. The data query method based on text vector according to claim 2, characterized in that: The step of modifying the initial data set according to the mapping relationship between the search term or the synonym and the associated term to obtain a triple comprises: When the mapping relationship between the search term or the synonym and the associated term is a text format unified relationship, determining whether there is a query record corresponding to the search term in the associated term; If it exists, modify the initial data group according to the definition formula corresponding to the query record to obtain a triple; or, When the mapping relationship between the search term or the synonym and the associated term is a content-associated relationship, a correlation degree analysis is performed on the initial data group to obtain an analysis result; According to the analysis result, a data group with a correlation degree greater than a preset threshold is screened out from the initial data group, and the screened data group is modified according to the content correlation relationship to obtain a triple.

4. The data query method based on text vector according to claim 1, characterized in that: The step of querying and obtaining a matching result from a vector data table according to the text feature vector comprises: Calculate the cosine similarity between the text feature vector and each data in the vector data table; The corresponding data whose cosine similarity is greater than the preset similarity is taken as the matching result.

5. The data query method based on text vector according to claim 1, characterized in that: Before the steps of obtaining the search term provided by the user, determining the synonyms of the search term from the pre-equipped word library, and determining the associated words that have a mapping relationship with the search term from the pre-equipped word library, the method further includes: Obtaining list data corresponding to the building cost list, and classifying the list data to obtain a first data table with a unified text format and a second data table with related text content; Extracting characteristic words that meet preset conditions according to the text content of the second data table, and determining replacement words with the same semantics as the characteristic words; The second data table is updated according to the replacement words and the characteristic words, and a pre-device selected word library is constructed according to the updated second data table and the first data table.

6. A data query device based on text vector, characterized in that: The data query device based on text vectors comprises: An acquisition module, used to acquire a search term provided by a user, and determine a synonym of the search term from a pre-equipped word library, and determine an associated term that has a mapping relationship with the search term from the pre-equipped word library; A generating module, used for generating index items according to the search terms, the synonyms and the associated terms; A query module is used to generate a text feature vector corresponding to the index item through a preset data matching model, and query a matching result from a vector data table based on the text feature vector; wherein the preset data matching model is obtained by iteratively training a language model to be trained based on text feature training samples and sample labels corresponding to the text feature training samples, and the vector data table is a feature vector data table obtained after the data matching model pre-performs text semantic analysis on the corresponding data in the construction cost list; The query module is further used to perform semantic recognition on the data content corresponding to the index item; extract the multi-level feature information corresponding to the index item according to the result of the semantic recognition, wherein the multi-level feature information is the feature information corresponding to the words corresponding to different index ranges in the index item, and the level to which each feature information belongs is determined according to the size of the index range in the multi-level feature information, and the index range is divided according to the amount of data indexed by the hierarchical feature information; generate the text feature vector corresponding to the index item according to the multi-level feature information through the data matching model, wherein the text feature vector includes a sub-vector corresponding to each hierarchical feature in the multi-level feature information; The query module is also used to calculate the feature pointing vectors between the feature information of each level according to the multi-level feature information through the data matching model; and generate the text feature vector composed of the multi-level feature information corresponding to the index item according to the feature pointing vector.

7. A data query device based on text vectors, characterized in that: The text vector-based data query device includes: a memory, a processor, and a text vector-based data query program stored in the memory and running on the processor, and the text vector-based data query program is configured to implement the steps of the text vector-based data query method as described in any one of claims 1 to 5.

8. A storage medium, characterized in that: A program for implementing a data query method based on a text vector is stored on a storage medium, and the program for implementing a data query method based on a text vector is executed by a processor to implement the steps of the data query method based on a text vector as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Video program searching method and apparatus

    CN106570196A

  • Intelligent semantic retrieval method and system and electronic equipment

    CN112035598A

  • Voice data search device

    JP2006040150A