A data query method and device, electronic equipment and storage medium
By filtering technical terms with high relevance to the query terms in the database, and combining the relevance and matching degree of enterprises in different business dimensions, the system automatically filters out target enterprises, solving the problems of large workload and inaccurate judgment caused by manual retrieval in existing technologies, and realizing efficient and accurate enterprise technology query.
Patent Information
- Application Number
- CN202210436608.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-25
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2042-04-25
AI Technical Summary
When searching for information on a company's technological development in a database, existing technologies require manually entering the company name and searching one by one, resulting in a large workload and inaccurate subjective judgment.
By extracting the keywords to be queried from the query terminal, filtering out technical terms whose relevance exceeds a threshold, determining the relevance and matching degree of enterprises under different business dimensions from the database, and combining the weights to calculate the third relevance, the target enterprises are automatically filtered out.
It expands the scope of queries, improves the accuracy of queries on enterprise technology development, reduces manual workload, and minimizes errors from subjective judgment.
Smart Images

Figure CN114780601B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data query, in particular to a data query method and device, an electronic device and a storage medium. BACKGROUND
[0002] With the continuous development of informatization, the speed of data generation is increasing, the amount of data to be processed is rapidly expanding, and the big data era is coming. Big data refers to the amount of data involved, which is so large that it cannot be processed within a reasonable time by mainstream software. When facing massive data, how to quickly find the desired data in the database has become a problem to be solved.
[0003] When querying data in the database, it is usually matched to the corresponding data by word-by-word comparison. Therefore, when searching for an enterprise name in a database containing information of multiple enterprises and companies, it is necessary to input the part of the text that coincides with the enterprise name, so that the target enterprise can be found according to the principle of word-by-word matching.
[0004] The inventor found in the research that in the prior art, if you want to know the development of a certain technology in various enterprises, you need to obtain the enterprise names of all enterprises related to the technology, input the enterprise names one by one in the search box, obtain the enterprise information, and subjectively judge the reserve situation of each enterprise for the technology, which is a large amount of manual work. SUMMARY
[0005] Therefore, the embodiments of the present application provide a data query method and device, an electronic device and a storage medium, which can help to solve the problem of large amount of manual work caused by the need to manually search for the technical situation of each enterprise.
[0006] In a first aspect, the embodiments of the present application provide a data query method, which comprises:
[0007] extracting a query word from the query content sent by the query terminal;
[0008] screening at least one technical word of the query word from the database; the first correlation degree between the technical word and the query word exceeds a first threshold value; the technical word is a pre-recorded word used to describe technical means;
[0009] For each preset business dimension, based on the plurality of first enterprises stored in the database, the business data of each first enterprise in the business dimension, the second correlation degree between the business data and the query word, and the matching degree between the business data and the technical word, at least one second enterprise is determined from the plurality of first enterprises; the business dimension includes an operation dimension, a technology dimension, a product dimension, and a trademark dimension;
[0010] based on the first correlation degree, the second correlation degree, the matching degree, and a weight pre-set for each of the business dimensions, calculating a third correlation degree of each of the second enterprises and the content to be queried;
[0011] sending a query result including each of the target enterprises and the third correlation degree of the target enterprises to the query terminal; the target enterprise is the second enterprise whose third correlation degree exceeds a second threshold.
[0012] In a feasible implementation, for each of the preset business dimensions, based on a plurality of first enterprises stored in the database, business data of each of the first enterprises in the business dimension, a second correlation degree of the business data and the query word, and a matching degree of the business data and the technical word, at least one second enterprise is determined from the plurality of first enterprises, including:
[0013] for each of the business dimensions, based on a business dimension label of each of the first enterprises stored in the database, filtering the first enterprises including the business dimension from the database; each of the business dimension labels corresponds to one of the business dimensions;
[0014] for each of the filtered first enterprises, extracting business data belonging to the business dimension from the database; the business dimension to which the business data belongs is pre-labeled;
[0015] extracting first data with a semantic correlation degree exceeding a third threshold from the business data and the query word;
[0016] for each of the first data, if a matching degree of the first data and a target technical word exceeds a fourth threshold, determining the first data as second data; the target technical word is at least one of the technical words;
[0017] based on the second data of each of the first enterprises, according to a data amount of the second data, a matching degree of each of the second data, and a semantic correlation degree of each of the second data, determining at least one second enterprise from the plurality of first enterprises.
[0018] In a feasible implementation, the labeling method of the business dimensions includes:
[0019] based on the enterprise data of the first enterprises stored in the database, for each of the first enterprises, extracting at least one enterprise feature from the enterprise data by an entity recognition algorithm; the enterprise data includes business data of the first enterprise in each of the business dimensions; the enterprise feature is used to describe the attributes of the first enterprise;
[0020] determine a target business dimension to which each of the enterprise features belongs by using a pre-trained dimension labeling model, and label each of the enterprise features with a first label corresponding to each of the business dimensions in the target business dimension; the target business dimension is at least one of the business dimensions;
[0021] Based on the first label labeled for each of the enterprise features of the first enterprise and the target business dimension corresponding to the first label, count the business dimension label of the first enterprise.
[0022] In a feasible implementation, before labeling each of the enterprise features with the first label corresponding to the target business dimension uniquely, the method further comprises:
[0023] clean the enterprise features by using an entity alignment method and an attribute alignment method;
[0024] identify the semantics of each of the cleaned enterprise features by using a semantic recognition algorithm, and label each of the enterprise features with a second label based on the semantics; the second label includes a technology label and an attribute label;
[0025] generate an enterprise portrait for each of the first enterprises based on the second label and the enterprise features.
[0026] In a feasible implementation, filtering at least one technology word of the to-be-queried vocabulary from a database comprises:
[0027] Based on the to-be-queried vocabulary, searching for a target graph of the to-be-queried vocabulary in the database;
[0028] extracting at least one technology word having a target relationship with the to-be-queried vocabulary from the target graph; the target relationship includes a subordinate relationship and an application relationship.
[0029] In a feasible implementation, before sending the query result to the query terminal, the method further comprises:
[0030] extracting a third relevance of each of the target enterprises from the query result;
[0031] based on the numerical value of the third relevance, sorting the target enterprises in the query result to obtain an enterprise list containing a sorting result;
[0032] storing the enterprise list in the query result.
[0033] In a second aspect, the embodiments of the present application further provide a data query device, and the device comprises:
[0034] a first extraction unit configured to extract a to-be-queried vocabulary from to-be-queried content sent by a query terminal;
[0035] a screening unit configured to screen at least one technical term from the database, the technical term having a first correlation degree with the query term exceeding a first threshold, and the technical term being a pre-recorded term for describing technical means;
[0036] a determining unit configured to determine at least one second enterprise from the plurality of first enterprises for each business dimension based on the plurality of first enterprises stored in the database, business data of each first enterprise in the business dimension, a second correlation degree of the business data with the query term, and a matching degree of the business data with the technical term, wherein the business dimension comprises an operation dimension, a technology dimension, a product dimension, and a trademark dimension;
[0037] a calculating unit configured to calculate a third correlation degree of each second enterprise with the query content based on the first correlation degree, the second correlation degree, the matching degree, and a weight pre-set for each business dimension;
[0038] a sending unit configured to send a query result comprising each target enterprise and the third correlation degree of the target enterprise to the query terminal, wherein the target enterprise is the second enterprise having the third correlation degree exceeding a second threshold.
[0039] In an available embodiment, the determining unit is configured to:
[0040] for each business dimension, screen the first enterprise comprising the business dimension from the database based on the business dimension label of each first enterprise stored in the database, wherein each business dimension label corresponds to one business dimension;
[0041] for each screened first enterprise, extract the business data belonging to the business dimension from the database, wherein the business dimension to which the business data belongs is pre-labeled;
[0042] extract first data having a semantic correlation degree with the query term exceeding a third threshold from the business data;
[0043] for each first data, if a matching degree of the first data with a target technical term exceeds a fourth threshold, determine the first data as second data, wherein the target technical term is at least one of the technical terms;
[0044] determine at least one second enterprise from the plurality of first enterprises based on the second data of each first enterprise, a data amount of the second data, the matching degree of each second data, and the semantic correlation degree of each second data.
[0045] In an implementable embodiment, the device further comprises:
[0046] A first identifying unit is configured to, when labeling the business dimension label, extract at least one enterprise feature from the enterprise data of each of the first enterprises by an entity recognition algorithm based on the enterprise data of the first enterprises stored in the database; the enterprise data comprises the business data of the first enterprises under each of the business dimensions; and the enterprise feature is used to describe the attributes of the first enterprises;
[0047] A first labeling unit is configured to determine a target business dimension to which each of the enterprise features belongs by a pre-trained dimension labeling model, and label the first label corresponding to each of the business dimensions in the target business dimension for the enterprise feature; the target business dimension is at least one of the business dimensions;
[0048] A statistical unit is configured to count the business dimension label of the first enterprises based on the first label labeled for each of the enterprise features of the first enterprises and the target business dimension corresponding to the first label.
[0049] In an implementable embodiment, the device further comprises:
[0050] A cleaning unit is configured to clean the enterprise feature by an entity alignment method and an attribute alignment method before labeling the first label corresponding to the target business dimension for the enterprise feature;
[0051] A second identifying unit is configured to identify the semantics of each of the cleaned enterprise features by a semantic recognition algorithm, and label the second label for the enterprise feature based on the semantics; the second label comprises a technology label and an attribute label.
[0052] A portrait generating unit is configured to generate the enterprise portrait for each of the first enterprises based on the second label and the enterprise feature.
[0053] In an implementable embodiment, the screening unit is configured to:
[0054] Based on the to-be-queried vocabulary, find the target graph of the to-be-queried vocabulary in the database;
[0055] Extract at least one technology word having a target relationship with the to-be-queried vocabulary from the target graph; the target relationship comprises a subordinate relationship and an application relationship.
[0056] In an implementable embodiment, the device further comprises:
[0057] a second extraction unit configured to extract a third relevance of each target enterprise from the query result before sending the query result to the query terminal;
[0058] a sorting unit configured to sort the target enterprises in the query result based on the numerical value of the third relevance to obtain an enterprise list containing a sorting result;
[0059] a storage unit configured to store the enterprise list into the query result.
[0060] In a third aspect, an electronic device is provided, which includes a processor, a storage medium and a bus. The storage medium stores machine readable instructions executable by the processor. When the electronic device is running, the processor communicates with the storage medium through the bus. The processor executes the machine readable instructions to perform the steps of the method according to any one of the first aspect.
[0061] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program. When the computer program is run by a processor, the steps of the method according to any one of the first aspect are performed.
[0062] The data query method, device, electronic device and storage medium provided by the embodiments of the present application can form an expanded word set of the to-be-queried word by screening at least one technical word with a first relevance to the to-be-queried word exceeding a first threshold value from a database, and query the target enterprise by the to-be-queried word and the expanded word set, thereby avoiding the problem of too narrow search range caused by querying only according to the content input by the user. The target enterprise is determined from the first enterprise by determining the second relevance of the business data of the first enterprise to the to-be-queried word, the matching degree of the business data to the technical word, the first relevance of the to-be-queried word to the technical word and the weight of each business dimension. Compared with the prior art which requires the user to manually search each enterprise name and retrieve the situation of each enterprise, the embodiments of the present application expand the query range by the above method, improve the accuracy of querying the technical development situation of each enterprise by combining the technical word in the case of limiting the first relevance and ensuring the matching degree of the first enterprise to the technical word, and help to solve the problems of large manual workload and inaccurate subjective judgment of the technical development situation caused by the need to manually retrieve the technical situation of each enterprise.
[0063] In order to make the above objectives, characteristics and advantages of the present application more apparent, more comprehensible and more clear, the following preferred embodiments are specifically described below with reference to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0064] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some of the embodiments of the present application, and therefore should not be regarded as limiting the scope, and for those skilled in the art, other related drawings can also be obtained without creative labor.
[0065] Figure 1 A flow chart of a data query method provided by an embodiment of the present application is shown.
[0066] Figure 2 A flow chart of a screening method provided by an embodiment of the present application is shown.
[0067] Figure 3 A structural schematic diagram of a data query device provided by an embodiment of the present application is shown.
[0068] Figure 4 A structural schematic diagram of an electronic device provided by an embodiment of the present application is shown. DETAILED DESCRIPTION
[0069] In order to make the purpose, technical solutions and advantages of the embodiments of the present application more clear, the following will combine the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. It should be understood that the drawings in the present application only play the purpose of illustration and description, and do not limit the protection scope of the present application. In addition, it should be understood that the schematic drawings are not drawn according to the actual proportion. The flow chart shows the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flow chart can not be implemented in sequence, and the steps without logical context relationship can be reversed in sequence or implemented simultaneously. In addition, one or more other operations can be added to the flow chart or removed from the flow chart under the guidance of the content of the present application by those skilled in the art.
[0070] In addition, the described embodiments are only some of the embodiments of the present application, not all the embodiments. The components of the embodiments of the present application described and shown in the drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0071] It should be noted in advance that the term "comprising" will be used in the embodiments of the present application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.
[0072] It needs to be pointed out in advance that the device or electronic equipment involved in the embodiments of the present application can be executed on a single server, or can be executed on a server group. The server group can be centralized or distributed. In some embodiments, the server can be local or remote relative to the terminal. For example, the server can access information and / or data stored in the service requester terminal, the service provider terminal, or the database, or any combination thereof via a network. As another example, the server can be directly connected to at least one of the service requester terminal, the service provider terminal, and the database to access the stored information and / or data. In some embodiments, the server can be implemented on a cloud platform; for example, the cloud platform can include a private cloud, a public cloud, a hybrid cloud, a community cloud, a distributed cloud, an inter-cloud, a multi-cloud, etc., or any combination thereof.
[0073] Figure 1 A flowchart of a data query method provided by the embodiments of the present application is shown in FIG. 1. As shown in FIG. 1, the method includes the following steps: Figure 1
[0074] Step 101: extracting a query word from the query content sent by the query terminal.
[0075] Specifically, the query terminal is any terminal device, including but not limited to a mobile phone, a tablet, a computer, etc. The query terminal provides a graphical user interface for inputting the query content. After the user inputs the query content through the graphical user interface, the query terminal sends the query content to the server, and the server processes the query content and returns the query result.
[0076] The query content includes text content (including keywords, short texts, long texts, etc.) input by the user. The text content can be manually input by the user, can be obtained by converting audio / video input by the user, or can be obtained by the user clicking a certain text. The query content includes but is not limited to the name of an enterprise that the user wants to query, a technical method that the user wants to query, the name of a product that the user wants to query, etc. The method for extracting the query word includes but is not limited to entity recognition, semantic recognition, etc. The query word is obtained by performing semantic analysis and sentence structure division on the text input by the user. There can be one or more query words.
[0077] Step 102: screening at least one technical word of the query word from the database; the first correlation degree between the technical word and the query word exceeds a first threshold value; and the technical word is a pre-recorded word used to describe a technical method.
[0078] Specifically, all the thresholds mentioned in the embodiments of the present application are preset, and the thresholds include a first threshold, a second threshold, a third threshold, and a fourth threshold. The thresholds can be adjusted according to the needs of the user, or can be adjusted by the model. The database includes a plurality of vocabularies, and the technical words are used to refer to specific technical means and technical methods, which are pre-recorded in the database. For example, data cleaning, high-temperature heating, high-temperature manufacturing process, high-temperature glass processing, unmanned driving, route planning, manufacturing, processing, etc. The first relevance of the to-be-queried vocabulary and the technical word can be determined according to the semantic relevance of the to-be-queried vocabulary and the technical word, or can be determined according to the inclusion relationship of the to-be-queried vocabulary and the technical word. For example, the to-be-queried vocabulary is “cup”, and the technical words with a high first relevance to the to-be-queried vocabulary include “ceramic processing”, “high-temperature processing”, “high-temperature manufacturing”, “high-temperature glass processing”, “plastic processing”, etc. The technical words with a low first relevance to the to-be-queried vocabulary include “mask production”, “charging”, “battery processing”, etc. The method for calculating the first relevance includes but is not limited to semantic recognition, word-by-word matching, trained model recognition, and various methods.
[0079] In step 103, for each preset business dimension, at least one second enterprise is determined from the plurality of first enterprises based on the plurality of first enterprises stored in the database, the business data of each first enterprise in the business dimension, the second relevance of the business data to the to-be-queried vocabulary, and the matching degree of the business data to the technical word. The business dimensions include business dimensions, technical dimensions, product dimensions, and trademark dimensions.
[0080] Specifically, the business dimensions are a plurality of business dimensions involved in the development of an enterprise. In the embodiments of the present application, the business dimensions include business dimensions, technical dimensions, product dimensions, and trademark dimensions. When the business dimension is the business dimension, the business data is the business data of the first enterprise. The business data includes the business scope. When the business dimension is the technical dimension, the business data includes the technical data published by the first enterprise. The technical data includes technical literature and technical text. When the business dimension is the product dimension, the business data includes the products operated by the first enterprise. When the business dimension is the trademark dimension, the business data is the trademark data publicly applied by the first enterprise. Each business data is all the data that can be obtained from a public platform, including but not limited to the data publicly disclosed by the enterprise homepage, the encyclopedia data, the data publicly disclosed by the news website, the papers, the literature, the periodicals, the files, and the keywords of each database.
[0081] According to the business data of the first enterprise, the second correlation degree between the business data and the to-be-queried word, and the matching degree between the business data and the technical word, a second enterprise that meets the second correlation degree standard and the matching degree standard can be screened out from the plurality of first enterprises. In a specific implementation, the second correlation degree of the second enterprise can be set to be greater than a fifth threshold value, and the matching degree can be set to be greater than a sixth threshold value. The fifth threshold value and the sixth threshold value can be preset. The calculation method of the second correlation degree and the matching degree is similar to the calculation method of the first correlation degree, and will not be described herein again.
[0082] According to the business data of each first enterprise stored in the database, the second correlation degree between the business data of each first enterprise and the to-be-queried word in each business dimension is calculated, the matching degree between the first enterprise and the to-be-queried word is determined through the second correlation degree, and the enterprise information can be more accurately determined according to the second correlation degree. The to-be-queried word is taken as a first keyword, and the technical word is taken as a second keyword. Since the first correlation degree between the technical word and the to-be-queried word is greater than the first threshold value, the matching degree between the first enterprise and the technical word is high, which means that the matching degree between the first enterprise and the to-be-queried word is also high. Through the query method combined with the technical word, the query range is expanded, so that the query accuracy is improved under the condition that the matching degree between the first enterprise and the technical word is ensured, and the problem that the query range is small due to the query based only on the to-be-queried word input by the user is avoided.
[0083] In step 104, a third correlation degree between each second enterprise and the to-be-queried content is calculated based on the first correlation degree, the second correlation degree, the matching degree, and a weight set in advance for each business dimension.
[0084] Specifically, after the matching degree and the second correlation degree of the second enterprise in each business dimension are calculated through step 103, the third correlation degree between the second enterprise and the to-be-queried word can be calculated according to the first correlation degree, the second correlation degree, and the matching degree of the second enterprise, and the weight determined for each business dimension. When the second correlation degree, the matching degree, and the weight are constant, the greater the first correlation degree, the greater the third correlation degree between the second enterprise and the to-be-queried word; when the first correlation degree, the matching degree, and the weight are constant, the greater the second correlation degree, the greater the third correlation degree between the second enterprise and the to-be-queried word; when the first correlation degree, the second correlation degree, and the matching degree are constant, the greater the weight of a certain business dimension, the greater the contribution of the second correlation degree and the matching degree of the second enterprise in the business dimension to the third correlation degree.
[0085] In step 105, a query result containing each target enterprise and the third correlation degree of the target enterprise is sent to the query terminal; the target enterprise is a second enterprise whose third correlation degree is greater than a second threshold value.
[0086] Specifically, after obtaining the third correlation degree of the second enterprise according to step 104, the target enterprise whose third correlation degree exceeds the second threshold is determined according to the size of the third correlation degree, the accuracy of the target enterprise is ensured, and the query result carrying the target enterprise and the third correlation degree is returned to the query terminal for display.
[0087] The data query method provided by the embodiment of the present application can form an expanded word set of the to-be-queried word by screening at least one technical word with a first correlation degree to the to-be-queried word exceeding a first threshold from the database, query the target enterprise by the to-be-queried word and the expanded word set, thereby avoiding the problem of too narrow search range caused by querying only according to the content input by the user, and determining the target enterprise from the first enterprise by determining the second correlation degree of the business data of the first enterprise to the to-be-queried word, the matching degree of the business data to the technical word, the first correlation degree of the to-be-queried word to the technical word, and the weight of each business dimension. Compared with the prior art which requires the user to manually search each enterprise name and retrieve the situation of each enterprise, the embodiment of the present application expands the query range by the above method, improves the accuracy of querying the technical development situation of each enterprise under the condition of limiting the first correlation degree and ensuring the matching degree of the first enterprise to the technical word, and helps to solve the problems of large manual workload and inaccurate subjective judgment of the technical development situation caused by the need to manually retrieve the technical situation of each enterprise.
[0088] Figure 2 A flowchart of a screening method provided by the embodiment of the present application is shown in FIG. 1. Figure 2 As shown in FIG. 1, in a feasible implementation, when step 103 is performed, the method further includes the following steps:
[0089] Step 201: For each business dimension, screening the first enterprise containing the business dimension from the database based on the business dimension label of each first enterprise stored in the database; each business dimension label corresponds to a business dimension.
[0090] Specifically, the database stores the business dimension label marked for each first enterprise, and whether the first enterprise contains the business dimension corresponding to the business dimension label is determined according to the business dimension label carried by the first enterprise. For example, the operating dimension corresponds to label 1, the technical dimension corresponds to label 2, the product dimension corresponds to label 3, and the trademark dimension corresponds to label 4. The first enterprise includes enterprise one, and enterprise one carries label 1 and label 3, so enterprise one contains the operating dimension and the product dimension.
[0091] Step 202, for each of the first enterprises screened, extracting business data belonging to the business dimension from the database; the business dimension to which the business data belongs is pre-labeled.
[0092] Specifically, the database contains business data of each first enterprise in each business dimension. After each first enterprise containing the business dimension is screened according to step 201, the business data belonging to the business dimension is extracted from the database.
[0093] Step 203, extracting first data with semantic relevance to the query vocabulary exceeding a third threshold from the business data.
[0094] Specifically, after determining the business data of the first enterprise in the business dimension according to step 202, the scope of the business data is narrowed, and the business data in the business dimension can be analyzed to calculate the semantic relevance of the query vocabulary to the business data. When the business data is a vocabulary, the semantic relevance of the two can be directly obtained; when the business data is a long text, the keywords (or core words) in the text can be extracted, the semantic similarity and correlation between the keywords and the query vocabulary are calculated, and the semantic relevance of the business data to the query vocabulary is determined according to the semantic similarity and correlation between the keywords and the query vocabulary. The first data consistent with the query vocabulary in the business data is determined by the above method.
[0095] Step 204, for each of the first data, if the matching degree of the first data to the target technical word exceeds a fourth threshold, the first data is determined as second data; the target technical word is at least one of the technical words.
[0096] Specifically, the more technical words that exist in the first data, the higher the value of the matching degree, or in other words, the more words with close semantic similarity to the technical words that exist in the first data, the higher the matching degree of the first data. The first data with a matching degree to the target technical word exceeding the fourth threshold is determined as the second data.
[0097] Step 205, based on the second data of each first enterprise, according to the data amount of the second data, the matching degree of each second data, and the semantic relevance of each second data, at least one second enterprise is determined from the plurality of first enterprises.
[0098] Specifically, after determining the second data according to step 204, the greater the data amount of the second data, or the higher the matching degree of the second data, or the higher the semantic correlation degree of the second data, the closer the first enterprise corresponding to the second data is to the target enterprise that the user wants to query. According to the above method, the second enterprise closest to the target enterprise that the user wants to query is selected from the plurality of first enterprises, for example, the total value of each first enterprise is calculated according to the data amount, the matching degree, and the semantic correlation degree, the plurality of first enterprises are sorted according to the total value of each first enterprise, a specific proportion of the second enterprises is selected from the plurality of first enterprises, or the first enterprise with a total value exceeding a preset total value is determined as the second enterprise. The preset total value and the specific proportion can be adjusted.
[0099] In a feasible implementation, the marking method of the business dimension includes:
[0100] Step 210, based on the enterprise data of the first enterprise stored in the database, at least one enterprise feature is extracted from the enterprise data for each first enterprise by an entity recognition algorithm; the enterprise data includes the business data of the first enterprise in each business dimension; the enterprise feature is used to describe the attributes of the first enterprise.
[0101] Specifically, the enterprise feature includes the basic information and the core technology of the enterprise, which can represent the characteristics of each enterprise that are different from other enterprises, and the content in the enterprise feature is used to describe: core product, core technology, enterprise name, establishment time, industry, geographic location, technology classification, technology efficacy, news dynamics. The enterprise feature is extracted from the business data by domain phrase discovery, named entity recognition, adaptive knowledge extraction and other technologies.
[0102] Step 211, determine the target business dimension to which each enterprise feature belongs by using a pre-trained dimension marking model, and mark the first label corresponding to each business dimension in the target business dimension for the enterprise feature; the target business dimension is at least one of the business dimensions.
[0103] Specifically, after extracting each enterprise feature, the target business dimension to which each enterprise feature belongs is determined by using a pre-trained dimension marking model. The first label can be a specific business dimension directly, for example, when the enterprise feature is “cup manufacturing technology”, the first label corresponding to the enterprise feature is “technology dimension”, and when the enterprise feature is “sell goods A”, the first label corresponding to the enterprise feature is “product dimension”. It can also be that the first label is a specific number, each number uniquely corresponds to a business dimension. Each first label uniquely corresponds to a business dimension.
[0104] Step 212, based on the first label marked for each enterprise feature of the first enterprise, the target business dimension corresponding to the first label, counting the business dimension label of the first enterprise.
[0105] Specifically, after marking the first label for each enterprise feature of the first enterprise through step 211, at least one business dimension label corresponding to the first label is counted, for example, the first enterprise has multiple enterprise features carrying the following first labels in turn: "product dimension", "technology dimension", "technology dimension", "technology dimension", "product dimension", "trademark dimension", regardless of the number of times "technology dimension" appears, the first enterprise is counted once "technology dimension", that is, the business dimension label of the first enterprise includes three, which are "product dimension", "technology dimension" and "trademark dimension".
[0106] In a feasible implementation, before performing step 211 to mark the first label corresponding to the target business dimension for the enterprise feature, the method further comprises:
[0107] Step 220, cleaning the enterprise features by an entity alignment method and an attribute alignment method.
[0108] Specifically, by using various methods such as entity alignment, attribute alignment and entity disambiguation, enterprise features with the same meaning are eliminated, and multiple enterprise features are merged, cleaned and disambiguated.
[0109] Step 221, identifying the semantics of each cleaned enterprise feature by a semantic recognition algorithm, and marking a second label for the enterprise feature based on the semantics; the second label includes a technology label and an attribute label.
[0110] Specifically, the semantics of each enterprise feature is identified to determine whether the enterprise feature belongs to a technology type or an attribute type, for example, the enterprise feature is "located in city A", which is an attribute type, and an "attribute label" is marked for the enterprise feature.
[0111] Step 222, generating an enterprise portrait for each of the first enterprises based on the second label and the enterprise features.
[0112] Specifically, the enterprise portrait of the first enterprise is generated according to the enterprise characteristics and the second labels carried by the enterprise characteristics. The enterprise portrait with different emphases can be generated according to the needs of different users. When the emphasis is on the enterprise technology, the enterprise characteristics with the second label of "technology label" are taken as the emphasis, and the enterprise characteristics with the second label of "attribute label" are taken as the secondary emphasis, and the enterprise portrait is generated. The content of the enterprise characteristics carrying the "attribute label" is used to describe the enterprise name, the establishment time, the industry, the geographical location, the news dynamics, the legal dynamics and the like; and the content of the enterprise characteristics carrying the "technology label" is used to describe the technical literature of the enterprise, the technical use and the like.
[0113] In a feasible implementation, in the step of screening at least one technical word of the to-be-inquired vocabulary from the database in step 102, the method comprises the following steps:
[0114] Based on the to-be-inquired vocabulary, the target graph of the to-be-inquired vocabulary is searched in the database; at least one technical word having a target relationship with the to-be-inquired vocabulary is extracted from the target graph; the target relationship includes a subordinate relationship and an application relationship.
[0115] Specifically, the target graph contains the association relationship between the to-be-inquired vocabulary and each other vocabulary. For example, when the to-be-inquired vocabulary is a computer, the target graph corresponding to the computer includes information such as that the computer contains a display (a subordinate relationship), the computer contains a keyboard (a subordinate relationship), the computer is applied to the field of data processing (an application relationship), and the like. Then the extracted technical word is "data processing".
[0116] When the to-be-inquired vocabulary is high-temperature sterilization technology, the target graph corresponding to the high-temperature sterilization technology includes information such as that the high-temperature sterilization technology is applied to the field of food manufacturing (an application relationship), the high-temperature sterilization technology includes a direct heating method and an indirect heating method (a subordinate relationship), and the like. Then the extracted technical words are "direct heating method" and "indirect heating method".
[0117] In a feasible implementation, before the step of sending the query result to the query terminal in step 105, the method further comprises the following steps:
[0118] A third correlation degree of each target enterprise is extracted from the query result; based on the numerical value of the third correlation degree, the target enterprises in the query result are sorted to obtain an enterprise list containing a sorting result; and the enterprise list is stored in the query result.
[0119] Specifically, the target enterprises are ranked by the calculated third correlation degree to obtain a target enterprise ranking list, and the third correlation degree with a high priority is displayed to facilitate the user to select the enterprise with a higher correlation degree in the query terminal according to the ranking.
[0120] Figure 3 A structure diagram of a data query device provided by an embodiment of the application is shown in the figure. As shown in the figure, the device comprises a first extraction unit 301, a screening unit 302, a determination unit 303, a calculation unit 304, and a sending unit 305. Figure 3
[0121] The first extraction unit 301 is configured to extract a query word from query content sent by a query terminal.
[0122] The screening unit 302 is configured to screen at least one technical word of the query word from a database; the first correlation degree between the technical word and the query word exceeds a first threshold value; and the technical word is a pre-recorded word used to describe technical means.
[0123] The determination unit 303 is configured to determine at least one second enterprise from a plurality of first enterprises in the database based on the plurality of first enterprises, business data of each first enterprise in each business dimension, a second correlation degree between the business data and the query word, and a matching degree between the business data and the technical word; and the business dimension comprises an operation dimension, a technology dimension, a product dimension, and a trademark dimension.
[0124] The calculation unit 304 is configured to calculate a third correlation degree between each second enterprise and the query content based on the first correlation degree, the second correlation degree, the matching degree, and a weight set in advance for each business dimension.
[0125] The sending unit 305 is configured to send a query result comprising each target enterprise and the third correlation degree of the target enterprise to the query terminal; and the target enterprise is a second enterprise whose third correlation degree exceeds a second threshold value.
[0126] In a feasible embodiment, the determination unit is configured to:
[0127] For each business dimension, the determination unit is configured to screen a first enterprise comprising the business dimension from the database based on a business dimension label of each first enterprise stored in the database; and each business dimension label corresponds to a business dimension.
[0128] For each screened first enterprise, the determination unit is configured to extract business data belonging to the business dimension from the database; and the business dimension to which the business data belongs is pre-marked.
[0129] The determination unit is configured to extract first data having a semantic correlation degree with the query word exceeding a third threshold value from the business data.
[0130] For each of the first data, if the matching degree of the first data with a target technical word exceeds a fourth threshold value, the first data is determined as second data; the target technical word is at least one of the technical words.
[0131] Based on the second data of each first enterprise, at least one second enterprise is determined from the plurality of first enterprises according to the data amount of the second data, the matching degree of each of the second data, and the semantic relevance of each of the second data.
[0132] In a feasible implementation, the device further comprises:
[0133] The first identification unit is configured to, when labeling the business dimension label, based on the enterprise data of the first enterprise stored in the database, for each of the first enterprises, extract at least one enterprise feature from the enterprise data by an entity recognition algorithm; the enterprise data comprises business data of the first enterprise under each of the business dimensions; and the enterprise feature is used to describe the attributes of the first enterprise.
[0134] The first labeling unit is configured to determine a target business dimension to which each of the enterprise features belongs by a pre-trained dimension labeling model, and label each of the first labels corresponding to each of the business dimensions in the target business dimension for the enterprise feature; the target business dimension is at least one of the business dimensions.
[0135] The statistical unit is configured to, based on the first label labeled for each of the enterprise features of the first enterprise and the target business dimension corresponding to the first label, count the business dimension label of the first enterprise.
[0136] In a feasible implementation, the device further comprises:
[0137] The cleaning unit is configured to, before labeling each of the enterprise features with a first label corresponding to the target business dimension uniquely, clean the enterprise features by an entity alignment method and an attribute alignment method.
[0138] The second identification unit is configured to identify the semantics of each of the cleaned enterprise features by a semantic recognition algorithm, and label each of the enterprise features with a second label based on the semantics; the second label comprises a technical label and an attribute label.
[0139] The portrait generation unit is configured to generate an enterprise portrait for each of the first enterprises based on the second label and the enterprise feature.
[0140] In a feasible implementation, the screening unit is configured to:
[0141] Based on the to-be-queried vocabulary, find a target graph of the to-be-queried vocabulary in the database.
[0142] Extract at least one technical word having a target relationship with the to-be-queried vocabulary from the target graph; the target relationship includes a subordinate relationship and an application relationship.
[0143] In a feasible implementation, the device further includes:
[0144] A second extraction unit is configured to extract a third correlation degree of each target enterprise from the query result before sending the query result to the query terminal.
[0145] A sorting unit is configured to sort the target enterprises in the query result based on the numerical value of the third correlation degree to obtain an enterprise list containing a sorting result.
[0146] A storage unit is configured to store the enterprise list in the query result.
[0147] The data query device provided by the embodiment can form an expanded vocabulary set of the to-be-queried vocabulary by screening at least one technical word having a first correlation degree exceeding a first threshold from the database, and can query target enterprises by using the to-be-queried vocabulary and the expanded vocabulary set, thereby avoiding the problem of too narrow search range caused by querying only according to the content input by the user. The target enterprises are determined from the first enterprise by determining the second correlation degree of the business data of the first enterprise in different business dimensions and the to-be-queried vocabulary, the matching degree of the business data and the technical word, the first correlation degree of the to-be-queried vocabulary and the technical word, and the weight of each business dimension. Compared with the prior art in which the user needs to manually search each enterprise name and retrieve the situation of each enterprise, the embodiment can expand the query range by using the above device and the method of querying in combination with the technical word, improve the accuracy of querying the technical development situation of each enterprise under the condition of limiting the first correlation degree and ensuring the matching degree of the first enterprise and the technical word, and help to solve the problems of large manual workload and inaccurate subjective judgment of the technical development situation caused by the need to manually retrieve the technical situation of each enterprise.
[0148] Figure 4 A structural schematic diagram of an electronic device provided by the embodiment is shown, which includes a processor 401, a storage medium 402, and a bus 403. The storage medium 402 stores machine-readable instructions executable by the processor 401. When the electronic device runs the data query method in the embodiment, the processor 401 and the storage medium 402 communicate through the bus 403. The processor 401 executes the machine-readable instructions to perform the steps in the embodiment.
[0149] In embodiments, the storage medium 402 can further execute other machine-readable instructions to perform methods as otherwise described in embodiments, see the description of embodiments for specific method steps and principles performed, which are not repeated in detail here.
[0150] The computer program is stored in the computer readable storage medium, and when the computer program is executed by the processor, the computer program is executed to perform the steps in the embodiments.
[0151] In embodiments, the computer program is executed by the processor, and the computer program can further execute other machine-readable instructions to perform methods as otherwise described in embodiments, see the description of embodiments for specific method steps and principles performed, which are not repeated in detail here.
[0152] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented by other manners. The apparatus embodiment described above is only schematic, for example, the division of the modules is only a logical function division, and actual implementation can have another division manner, for example, a plurality of modules or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed each other can be indirect coupling or communication connection through some communication interface, apparatus or module, which can be electrical, mechanical or other forms.
[0153] The modules described as separate components can or can not be physically separated, and the components displayed as modules can or can not be physical units, that is, can be located in one place, or can be distributed to a plurality of network units. According to actual needs, part or all of the units can be selected to achieve the purpose of the embodiment scheme.
[0154] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit.
[0155] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a nonvolatile computer readable storage medium executable by a processor. Based on this understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a ROM, a RAM, a magnetic disk or an optical disk, and various media that can store program codes.
[0156] The above merely describes the specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A data query method, characterized by, The method comprises: extracting a to-be-queried vocabulary from to-be-queried content sent by a query terminal; screening at least one technical word of the to-be-queried vocabulary from a database; the first correlation degree of the technical word with the to-be-queried vocabulary exceeds a first threshold value; the technical word is a vocabulary pre-recorded for describing technical means; for each preset business dimension, determining at least one second enterprise from a plurality of first enterprises stored in the database based on business data of each first enterprise in the business dimension, a second correlation degree of the business data with the to-be-queried vocabulary, and a matching degree of the business data with the technical word; the business dimension comprises an operation dimension, a technology dimension, a product dimension, and a trademark dimension; the first enterprise is obtained by querying the to-be-queried vocabulary as a first keyword and the technical word as a second keyword; calculating a third correlation degree of each second enterprise with the to-be-queried content based on the first correlation degree, the second correlation degree, the matching degree, and a weight preset for each business dimension; the weight is adjustable; sending a query result comprising each target enterprise and the third correlation degree of the target enterprise to the query terminal; the target enterprise is the second enterprise whose third correlation degree exceeds a second threshold value; screening at least one technical word of the to-be-queried vocabulary from the database comprises: finding a target graph of the to-be-queried vocabulary in the database based on the to-be-queried vocabulary; extracting at least one technical word having a target relationship with the to-be-queried vocabulary from the target graph; the target relationship comprises a subordinate relationship and an application relationship; for each preset business dimension, determining at least one second enterprise from a plurality of first enterprises stored in the database based on business data of each first enterprise in the business dimension, a second correlation degree of the business data with the to-be-queried vocabulary, and a matching degree of the business data with the technical word comprises: for each business dimension, screening a first enterprise comprising the business dimension from the database based on a business dimension label of each first enterprise stored in the database; each business dimension label corresponds to one business dimension; for each screened first enterprise, extracting business data belonging to the business dimension from the database; the business dimension to which the business data belongs is pre-marked; extracting first data having a semantic correlation degree with the to-be-queried vocabulary exceeding a third threshold value from the business data; for each first data, if a matching degree of the first data with a target technical word exceeds a fourth threshold value, determining the first data as second data; the target technical word is at least one of the technical words; the matching degree is directly proportional to the number of technical words; determining at least one second enterprise from the plurality of first enterprises based on second data of each first enterprise according to a data amount of the second data, a matching degree of each second data, and a semantic correlation degree of each second data.
2. The method of claim 1, wherein, The marking method of the business dimension comprises: extracting, for each of the first enterprises, at least one enterprise feature from the enterprise data of the first enterprise stored in the database by an entity recognition algorithm; the enterprise data comprises business data of the first enterprise under each of the business dimensions; the enterprise feature is used to describe the attribute of the first enterprise; determining, by a pre-trained dimension marking model, the target business dimension to which each of the enterprise features belongs, and marking the enterprise feature with a first label corresponding to each of the business dimensions in the target business dimension; the target business dimension is at least one of the business dimensions; based on the first label marked for each of the enterprise features of the first enterprise and the target business dimension corresponding to the first label, counting the business dimension label of the first enterprise.
3. The method of claim 2, wherein, Before marking the enterprise feature with the first label corresponding to each of the business dimensions in the target business dimension, the method further comprises: cleaning the enterprise feature by an entity alignment method and an attribute alignment method; recognizing the semantics of each of the cleaned enterprise features by a semantic recognition algorithm, and marking the enterprise feature with a second label based on the semantics; the second label comprises a technology label and an attribute label; generating an enterprise portrait for each of the first enterprises based on the second label and the enterprise feature.
4. The method of claim 1, wherein, Before sending the query result to the query terminal, the method further comprises: extracting a third relevance of each of the target enterprises from the query result; based on the numerical value of the third relevance, sorting the target enterprises in the query result to obtain an enterprise list containing a sorting result; storing the enterprise list in the query result.
5. A data query apparatus, characterized by comprising: The device comprises: a first extraction unit configured to extract a query word from query content sent by a query terminal; a screening unit configured to screen at least one technology word of the query word from a database; the first relevance of the technology word to the query word exceeds a first threshold value; the technology word is a pre-recorded word used to describe a technical means; a determination unit configured to determine at least one second enterprise from a plurality of first enterprises based on the database, business data of each of the first enterprises under each of the preset business dimensions, a second relevance of the business data to the query word, and a matching degree of the business data to the technology word; the business dimensions include an operation dimension, a technology dimension, a product dimension, and a trademark dimension; the first enterprise is obtained by querying the query word as a first keyword and the technology word as a second keyword; a calculation unit configured to calculate a third relevance of each of the second enterprises to the query content based on the first relevance, the second relevance, the matching degree, and a weight preset for each of the business dimensions; the weight is adjustable; a sending unit configured to send a query result containing each of the target enterprises and the third relevance of the target enterprises to the query terminal; the target enterprise is a second enterprise whose third relevance exceeds a second threshold value; the screening unit is configured to: Based on the to-be-queried vocabulary, find a target graph of the to-be-queried vocabulary in the database; Extract at least one technical word having a target relationship with the to-be-queried vocabulary from the target graph; the target relationship includes a subordinate relationship and an application relationship; The determination unit is configured to: For each business dimension, filter a first enterprise containing the business dimension from the database based on a business dimension label of each first enterprise stored in the database; each business dimension label corresponds to one business dimension; For each filtered first enterprise, extract business data belonging to the business dimension from the database; the business dimension to which the business data belongs is pre-labeled; Extract first data having a semantic correlation degree with the to-be-queried vocabulary exceeding a third threshold from the business data; For each first data, if a matching degree of the first data with a target technical word exceeds a fourth threshold, determine the first data as second data; the target technical word is at least one of the technical words; the matching degree is directly proportional to the number of technical words; Based on the second data of each first enterprise, determine at least one second enterprise from the plurality of first enterprises according to a data amount of the second data, a matching degree of each second data, and a semantic correlation degree of each second data.
6. An electronic device, comprising: Comprise: A processor, a storage medium, and a bus, the storage medium stores machine readable instructions executable by the processor, when the electronic device is running, the processor and the storage medium communicate through the bus, the processor executes the machine readable instructions to perform the steps of the data query method in any one of claims 1 to 4.
7. A computer readable storage medium characterized in that, The computer readable storage medium stores a computer program, the computer program is run by the processor to perform the steps of the data query method in any one of claims 1 to 4.
Citation Information
Patent Citations
Product text determination method and device, computer equipment and medium
CN111104485A