A natural language semantic extraction method and system
By performing word segmentation, grammar analysis and semantic extraction of the target files, building a grammar tree and extracting effective keywords and semantic tags, the problem of inability to accurately understand the intention of job seekers/recruiters in the existing technology is solved, and more accurate job recommendations are achieved.
Patent Information
- Application Number
- CN202111220443.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-20
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2041-10-20
AI Technical Summary
The existing job recommendation method cannot accurately understand the job seeker/recruiter’s intentions, resulting in the recommended position not fully meeting the job seeker/recruiter’s requirements.
By performing word segmentation, grammar analysis and semantic extraction on the target file, a grammar tree is constructed, and effective keywords and semantic tags are extracted to improve the understanding of user intentions.
Improves understanding of user target needs and ensures that the recommended positions more accurately meet job seekers/recruiters' requirements.
Smart Images

Figure CN113886527B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of semantic recognition, and particularly to a natural language semantic extraction method and system applied to an information recommendation system. Background Art
[0002] In the current information age, various professional or comprehensive information platforms can provide the information needed by both users with supply and demand relationships. For example, professional recruitment platforms provide a large number of recruitment information and job hunting information for both recruiters as suppliers and job seekers as demanders; some comprehensive websites provide a large amount of advertising bidding information and advertising placement demand information, etc. Taking the recruitment market as an example, the vast majority of job seekers and recruiters will choose to look for suitable positions and talents on some online recruitment platforms. Usually, job seekers and recruiters will register on platforms such as recruitment websites and recruitment APPs. Job seekers fill in their resumes on them, which record personal information and the positions they hope to obtain, while recruiters fill in recruitment information, which record company information, specific positions to be recruited, and position requirements and other information. Or, job seekers directly search for positions by filling in keywords in the search bar in the unlogged state. Since a large amount of information is gathered on the recruitment platform, it will be a time-consuming and very difficult thing to find a suitable position or talent for oneself in the vast amount of information by simply relying on job seekers and recruiters to manually search. Therefore, in order to increase the success rate of job hunting or recruitment on the recruitment platform and help job seekers and recruiters improve efficiency, some recruitment platforms have launched a job recommendation service, that is, according to algorithms, recommend job information to job seekers. For example, the Chinese patent application with the application number 201811208036.3 and the name "A Job Recommendation Method and System" provides a method of generating a data matrix from the user's access data, using a deep learning algorithm to predict the data matrix, and generating job recommendation data based on the prediction result and the portrait data of the person. The Chinese patent application with the application number 201710947915.7 and the name "Processing Method and Device for Job Recommendation" provides another method of extracting the resume features of job seekers to obtain job seeker feature information, extracting the resume features of the resumes submitted to the recruitment project to obtain job feature information, and by matching these two features, obtaining relevant positions that can be recommended to job seekers according to the matching degree between the two. There are also some other methods for realizing job recommendation, which will not be elaborated one by one here.
[0003] By analyzing the existing job recommendation methods, it is found that most of the recommended jobs do not really meet the requirements of job seekers in some aspects. There may be multiple reasons for this result. One of the important reasons is that the recommendation algorithm cannot accurately understand the intentions of job seekers / recruiters. For example, the method provided in the Chinese invention patent application with the application number 201811208036.3 understands the job seeking intentions of users through user access data and user portrait data. The obtained job seeking intentions do not directly come from users, and it is easy to have understanding deviations. For the Chinese invention patent application with the application number 201710947915.7, although the method it provides is to extract features based on the resumes of users, the accuracy of feature extraction and whether the extracted features can truly reflect the job seeking intentions of users remain to be verified.
[0004] Job seekers / recruiters will clearly write their job seeking / recruitment intentions into their resumes / recruitment information. In theory, the intentions of job seekers / recruiters can be obtained through resumes / recruitment information. However, from practical experience, the jobs finally obtained by job seekers do not completely match the positions indicated in their resumes, and the personnel finally recruited by recruiters do not completely match the recruitment information. Therefore, if only sticking to the content in the resume or recruitment information, many positions / talents that meet the intentions of job seekers / recruiters will obviously be missed. Therefore, being able to correctly understand the resume / recruitment information and truly obtain the intentions of job seekers / recruiters through the deep semantics hidden in the resume / recruitment information is an important factor in providing accurate recommendation information. Unfortunately, there is no such solution yet. Summary of the Invention
[0005] In view of the technical problems existing in the prior art, the present invention proposes a natural language semantic extraction method and system, which improves the understanding of the intentions of user-related description files by performing semantic extraction on user-related description files.
[0006] To solve the above technical problems, according to one aspect of the present invention, the present invention provides a natural language semantic extraction method, which includes the following steps:
[0007] Segment the target file into words in units of sentences to obtain multiple word segmentation units; analyze whether the multiple word segmentation units in a sentence form keywords; in response to the multiple word segmentation units in a sentence forming one or more keywords, extract the sentence containing the keywords; perform grammatical analysis on the sentence containing the keywords to obtain a syntax tree; and extract valid keywords from the syntax tree.
[0008] According to another aspect of the present invention, the present invention further provides a natural language semantic extraction system, which includes a word segmentation module, a sentence extraction module, a grammar analysis module, and a keyword extraction module. Among them, the word segmentation module is configured to segment a target file sentence by sentence to obtain a plurality of word segmentation units; the sentence extraction module is connected to the word segmentation module and is configured to analyze whether the plurality of word segmentation units form keywords and extract sentences containing keywords; the grammar analysis module is connected to the sentence extraction module and is configured to analyze the sentences containing the keywords and construct a grammar tree according to the analysis results; the keyword extraction module is connected to the grammar analysis module and is configured to extract effective keywords from the grammar tree.
[0009] The method and system provided by the present invention extract keywords that meet the requirements of information recommendation from the content of relevant target files according to the needs of information recommendation, and obtain effective keywords through grammar analysis of the sentences containing the keywords, splitting and reconstructing the grammar tree, so as to obtain the true intention of the target file and improve the understanding of the user's target needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Next, the preferred embodiments of the present invention will be further described in detail with reference to the drawings, where:
[0011] Figure 1 is a flowchart of a natural language semantic extraction method applied to an information recommendation system according to an embodiment of the present invention;
[0012] Figure 2 is a flowchart of word segmentation according to an embodiment of the present invention;
[0013] Figure 3 is a schematic diagram of a grammar tree according to an embodiment of the present invention;
[0014] Figure 4A is a schematic flowchart of extracting effective keywords and corresponding semantic tags from the grammar tree according to an embodiment of the present invention;
[0015] Figure 4B is a schematic flowchart of generating a directed acyclic graph according to an embodiment of the present invention;
[0016] Figure 5 is a schematic flowchart of generating a main label according to an embodiment of the present invention;
[0017] Figure 6 is a schematic block diagram of the principle of a natural language semantic extraction system according to an embodiment of the present invention;
[0018] Figure 7It is a schematic block diagram of a sentence extraction module according to an embodiment of the present invention;
[0019] Figure 8 It is a schematic block diagram of a syntax analysis module according to an embodiment of the present invention; and
[0020] Figure 9 It is a schematic block diagram of a keyword extraction module according to an embodiment of the present invention. Detailed implementation manners
[0021] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0022] In the following detailed description, reference may be made to the accompanying drawings that form a part hereof, and in which are shown by way of illustration specific embodiments in which the application may be practiced. In the drawings, like reference numerals describe substantially similar components in different figures. The specific embodiments of the present application have been described in sufficient detail below to enable those of ordinary skill in the art with relevant knowledge and technology to implement the technical solutions of the present application. It should be understood that other embodiments may be utilized or structural, logical, or electrical changes may be made to the embodiments of the present application.
[0023] In order to clearly illustrate the solution of the present invention, the present invention defines the specific meanings of the following terms:
[0024] Keyword: A phrase with diverse expression forms, such as a phrase composed of two nouns, such as "automobile sales", "vehicle sales", "automobile promotion", and "vehicle promotion", etc.
[0025] Word segmentation unit: After performing word segmentation on a sentence during the semantic extraction process, the word segments obtained, which are single characters, two-character words, or phrases composed of multiple single characters, such as "I", "am", "Java", and "engineer" in the segmented sentence "I am a Java engineer".
[0026] Word unit: A word collected in a dictionary, having an independent semantic meaning and not being further divisible, such as "automobile" and "sales".
[0027] Semantic tag: A standardized version of a keyword. For example, the semantic tag "automobile sales" is used as the standardized version of the keywords "automobile sales", "vehicle sales", "automobile promotion", and "vehicle promotion".
[0028] Prefix word: The first of two or more word segmentation units that make up a keyword. For example, "automobile" in "automobile sales" and "vehicle" in "vehicle sales".
[0029] Suffix word: The last of two or more word units that make up a keyword. For example, "sales" in "vehicle sales" and "purchase" in "automobile purchase".
[0030] Figure 1 It is a flowchart of a natural language semantic extraction method applied to an information recommendation system according to an embodiment of the present invention. The method includes:
[0031] Step S1: Segment the target file sentence by sentence to obtain multiple word segmentation units.
[0032] Step S2: Analyze whether the multiple word segmentation units form keywords.
[0033] Step S3: In response to the multiple word segmentation units forming one or more keywords, extract the sentences containing the keywords.
[0034] Step S4: Perform syntactic analysis on the sentences containing the keywords to obtain a syntax tree.
[0035] Step S5: Extract valid keywords and corresponding semantic tags from the syntax tree.
[0036] Among them, in step S1, any existing word segmentation method can be used for word segmentation. For example, the mechanical word segmentation method based on a dictionary, such as the forward maximum matching method, the backward maximum matching method, or the bidirectional matching method, etc. Or it is a word segmentation method based on statistics, which determines the probability of a word being formed by calculating the co-occurrence frequency of a character and the adjacent characters in its context. In one embodiment, the Hierarchical Hidden Markov Model (HHMM) integrates lexical analysis tasks such as Chinese word segmentation, segmentation disambiguation, recognition of out-of-vocabulary words, and part-of-speech tagging into a relatively unified model, and realizes synonym replacement, Chinese word segmentation, and part-of-speech tagging for the character string in a sentence, so as to segment the sentence. As Figure 2 shown, it is a word segmentation flowchart according to an embodiment of the present invention, which specifically includes the following steps:
[0037] Step S11, synonym replacement is performed on the original character string in the sentence. Before word segmentation, synonym replacement is performed on the original character string in the sentence. A plurality of entries and synonym entries that can be replaced are stored in the synonym dictionary. In the original character string, the synonym of the maximum length is searched, and if found, it is replaced with the corresponding replacement word. In the processing process, the uppercase and lowercase conversion of English, the conversion of punctuation marks between Chinese and English, and the conversion of full-width characters and half-width characters can also be performed. For example, "sales manager" is a synonym entry for replacement, and the corresponding entries include "business manager" and "salesmanager". For another example, "HP" is a synonym entry for replacement, and the corresponding entries include "HP". "DBA" is a synonym entry for replacement, and the corresponding entry is "database engineer". When the original sentence is "HP company recruits salesmanager and database engineer positions", it becomes "HP company recruits sales manager and DBA positions" after synonym replacement. Through the above synonym replacement, the non-standardization of semantic analysis corpus can be effectively reduced and the accuracy of semantic analysis can be improved.
[0038] Step S12, adopt the K-Best shortest path method to perform preliminary segmentation on the sentence, so as to obtain K best segmentation results that can cover ambiguity. Query the core data dictionary, search for each word in the sentence to be segmented and all possible situations of it becoming a word, save the query word results in a sparse matrix, and record the frequency of the corresponding word. If the entry is an unsegmentable word (such as a word in the funclist dictionary), the predetermined segmentation and annotation results are output, and the other entries contained in the entry are deleted; if the entry "automatic test" is an unsegmentable word, then the entries "self", "automatic", "automation", "dynamic", "chemical", "test", "test" and "test" are deleted in the matrix accordingly, and only "automatic test" is retained. Traverse all nodes in the sparse matrix, and the front and back words are connected with @, such as: "say @ of", "say @ indeed", query the BigramDict data dictionary, obtain the corresponding probability value, and perform smoothing, and calculate the result as the probability of each edge, so as to obtain a word segmentation graph. Use the K-Best optimal path algorithm to find the optimal k paths in the existing m paths in the word segmentation graph.
[0039] Step S13: Identify out-of-vocabulary words using the underlying Hidden Markov Model. After the preliminary segmentation stage, K optimal paths have been generated, but they may contain some unrecognized out-of-vocabulary words, such as personal names and place names. The purpose of this step is to identify these out-of-vocabulary words. In the present invention, all words in a sentence are divided into three categories: the internal components of personal names, context, and irrelevant words. The words or phrases divided according to this classification rule are called roles. After role division, the current sentence is converted into a role sequence. In one embodiment, the Viterbi algorithm is used to perform role annotation on the segmentation result of a sentence in the preliminary segmentation stage to obtain an optimal role sequence. According to the preset role strings and the optimal role sequence for matching, if a role string is matched in the optimal role sequence, it is determined that the role string is an out-of-vocabulary personal name or place name, and the out-of-vocabulary personal name or place name is used as a node, and its probability is calculated and added to the segmentation graph.
[0040] Step S14: Solve the K-Best shortest path again to obtain an optimized segmentation result. Since the segmentation graph has changed, it is necessary to solve the K-Best shortest path again to obtain an optimized segmentation result. For example, the original "Zhang Huaping said is really reasonable" after optimized segmentation is processed as: "Zhang Huaping said is really reasonable".
[0041] Step S15: Perform Hidden Markov annotation of part-of-speech on the optimized segmentation result, that is, annotate the part-of-speech for each segmentation, such as noun, verb, adjective, etc. After annotation, we get "Zhang Huaping / nr said / v of / ad really / adj reasonable / vt", which provides a basis for the next semantic recognition.
[0042] In step S2, in order to analyze whether the multiple segmentation units can form keywords, the steps include:
[0043] Query the word unit dictionary. When a segmentation unit in the sentence is found in the word unit dictionary, it is determined that the segmentation unit is a word unit. The word unit dictionary contains common word units in the recruitment field. Such as word units in the industry dimension "software", "hardware", word units in the function dimension "engineer", "sales", "customer service", word units in the skill dimension "Java", "floriculture", word units in the language dimension "Japanese", "English", etc.
[0044] When multiple word units are obtained by querying, perform permutation and combination on the multiple word units to obtain multiple phrases. Query the phrase dictionary. When the phrase is found in the phrase dictionary, it is determined that the phrase is a keyword.
[0045] For example, after word segmentation, the sentence is "I am a Java and C++ engineer". By searching the word unit dictionary, the word units "Java", "C++", and "engineer" in the sentence are found, thus confirming that the word units of the sentence are "Java", "C++", and "engineer". By permuting and combining these three word units, phrases such as "Java", "C++", "engineer", "Java engineer", "C++ engineer", "JavaC++", and "JavaC++ engineer" can be obtained. Then, by searching the phrase dictionary, "Java", "C++", "engineer", "Java engineer", and "C++ engineer" can be found. Therefore, the keywords can be confirmed as "Java", "C++", "engineer", "Java engineer", and "C++ engineer", while the two phrases "JavaC++" and "JavaC++ engineer" do not appear in the phrase dictionary, so these two phrases are excluded.
[0046] After the above processing, it can be determined that the sentence "I am a Java and C++ engineer." contains keywords, so the next step of syntactic analysis can be carried out. If there are no keywords in the sentence, the next step of analysis is not performed.
[0047] Through step S2, multiple sentences containing meaningful keywords are screened out from the target file, so that only the sentences containing keywords are syntactically analyzed during syntactic analysis, avoiding syntactic analysis of the entire target file, which improves efficiency and reduces interference information.
[0048] In one embodiment, a syntax tree is generated for each sentence. A syntax tree is provided with a root node Root, and each word unit in the sentence is a node. Among them, the root node Root points to the word unit serving as the predicate (whose part of speech is usually a verb), and then the word unit serving as the predicate points to other word units. Two nodes with a pointing relationship form a certain syntactic relationship. For example, a word unit with the part of speech of a verb points to the word unit serving as the subject, and these two word units form a subject-predicate relationship (nsubj); a word unit with the part of speech of a verb points to the word unit serving as the object of the verb, and these two word units form a verb-object relationship (dobj). In step S4, first, according to the sorting of the word units in a sentence in the sentence, starting from the beginning of the sentence, the pointing relationship and syntactic relationship between two word units are sequentially obtained according to the preset syntactic rules; then, a syntax tree is established based on the pointing relationship and syntactic relationship of the word units.
[0049] In one embodiment, a neural network syntactic relationship analysis model is used to calculate the pointing relationship and syntactic relationship between two word units in each sentence. In this network model, the transfer analysis method is used to obtain the pointing relationship and syntactic relationship between two word units.
[0050] In one embodiment, a configuration structure is constructed. The configuration structure includes three structures: a buffer, a stack, and a dependency. Among them, the buffer is used to store the tokenized units in a sentence, which is equivalent to a queue and follows the first-in, first-out principle. The stack is used to store the tokenized units in a sentence, which is also equivalent to a queue and follows the last-in, first-out principle. The "Root" node is stored at the bottom of the stack first. During each judgment, only the relationship between the two topmost tokenized units in the stack is judged. Therefore, the stack stores the tokenized units whose relationships are to be judged, or the tokenized units that failed the previous judgment and are waiting to be judged again. The dependency is used to store the pointing relationship and grammatical relationship between two tokenized units.
[0051] By transferring the tokenized units in the buffer into the stack one by one, the pointing relationship and grammatical relationship between the two topmost tokenized units in the stack are determined based on the three topmost tokenized units in the stack and the first three tokenized units in the buffer.
[0052] Among them, the transfer operations in the configuration are defined as the following three types:
[0053] Left-Arc: It is determined that there is a grammatical relationship xxx from the topmost element S1 (the first tokenized unit) in the stack to the element S2 (the second tokenized unit) below it (S1 -> S2) (it can also be said that S2 is the child node of S1 and has the grammatical relationship xxx), and S2, which is the child node, is removed from the stack. It is necessary to ensure that there are at least two tokenized units in the stack.
[0054] Right-Arc: It is determined that there is a grammatical relationship xxx from the element S2 below the topmost element S1 in the stack to S1 (S2 -> S1) (or it can be said that S1 is the child node of S2 and has the grammatical relationship xxx), and S1, which is the child node, is removed from the stack. It is necessary to ensure that there are at least two tokenized units in the stack.
[0055] Shift: When it is determined that there is no grammatical relationship between the topmost element S1 in the stack and the element S2 below it, the tokenized unit at the head of the buffer is pushed into the stack.
[0056] Taking the sentence "I am familiar with Java and C++ development." as an example, after tokenization, the relationship between the word units, their sorting numbers in the sentence, and their parts of speech is as shown in Table - 1:
[0057] Table - 1
[0058] Lexical Unit I be familiar with Java and C++ development 。 Serial Number 1 2 3 4 5 6 7 Part of Speech PN VV NR CC NR NN PU
[0059] Among them, PN is a pronoun, VV is other verb, NR is a proper noun, CC is a conjunction, NN is other noun, and PU is punctuation.
[0060] Each row in the following table is a Configuration, including <Stack, Buffer, Dependency> (i.e., the syntactic structures of stack, buffer, and storage). The initial Configuration state is as follows: there is only Root in the stack (Stack), and the token units "I", "familiar with", "Java", "and", "C++", "development", "." are sequentially placed in the buffer (Buffer), and the syntactic structure in the storage is empty. Then, perform a transition operation (Transition) on this Configuration to update the Configuration. Specifically, gradually move the token units in the buffer to the stack, and use the top three words in the stack and the first three words in the buffer to determine the syntactic relationship between the top two words in the stack. If there is a left assignment or right assignment, store the obtained syntactic relationship in the syntactic structure and remove the corresponding token unit from the stack. And so on, the i-th transition operation depends on the neural network prediction structure of the (i - 1)-th Configuration. After multiple iterative updates, until the buffer in the Configuration is emptied, the syntactic tree analysis of the entire sentence is completed.
[0061] Among them, the above operation process is shown in Table - 2 as follows:
[0062] Table - 2
[0063]
[0064] Through the above operations, the pointing relationship and syntactic relationship between two word units in a sentence are obtained, thus obtaining a syntactic tree as Figure 3 shown.
[0065] The above judgment of the pointing relationship and syntactic relationship between two word units in the Stack after the transition operation on a configuration structure can be completed by a neural network prediction model. This will not be elaborated here.
[0066] After obtaining the syntactic tree, in step S5, through Figure 4A the following steps shown in
[0067] Step S51, obtain pairs of token units with various syntactic relationships from the syntactic tree and classify the pairs of token units according to the syntactic relationship.
[0068] As can be seen from the syntax tree, there are various syntactic relationships in a syntax tree that connect two token units. However, when understanding the intention of the target file, some token units and their syntactic relationships are meaningless and even interfere with the understanding of the intention of the target file. For example, for the syntactic relationship between the two token units "closed deal" and "yuan" in the token unit combination "closed deal of over 100 million yuan of drugs" is "dative object as a quantity (nmod:range)". The quantifier meaning expressed in this syntactic relationship is not very helpful for understanding the intention of the target file, so the "dative object as a quantity (nmod:range)" is set as an invalid syntactic relationship. While the syntactic relationship of "conjunctive relationship (conj)" formed by "Java" and "C++" is very likely to represent the intention of the target file, so it is set as a valid syntactic relationship. Therefore, in one embodiment, all syntactic relationships are divided into the following five categories as shown in Table-3:
[0069] Table-3
[0070]
[0071] Among them, the descriptions of syntactic relationship tags not explained in Table-3 are as follows:
[0072] Dobj: Direct object
[0073] Nsubj: Noun as subject
[0074] nmod:topic: Noun as topic
[0075] ccomp: Clausal complement
[0076] compound:nn: Compound noun as a modifier
[0077] amod: Adjective as a modifier
[0078] nmod:assmod: Correlative as a modifier
[0079] amod:ordmod: Ordinal number as a modifier
[0080] compound:vc: Compound of verb clause
[0081] nmod: Noun as a modifier
[0082] advmod: Adverb as a modifier
[0083] nummod: Number as a modifier
[0084] mark:clf (classifier modifier) Classifier as a modifier
[0085] In this embodiment, the parallel relationship (conj), the attributive clause (Acl), the verb-noun relationship, and the noun-noun / adverbial relationship are valid syntactic relationships, and the remaining syntactic relationships are invalid syntactic relationships.
[0086] Taking the sentence "I am familiar with Java and C++ development." as an example, according to the syntactic relationships and corresponding tokenization unit pairs in the sentence determined in the previous step, as shown in Table-4:
[0087]
[0088] According to the classification of syntactic relationships, the invalid syntactic relationships are removed, and the valid syntactic relationships and tokenization unit pairs are shown in Table-5.
[0089] Step S52, reorganize the classified tokenization unit pairs to construct language groups, where the language groups include multiple tokenization units connected step by step in the same path.
[0090] In one embodiment, the classified tokenization units are reconnected by constructing a directed acyclic graph. Among them, when reconnecting, the tokenization unit pairs are split and reconnected according to the syntactic relationships of the tokenization unit pairs themselves and the syntactic relationships of the tokenization unit pairs connected to them. For example: Reorganize the valid tokenization unit pairs in Table-5, and the process is as Figure 4B shown:
[0091] Step S521, first process the tokenization unit pairs with a parallel relationship of category 1. In this embodiment, the tokenization unit pair with a parallel relationship is (C++, Java), and the associated tokenization unit pair with a noun-noun relationship is (development, C++). According to the preset rule: when the head in the tokenization unit pair with a parallel relationship (i.e., the parent node in the tokenization unit pair) is connected to category 3 or 4 in Table-3, the connection of the original tokenization unit pair with a parallel relationship is interrupted, and the two tokenization units in the original parallel relationship are used as child nodes and respectively connected to the parent node in category 3 or 4. Therefore, in this embodiment, first interrupt the connection between "C++" and "Java", and then connect them respectively to "development" as the child nodes of "development".
[0092] Step S522, next process the tokenization unit pairs with an attributive clause syntactic relationship of category 2. Since there are no tokenization unit pairs with an attributive clause relationship in this embodiment, this step is skipped and the tokenization unit pairs with a verb-noun relationship of category 3 are processed. Among them, for the tokenization unit pair with an attributive clause syntactic relationship (from a noun pointing to a verb, the noun is the parent node of the verb), the verb node is connected to the root node (Root) as the child node of the root node (Root), and the original pointing of the tokenization unit pair is changed, that is, from the verb pointing to the noun, and the verb becomes the parent node of the noun.
[0093] The participle unit pairs of verb-object relationships in this embodiment are (familiarize, develop) with verb-object grammatical relationships and (familiarize, me) with subject-predicate grammatical relationships. According to the preset grammatical rules: connect the parent node in the verb-object grammatical relationship to the root node (Root), that is, connect "familiarize" to "Root" as a child node of "Root", thus including the participle unit pair (Root, familiarize) with Root grammatical relationship in this embodiment.
[0094] Step S523: Use the root node "Root" as the start node, and connect each node without a lower-level node to the end node "End", thereby obtaining a complete directed acyclic graph.
[0095] Among them, all the participle units on the same path from the start node to the end node form a language group. As Figure 4B shown, there are three language groups, namely "familiarize" + "develop" + "C++", "familiarize" + "develop" + "Java", and "familiarize" + "me". The participle units on each path are connected step by step from the start node to the end node.
[0096] Step S53: Query whether all the word units that make up the keyword are in the same language group. This step is used to verify whether the keyword determined in step S3 is a valid keyword. The verification method is to query whether all the word units that make up the keyword are in the same language group. If they are in the same language group, it is a valid keyword; otherwise, it is not a valid keyword and should be discarded. For example, the prefix word "develop" and the suffix word "Java" in the keyword "develop Java" extracted from "I am familiar with Java and C++ development." are in the same language group, and the prefix word "develop" and the suffix word "C++" in the keyword "develop C++" are in the same language group, thus determining that "develop Java" and "develop C++" are valid keywords. In step S2, the keyword "purchase glass" is obtained from "I am responsible for cleaning the glass every day and occasionally go back to the purchasing department to clean". When querying the language group, the prefix word "purchase" and the suffix word "glass" are not in the same language group, so "purchase glass" is not a valid keyword and is discarded. Thus, it can be seen that by constructing language groups, mis-extracted keywords can be excluded, and the target file can be understood more accurately.
[0097] Step S54: In response to all the word units that make up the keyword being in the same language group, determine the keyword as a candidate keyword.
[0098] In step S55, filter the candidate keywords, and the filtered ones are valid keywords.
[0099] In one embodiment of the present invention, the word unit has corresponding attributes. For example, for a normal word unit in the white list, the value of its attribute identification flag is set to 0. For a word unit located inside a bracket and after a special symbol, the value of its attribute identification flag is 1; the value of the attribute identification flag of a word unit in the black list is 2; the value of the attribute identification flag of the default word unit added in the title (such as "盇盉") is 3. After the word segmentation is performed in step 1 to obtain the word segmentation unit, while determining that the word segmentation unit is a word unit in the dictionary, the attribute value is set for each word segmentation unit as a word unit according to its target file and the aforementioned attribute rules. For example, for "energy" in "front-end engineer (energy)" in the target file, it is located in brackets, so the attribute identification flag value of "energy" is set to 1, and the attribute identification flag value of "engineer" is set to 0. If the default word unit "盇盉" is extracted from the title of the target file, its attribute identification flag value is set to 3.
[0100] Get the attribute flag value of the word unit of the candidate keyword. When the attribute flag value of all the word units in the candidate keyword is 0, set the attribute flag value of the candidate keyword to 0, that is, the first category of candidate keywords. When there are word units with attribute flag values of 1 in the candidate keywords, it is a second category of candidate keywords. When there are word units with attribute flag values of 2 in the candidate keywords, it is a third category of candidate keywords; when there are word units with attribute flag values of 3 in the candidate keywords, it is a fourth category of candidate keywords.
[0101] When filtering, first perform an inclusion relationship filter on the entire first category of candidate keywords, that is, when all the word units constituting the first candidate keyword are all included in the second candidate keyword, delete the first candidate keyword and mark it. For example, when four keywords "engineer", "backend research and development", "backend engineer" and "backend research and development engineer" are extracted from "backend research and development engineer", since "engineer", "backend research and development" and "backend engineer" are all included in "backend research and development engineer", the first three are deleted and only the keyword "backend research and development engineer" is retained, which not only avoids duplication but also reduces the calculation amount of subsequent matching knowledge nodes.
[0102] When the second category of candidate keywords contains the deleted first category candidate keywords, they will be deleted. For example, three keywords "engineer", "front-end engineer" and "energy engineer" are extracted from "front-end engineer (energy)", and it can be clearly seen that "energy engineer" is an incorrect keyword. According to the filtering process of the aforementioned inclusion relationship, the keyword "engineer" is deleted and "front-end engineer" is retained. Since "energy" is in brackets and its attribute identifier flag value is 1, "energy engineer" belongs to the second category of candidate keywords, which includes the deleted "engineer", so it needs to be deleted, thereby solving the problem of extracting errors when extracting keywords.
[0103] The third category of candidate keywords contains word units in the blacklist, so they need to be deleted.
[0104] When there are word units in the fourth category of candidate keywords that are included in the first and second category candidate keywords, they are deleted. For example, when two keywords "盇盉人事" and "人事经理" are extracted from "盇盉人事经理", "人事经理" is the first category candidate keyword and "盇盉人事" is the third category candidate keyword. Since the word unit "人事" in "盇盉人事" is included in the first category candidate keyword "人事经理", the third category candidate keyword "盇盉人事" is deleted, and only "人事经理" is retained.
[0105] After the above filtering operation, repeated, ambiguous, and incorrectly extracted keywords are filtered out from the candidate keywords, and the remaining keywords are valid keywords.
[0106] Step S56, standardize the valid keywords to obtain corresponding semantic tags. The present invention provides a configuration file, which includes a prefix table and a suffix table, wherein the prefix / suffix table records multiple prefixes / suffixes constituting the keywords, and each prefix / suffix has a corresponding standardized version constituting the semantic tag. For example, the prefix "socket" is a standardized version of the prefix "socket". The standardized suffix of the suffix "selling" is "sales". According to the prefix and suffix of the valid keyword, the prefix table and the suffix table are queried respectively, and the prefix and suffix of the valid keyword are mapped to the standard prefix and standard suffix. For example, the keyword "real estate promotion" is mapped to "real estate sales", "Java development" is mapped to "Java R&D engineer", "business specialist" is mapped to "business personnel", and so on.
[0107] The target file mentioned above can be an entire file, or a certain chapter or paragraph of a file. Through the above steps, semantic tags representing various categories of information are mentioned in a target file. Some of them are very important, while some used may not be important. In order to make the information recommendation system more efficient and accurate when querying and recommending targets based on the target file, the present invention also includes a process for extracting main tags (Postag). Based on multiple semantic tags of the target file, the knowledge graph is used to split and generalize the semantic tags, and then filtered to obtain node tags with standardized information. Multiple main tags are then combined from the node tags. Specifically, as shown in Figure 5 the process shown below.
[0108] Step S61: Input the semantic label into the knowledge graph library. According to the mapping relationship between the semantic label and the node label, multiple nodes are obtained for the semantic label. The knowledge graph library includes multiple associated nodes. Each node includes a node label and one or more corresponding attributes. The nodes are connected to the nodes with which they have a mapping relationship according to different attributes. The mapping relationship is an inclusion relationship or a similarity relationship. For each attribute of each knowledge node in the knowledge graph, according to the inclusion relationship, it can be connected to both the upper-level node and the lower-level node. Therefore, the mapping relationship of one attribute is a chain. Among the nodes on the multi-level chain of this mapping relationship, some are relatively abstract and some are more specific. Since the main label is used for search and matching during job recommendation, if the content of the main label is too specific, it is not conducive to job recommendation. Therefore, some nodes in the mapping relationship of one attribute are not suitable to be combined into the main label, while some can be combined into the main label. In the present invention, the nodes suitable to be combined into the main label are called valid nodes (Fclass). For example, the node "Hibernate" is too detailed and not suitable to be used to combine into the main label, while the node label "Java" can be used to combine the main label. Therefore, the node "Java" belongs to the valid node Fclass. According to the mapping relationship, multiple valid nodes also form a multi-level chain. At the end of this multi-level chain, the valid node without an upper-level node is called the valid root node (TFclass). When the semantic label is input into the knowledge graph library, the prefix and suffix of the semantic label are respectively used to match the nodes in the knowledge graph library. By using the prefix and suffix respectively, a node is obtained, and then according to the mapping relationship between the nodes, multiple associated nodes can be obtained. For example, through the semantic label "Hibernate R & D Engineer", the nodes "Hibernate" and "R & D Engineer" can be obtained. Among them, the node "Hibernate" has skill attributes and industry attributes. On the skill attribute, the upper-level node "Java" can be obtained, and on the industry attribute, the upper-level node "Software" can be obtained. The attribute of the node "R & D Engineer" is function, and the upper-level "Engineer" can be obtained in the mapping relationship of the function attribute.
[0109] Step S62: Merge and filter the obtained multiple nodes to obtain multiple candidate nodes. Among the nodes obtained after the previous step S61, some nodes may be repeatedly matched according to different attributes because they have different attributes at the same time. The repeated nodes are merged into the same node, and the unsuitable nodes are filtered out. For example, when the user's current job title is Energy Engineer and the selected industry is "Oil / Chemical / Mining". Since there is no node "Energy" in the mapping relationship of the valid root nodes with "Chemical" and "Mining" respectively, that is, the valid root nodes do not match the node "Energy", all the nodes with "Chemical" and "Mining" as the valid root nodes need to be filtered out.
[0110] Step S63: Obtain valid nodes corresponding to the candidate nodes. Herein, the valid node Fclass refers to a node that can be combined into a main label. In this embodiment, each attribute (Cube) of each node Class has a corresponding valid node Fclass. After the candidate nodes are determined, the corresponding valid nodes can be obtained according to their mapping relationship.
[0111] Step S64: Combine the valid node labels in pairs to obtain multiple main labels. For example, "Java" and "engineer" are combined into "Java engineer", "real estate" and "salesperson" are combined into "real estate salesperson", and "environmental protection" and "investigator" are combined into "environmental protection investigator".
[0112] The multiple main labels are semantically extended accordingly based on the information of the target file, fully summarizing the intention of the target file and providing a basis with standard information for the search and matching of positions.
[0113] Figure 6 is a principle block diagram of a natural language semantic extraction system according to an embodiment of the present invention, which includes a word segmentation module 1, a sentence extraction module 2, a syntax analysis module 3, a keyword extraction module 4, and a standardization module 5. Among them, the word segmentation module 1 is configured to segment the target file into sentences to obtain multiple word segmentation units. The sentence extraction module 2 is connected to the word segmentation module 1 to analyze whether the multiple word segmentation units form keywords and extract the sentences containing keywords. Specifically, as Figure 7 shown, the sentence extraction module 2 includes a word unit determination unit 21, a phrase combination unit 22, a keyword determination unit 23, and a sentence determination unit 24. Among them, the word unit determination unit 21 queries the word unit dictionary based on the word segmentation units in a sentence. If the word segmentation unit is found in the word unit dictionary, it is determined that the word segmentation unit is a word unit that can form a keyword. If the word segmentation unit is not found in the word unit dictionary, the current word segmentation unit cannot form a keyword. After the word unit determination unit 21 queries each word segmentation unit in a sentence one by one, all the word units that can form keywords in the current sentence can be determined. For example, for the sentence "I am responsible for cleaning the glass every day and occasionally go back to the purchasing department to clean." after word segmentation, the word units "purchasing" and "glass" are obtained. Another example is that for the sentence "I am familiar with Java and C++ development.", the word units "Java", "C++", and "development" can be obtained.
[0114] The phrase combination unit 22 is connected to the word unit determination unit 21 and is configured to perform permutations and combinations on multiple word units to obtain multiple phrases. For example, based on the word units "purchase" and "glass", keywords "purchase", "glass", and "purchase glass" are combined. Based on the word units "Java", "C++", and "development", keywords "Java", "C++", "development", "Java C++", "C++ development", "Java development", and "JavaC++ development" are combined.
[0115] The keyword determination unit 23 is connected to the phrase combination unit 22 and is used to query a phrase dictionary. In response to the phrase being found in the phrase dictionary, the phrase is determined as a keyword. For example, "purchase", "glass", and "purchase glass" can all be found in the phrase dictionary, so they are all keywords. However, "JavaC++" and "JavaC++ development" cannot be found in the phrase dictionary, so these two are not keywords, and the other ones are keywords.
[0116] The sentence determination unit 24 is connected to the keyword determination unit 23 and is used to extract the sentences containing keywords in the target file as the basic corpus for syntactic analysis.
[0117] The syntactic analysis module 3 is connected to the sentence extraction module 2 and is used to analyze the sentences containing the keywords and construct a syntax tree based on the analysis results. In one embodiment, as Figure 8 shown, the syntactic analysis module 3 includes a part-of-speech tagging unit 31, a participle unit pair determination unit 32, and a syntax tree construction unit 33. Among them, the part-of-speech tagging unit 31 tags the part of speech for each participle unit in the sentence; the participle unit pair determination unit 32 is connected to the part-of-speech tagging unit 31 and determines the syntactic relationship and pointing relationship between two participle units based on preset syntactic rules and the part of speech of the participle units. In one embodiment, the participle unit pair determination unit 32 uses a transfer analysis model with a neural network to predict the syntactic relationship and pointing relationship between two participle units. For the specific process, see the process shown in Table 2. This will not be elaborated here. The syntax tree construction unit 33 is connected to the participle unit pair determination unit 32 and connects the corresponding participle units in the sentence according to the pointing relationship in the participle unit pair to establish a syntax tree. As Figure 3 the syntax tree shown.
[0118] The keyword extraction module 4 is connected to the syntactic analysis module 3 and is used to extract effective keywords from the syntax tree. In one embodiment, as Figure 9As shown, the keyword extraction module 4 includes a first filtering unit 41, a language group construction unit 42 and a valid keyword determination unit 43, wherein the first filtering unit 41 is connected to the grammatical analysis module 3, and filters out the word segmentation unit pairs with invalid grammatical relations from the grammatical tree, and the classification of grammatical relations is shown in Table-3. The word segmentation units with invalid grammatical relations are filtered out according to Table-3. The language group construction unit 42 is connected to the first filtering unit 41, and multiple word segmentation units with valid grammatical relations are reorganized to obtain one or more language groups. In one embodiment, the language group construction unit 42 determines different language groups by constructing a directed acyclic graph. For details, please refer to the above description, which will not be repeated here. The valid keyword determination unit 43 is connected to the language group construction unit 42, and determines the keywords whose all word units constituting the keywords are located in the same language group as valid keywords. In this way, keywords with incorrect combinations can be filtered out, for example, "purchase" and "glass" in "purchase glass" are located in different language groups, so it can be determined that "purchase glass" is not a valid keyword.
[0119] Even if all the word units constituting the keywords are in the same language group, not all the keywords are suitable. For example, from the sentence "I am familiar with Java and C++ development.", five keywords such as "Java", "C++", "development", "C++ development" and "Java development" are obtained. However, some of the keywords here have a containment relationship. For job search and matching, these keywords with a containment relationship are not conducive to improving the accuracy rate, but also increase the workload. Therefore, in a better embodiment, these keywords with a containment relationship are deleted. In addition, for some situations that may cause errors in constructing keywords, such as some punctuation marks, added default words, etc., or blacklisted word units, in order to avoid these problems, in a further embodiment, the keyword extraction module 4 also includes a second filtering unit 44, which can filter the valid candidate keywords obtained by the valid keyword determination unit 43. Please refer to the description of step S55 in Figure 4 for details, which will not be repeated here.
[0120] The standardization module 5 is connected to the keyword extraction module 4 and is used to standardize the effective keywords to obtain corresponding semantic tags.
[0121] In a further embodiment, the system further includes a main label construction module 6, which includes a knowledge node matching unit 61 and a main label combination unit 62. Among them, the knowledge node matching unit 61 is connected to the standardization module, uses the knowledge graph library to match semantic labels to obtain corresponding multiple knowledge nodes, and merges and filters the multiple nodes to extract valid nodes therefrom. The main label combination unit 62 is connected to the knowledge node matching unit 61, combines the multiple valid nodes in pairs to generate one or more main labels. Among them, the attributes of the knowledge nodes are used as the main label types, and the label contents of two nodes are used as the corresponding main label contents.
[0122] The method and system provided by the present invention extract keywords that meet the requirements of information recommendation from the content of the target file according to the needs of information recommendation, filter out incorrect and inappropriate keywords by performing syntactic analysis on the sentences containing the keywords, and then standardize the keywords. Multiple main labels that can understand the true semantics of the entire text expression are obtained by splitting and reconstructing the standardized keywords. Since the present invention is not limited to the keywords in the file, but appropriately splits, generalizes, and reorganizes the keywords, it can understand the deep semantics hidden in the target file information. And after information standardization, it can effectively simplify the complexity of information search and matching, improve the speed and accuracy of information search and matching, and provide a good semantic basis for information search, matching, and recommendation.
[0123] The above embodiments are only for illustrating the present invention and are not intended to limit the present invention. Those of ordinary skill in the relevant technical fields can also make various changes and modifications without departing from the scope of the present invention. Therefore, all equivalent technical solutions should also fall within the scope of the disclosure of the present invention.
Claims
1. A natural language semantic extraction method, which includes: Performing word segmentation on a target file sentence by sentence to obtain a plurality of word segmentation units; Analyzing whether a plurality of word segmentation units in a sentence constitute keywords; In response to a plurality of word segmentation units in a sentence constituting one or more keywords, extracting sentences containing the keywords from the target file; Performing syntactic analysis on the sentences containing the keywords to obtain a syntax tree; And Extracting valid keywords from the syntax tree; Wherein the step of extracting valid keywords from the syntax tree includes: Obtaining pairs of word segmentation units with valid syntactic relationships from the syntax tree, and classifying the pairs of word segmentation units according to the syntactic relationships; Recombining the classified pairs of word segmentation units by constructing a directed acyclic graph to construct language groups. When recombining, splitting and reconnecting the pairs of word segmentation units according to the syntactic relationships of the pairs of word segmentation units themselves and the syntactic relationships of the pairs of word segmentation units connected to them, wherein the language groups include a plurality of word segmentation units connected step by step on the same path; Querying whether all the word units constituting the keywords are located in the same language group; and In response to all the word units constituting the keywords being located in the same language group, determining the keyword as a valid keyword.
2. The method according to claim 1, wherein the step of analyzing whether a plurality of word segmentation units in a sentence constitute keywords includes: Querying a word unit dictionary, and in response to querying one or more word segmentation units in the sentence in the word unit dictionary, determining the word segmentation unit as a word unit; In response to obtaining a plurality of word units in the sentence, performing permutation and combination on the plurality of word units to obtain a plurality of word groups; And Querying a word group dictionary, and in response to querying the word group in the word group dictionary, determining the word group as a keyword.
3. The method according to claim 1, wherein the step of performing syntactic analysis on the sentence containing the keyword includes: According to the sorting of the word segmentation units in the sentence, starting from the beginning of the sentence, sequentially obtaining a plurality of pairs of word segmentation units and their syntactic relationships according to preset syntactic rules, wherein the pair of word segmentation units includes two word segmentation units with a pointing relationship; and Establishing a syntax tree based on the pointing relationship and syntactic relationship in the pair of word segmentation units.
4. The method according to claim 3, further comprising: Adopting a neural network syntactic relationship analysis model and using the transfer analysis method to obtain the pointing relationship and syntactic relationship of a plurality of pairs of word segmentation units in each sentence.
5. The method according to claim 1, wherein in response to all the word units constituting the keyword being located in the same language group, determining the keyword as a candidate keyword; filtering the candidate keyword to obtain a valid keyword.
6. The method according to claim 5, wherein the step of filtering the candidate keyword further includes: Classify the candidate keywords into the first - type candidate keywords, the second - type candidate keywords, the third - type candidate keywords, and the fourth - type candidate keywords according to the attributes of the word units that make them up; among them, the attributes of all word units in the first - type candidate keywords are normal word units in the whitelist; the second - type candidate keywords include word units located inside parentheses or behind special symbols; the third - type candidate keywords include word units in the blacklist; the fourth - type candidate keywords include default word units added in the title; Among the first - type candidate keywords, when all word units that make up the first - type candidate keyword are included in the second - type candidate keyword, delete the first - type candidate keyword and mark it; When the second - type candidate keyword contains the deleted first - type candidate keyword, delete the second - type candidate keyword; Delete the third - type candidate keywords; and When there are word units in the fourth - type candidate keyword that are included in the first - type and second - type candidate keywords, delete the fourth - type candidate keyword.
7. The method according to claim 1, further comprising: Standardize the valid keywords to obtain corresponding semantic tags.
8. The method according to claim 7, further comprising: Using a knowledge graph library to match corresponding multiple knowledge nodes for the semantic tags, the knowledge graph library includes multiple associated knowledge nodes, the knowledge nodes include node tags and corresponding one or more attributes, and the knowledge nodes are connected to knowledge nodes with a mapping relationship with them according to different attributes; And Combining the multiple matched knowledge nodes in pairs to generate one or more main tags, where the attributes of the knowledge nodes are used as the main tag types, and the contents of the two node tags are combined together as the corresponding main tag contents.
9. A natural language semantic extraction system, comprising: A word - segmentation module, configured to segment a target file into sentences to obtain multiple word - segmentation units; A sentence - extraction module, connected to the word - segmentation module, configured to analyze whether the multiple word - segmentation units form keywords and extract sentences containing keywords; A syntax - analysis module, connected to the sentence - extraction module, configured to analyze the sentences containing the keywords and construct a syntax tree according to the analysis results; And A keyword - extraction module; Connected to the syntax - analysis module, configured to extract valid keywords from the syntax tree; Among them, the keyword - extraction module includes: A first - filtering unit, connected to the syntax - analysis module, configured to filter out pairs of word - segmentation units with invalid syntax relationships from the syntax tree; A language - group construction unit, connected to the first - filtering unit, reorganizes multiple word - segmentation units with valid syntax relationships to obtain one or more language groups. Among them, a directed acyclic graph is constructed to reorganize the classified pairs of word - segmentation units to construct language groups. When reorganizing, the pairs of word - segmentation units are split and re - connected according to their own syntax relationships and the syntax relationships of the pairs of word - segmentation units connected to them. The language group includes multiple word - segmentation units connected step - by - step on the same path; and An effective keyword determination unit, which is connected to the language group construction unit, determines keywords whose all word units constituting the keyword are in the same language group as effective keywords.
10. The system according to claim 9, wherein the sentence extraction module includes: A word unit determination unit, which is connected to a word unit dictionary and is configured to query the word unit dictionary to determine whether the word segmentation unit in the current sentence is a word unit in the word unit dictionary; A phrase combination unit; which is connected to the word unit determination unit and is configured to perform permutation and combination on multiple word units determined in the sentence to obtain multiple phrases; A keyword determination unit, which is connected to the phrase combination unit and a phrase dictionary and is configured to query the phrase dictionary to determine whether the phrase obtained by permutation and combination is a phrase in the phrase dictionary, and determines the phrase obtained by permutation and combination and located in the phrase dictionary as a keyword; and A sentence determination unit, which is connected to the keyword determination unit and is configured to extract sentences containing keywords.
11. The system according to claim 9, wherein the syntax analysis module includes: A part-of-speech tagging unit, which is configured to tag the part of speech for each word segmentation unit in the sentence; A word segmentation unit pair determination unit, which is connected to the part-of-speech tagging unit and is configured to determine the syntactic relationship and pointing relationship between two word segmentation units based on preset syntax rules and the part of speech of the word segmentation units; And A syntax tree construction unit, which is connected to the word segmentation unit pair determination unit and is configured to connect the corresponding word segmentation units in the sentence according to the pointing relationship in the word segmentation unit pair to construct a syntax tree.
12. The system according to claim 9, wherein, The effective keyword determination unit determines the keyword whose all word units constituting the keyword are in the same language group as a candidate keyword, and the keyword extraction module further includes a second filtering unit, which is configured to filter the candidate keyword to obtain an effective keyword.
13. The system according to claim 9, further including a standardization module, which is connected to the keyword extraction module and is configured to standardize the effective keyword to obtain a corresponding semantic label.
14. The system according to claim 13, further including a main label construction module, which is configured to include: A knowledge node matching unit, which is connected to the standardization module and is configured to use a knowledge graph database to match corresponding multiple knowledge nodes for the semantic label, and the knowledge node includes a node label and one or more corresponding attributes; And A main label combination unit, which is connected to the knowledge node matching unit and is configured to generate one or more main labels by combining two by two the multiple nodes obtained by matching, wherein the attribute of the knowledge node is used as the main label type, and the content of the two node labels is used as the corresponding main label content.
Citation Information
Patent Citations
Position recommendation processing method and device
CN107832776A
A method and system for position recommendation
CN109241446A
Processing method and device for searching of natural language by remote sensing data
CN103092979A
Keyword information extraction method based on semantic role analysis
CN111368540A