Construction Method, System, Device and Storage Medium of Vehicle Knowledge Graph Outline
Through automated word segmentation, screening and return to the bidding steps, a knowledge graph outline is constructed, which solves the problems of high labor costs and low efficiency in the existing technology, and realizes the efficient and automated knowledge graph outline construction.
Patent Information
- Application Number
- CN202210668666.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-14
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-06-14
AI Technical Summary
The construction of knowledge graph outlines in the prior art is high and inefficient. Especially in the automotive industry user manual, due to the large number of pictures, the traditional triple extraction method leads to sparse results and excessive vacancy values, and the cost of manual sorting is high.
By obtaining the statement collection, using the preset word segmentation model for word segmentation, determining the macro entity set and relationship set, filtering to obtain the micro entity set and attribute set, and returning the mark based on the relationship set to construct a knowledge graph outline. This method does not require manual intervention and realizes the establishment of a knowledge graph outline through an objective way.
It reduces the labor cost of building a knowledge graph outline, improves efficiency, can effectively process unstructured statement collections, generates a complete knowledge graph outline, and is suitable for fields such as the automotive industry.
Smart Images

Figure CN115062160B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence and natural language processing, and particularly to a method, system, device, and storage medium for constructing a vehicle knowledge graph outline. Background Art
[0002] In the process of realizing human-machine dialogue interaction based on the cockpit voice system of the user manual, a knowledge graph is usually used to macroscopically abstract the outline, and at the same time, objectively existing relationships are set to establish the background knowledge system of the dialogue, so as to assist in the causal relationship explanation of the human-machine dialogue in business.
[0003] However, in the prior art, it is not easy to unify the outline design of the knowledge graph with a set of frameworks, and there are respective domain characteristics in the business scenarios of different vertical domains. The current methods in the industry mainly include: extracting entities, relationships, entities, entities, attributes, and attribute values in the user manual through triples. Although the traditional triple extraction method has a rich theoretical basis, it is not applicable to the extraction of the knowledge graph outline of the user manual in the automotive industry. The reason is that in order to make the user understand more intuitively, the user manual mostly uses pictures to display the usage instructions. When extracting with triples, the results may be very sparse, with a large number of missing values. Usually, the manual method is used, and crowdsourcing is adopted to design in the form of question-and-answer pairs. This method can greatly realize the advantages of collective wisdom, but it does not have the value of industry promotion, with high labor costs and low efficiency. Summary of the Invention
[0004] The main purpose of this application is to provide a method, system, device, and storage medium for constructing a vehicle knowledge graph outline, aiming to solve the technical problems of high labor cost and low efficiency in the construction of the knowledge graph outline in the prior art.
[0005] To achieve the above purpose, this application provides a method for constructing a vehicle knowledge graph outline, and the method for constructing the vehicle knowledge graph outline includes:
[0006] Obtain a set of sentences;
[0007] Input the set of sentences into a preset word segmentation model, and based on the preset word segmentation model, segment the set of sentences to determine a set of macroscopic entities and a set of relationships;
[0008] Screen the set of macroscopic entities to determine a set of microscopic entities and a set of attributes;
[0009] Based on the set of relationships, perform back-labeling on the set of microscopic entities to obtain a set of microscopic entities connected by relationships, and based on the set of microscopic entities connected by relationships and the set of attributes, construct a knowledge graph outline.
[0010] Optionally, the step of screening the set of macro entities to determine the set of micro entities and the set of attributes includes:
[0011] Calculate the similarity of the set of macro entities to determine the set of micro entities, where the set of micro entities includes a set of similar entities and a set of standard entities;
[0012] Determine the attributes of the set of similar entities and the set of standard entities in the set of statements to obtain the set of attributes.
[0013] Optionally, the step of calculating the similarity of the set of macro entities to determine the set of micro entities includes:
[0014] Calculate the similarity coefficient and distance of the set of macro entities to obtain the set of similar entities;
[0015] Determine the standard entities in the set of similar entities;
[0016] Align the set of similar entities based on the standard entities to obtain the set of standard entities;
[0017] Obtain the set of micro entities based on the set of similar entities and the set of standard entities.
[0018] Optionally, the step of determining the attributes of the set of similar entities and the set of standard entities in the set of statements to obtain the set of attributes includes:
[0019] Traverse the preset statement positions of the set of similar entities and the set of standard entities in the set of statements to determine the set of attributes.
[0020] Optionally, the step of tokenizing the set of statements based on the preset tokenization model to determine the set of macro entities and the set of relationships includes:
[0021] Tokenize the set of statements based on the preset tokenization model to obtain a set of words;
[0022] Filter the set of words to obtain a set of words that meet the entity components and a set of words that meet the relationship components in the set of words;
[0023] Store the set of words that meet the entity components in the set of macro entities, and store the set of words that meet the relationship components in the set of relationships to obtain the set of macro entities and the set of relationships.
[0024] Optionally, the step of screening the set of words to obtain the set of words that conform to entity components and the set of words that conform to relationship components in the set of words includes:
[0025] Screen the set of words based on the part-of-speech combined syntactic tree in the preset word segmentation model;
[0026] Classify the word segments in the set of words whose part-of-speech is a noun or a pronoun and whose syntactic components are subject, object, attributive or adverbial as words that conform to entity components, to obtain the set of words that conform to entity components;
[0027] Classify the words in the set of words whose part-of-speech is a verb and whose syntactic components are predicate, adverbial, or complement as words that conform to relationship components, to obtain the set of words that conform to relationship components.
[0028] Optionally, the step of performing back-labeling on the set of microscopic entities based on the set of relationships to obtain the set of microscopic entities connected by relationships includes:
[0029] Input the set of relationships and the set of microscopic entities into a preset back-labeling model;
[0030] Based on the preset back-labeling model, perform back-labeling processing on the relationships between the standard entities to obtain the set of microscopic entities connected by relationships.
[0031] This application also provides a construction system for a vehicle knowledge graph outline, and the construction system for the vehicle knowledge graph outline includes:
[0032] An acquisition module, configured to acquire a set of statements;
[0033] A word segmentation module, configured to input the set of statements into a preset word segmentation model, and based on the preset word segmentation model, perform word segmentation on the set of statements to determine a set of macroscopic entities and a set of relationships;
[0034] A screening module, configured to screen the set of macroscopic entities to determine a set of microscopic entities and a set of attributes;
[0035] A back-labeling module, configured to perform back-labeling on the set of microscopic entities based on the set of relationships to obtain a set of microscopic entities connected by relationships, and based on the set of microscopic entities connected by relationships and the set of attributes, construct a knowledge graph outline.
[0036] This application also provides a construction device for a vehicle knowledge graph outline, and the construction device for the vehicle knowledge graph outline includes: a memory, a processor, and a program stored on the memory for implementing the construction method of the vehicle knowledge graph outline,
[0037] The memory is used to store a program for implementing the method for constructing a vehicle knowledge graph outline;
[0038] The processor is used to execute the program for implementing the method for constructing a vehicle knowledge graph outline to implement the steps of the method for constructing a vehicle knowledge graph outline.
[0039] This application also provides a storage medium, on which a program for implementing the method for constructing a vehicle knowledge graph outline is stored, and the program for implementing the method for constructing a vehicle knowledge graph outline is executed by a processor to implement the steps of the method for constructing a vehicle knowledge graph outline.
[0040] A method, system, device, and storage medium for constructing a vehicle knowledge graph outline provided in this application, compared with the prior art in which the construction of a knowledge graph outline has high labor costs and low efficiency, in this application, a set of statements is obtained; the set of statements is input into a preset word segmentation model, and based on the preset word segmentation model, word segmentation is performed on the set of statements to determine a set of macro entities and a set of relationships; the set of macro entities is screened to determine a set of micro entities and a set of attributes; based on the set of relationships, the set of micro entities is back-labeled to obtain a set of micro entities connected by relationships, and based on the set of micro entities connected by relationships and the set of attributes, a knowledge graph outline is constructed. That is, in this application, based on a preset word segmentation model, an unstructured set of statements is sorted out to obtain the entity and attribute information of standard entities, and then the knowledge of triples is constructed through the set of relationships to establish a knowledge graph outline without additional manual intervention, and the establishment of the knowledge graph outline is achieved objectively. Description of the Drawings
[0041] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application. In order to more clearly illustrate the embodiments of this application or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0042] Figure 1 It is a schematic diagram of the device structure of the hardware operating environment involved in the solution of the embodiment of this application;
[0043] Figure 2 It is a schematic flowchart of the first embodiment of the method for constructing a vehicle knowledge graph outline of this application;
[0044] Figure 3 It is a schematic flowchart of the second embodiment of the method for constructing a vehicle knowledge graph outline of this application.
[0045] The realization of the purpose, functional features and advantages of this application will be further described with reference to the accompanying drawings in combination with the embodiments. Detailed implementation manners
[0046] It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0047] As Figure 1 shown, Figure 1 is a schematic diagram of the terminal structure of the hardware operating environment involved in the embodiment solution of this application.
[0048] The terminal in the embodiment of this application can be a PC, or a smart phone, a tablet computer, an e-book reader, an MP3 (Moving Picture Experts Group Audio Layer III) player, an MP4 (Moving Picture Experts Group Audio Layer IV) player, a portable computer and other movable terminal devices with a display function.
[0049] As Figure 1 shown, the terminal may include: a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0050] Optionally, the terminal may further include a camera, an RF (Radio Frequency) circuit, sensors, an audio circuit, a WiFi module, etc. Among them, the sensors such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display screen according to the brightness of the ambient light, and the proximity sensor can turn off the display screen and / or the backlight when the mobile terminal is moved to the ear. As a kind of motion sensor, the gravity acceleration sensor can detect the magnitude of the acceleration in each direction (generally three axes), and can detect the magnitude and direction of gravity when stationary, and can be used for applications that identify the posture of the mobile terminal (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; of course, the mobile terminal can also be configured with other sensors such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, etc., which will not be elaborated here.
[0051] Those skilled in the art can understand that Figure 1 the terminal structure shown in
[0052] does not limit the terminal, and may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements. Figure 1 As shown in
[0053] In the Figure 1 shown terminal, the network interface 1004 is mainly used to connect to the background server and communicate with the background server for data; the user interface 1003 is mainly used to connect to the client (user side) and communicate with the client for data; and the processor 1001 can be used to call the vehicle knowledge graph outline construction program stored in the memory 1005.
[0054] Referring to Figure 2 , an embodiment of the present application provides a method for constructing a vehicle knowledge graph outline, and the method for constructing the vehicle knowledge graph outline includes:
[0055] Step S100, obtaining a set of statements;
[0056] Step S200, inputting the set of statements into a preset word segmentation model, and based on the preset word segmentation model, performing word segmentation on the set of statements to determine a set of macro entities and a set of relationships;
[0057] Step S300, based on the set of macro entities, determining a set of standard entities and a set of attributes;
[0058] Step S400: Based on the relation set, perform back-labeling on the micro-entity set to obtain a micro-entity set with relation connections, and construct a knowledge graph outline based on the micro-entity set with relation connections and the attribute set.
[0059] In this embodiment, a specific application scenario may be:
[0060] Establish a knowledge graph for a specified unstructured data set. There is no standardized design method for the outline design of the knowledge graph. The traditional method of establishing it by triple extraction is not applicable to the user manuals in the vehicle industry. Manually sorting out the knowledge graph outline also has certain difficulties in migration due to the domain differences between industries, and there is no good generalization and promotion ability.
[0061] The specific steps are as follows:
[0062] Step S100: Obtain a set of statements;
[0063] In this embodiment, the set of statements is a set of statements in natural language. The statements can be in Chinese, English, or other natural languages, and no specific limitation is made here. The set of statements can be a set of statements in industry books / booklets. The industry books / booklets can be physical books or e-books. For example, the user manual in the vehicle industry, which is provided by the developer or marketer of the produced vehicle to the vehicle owner for the vehicle owner to understand various contents of the vehicle. The way to obtain the set of statements can be that the system receives the uploaded electronic version of the user manual in the vehicle industry and the system recognizes the text in it to obtain the set of statements.
[0064] Step S200: Input the set of statements into a preset word segmentation model, and based on the preset word segmentation model, perform word segmentation on the set of statements to determine a macro-entity set and a relation set;
[0065] In this embodiment, the preset word segmentation model is a pre-trained word segmentation model. The system inputs the set of statements into the preset word segmentation model, and the preset word segmentation model performs word segmentation processing on the set of statements. Among them, the model performs word segmentation on the set of statements by the bidirectional maximum matching method, separates the words in the statements to obtain a set of words, where the set of words contains words of various parts of speech and syntactic components, and determines the macro-entity set and the relation set from the words of various parts of speech and syntactic components.
[0066] In this embodiment, the set of macro entities is a set of words whose part of speech is a noun or a pronoun and whose syntactic components are the subject, object, attributive or adverbial. It is a set of entities with a large granularity. That is, in the set of macro entities, in addition to including the set of entities, it also contains components such as attributes and other elements. For example, windshield wipers, sun visors, indicator lights, traffic rules, etc. all belong to the set of macro entities.
[0067] In this embodiment, the set of relationships is a set of words whose part of speech is a verb and whose sentence components are the predicate, adverbial, or complement. For example, automotive lights include interior reading lights and indicator lights. Among them, the word "include" belongs to the set of relationships.
[0068] Specifically, step S200 includes the following steps S210 - S230:
[0069] Step S210: Based on the preset word segmentation model, segment the set of sentences to obtain a set of words.
[0070] In this embodiment, the preset word segmentation model segments the set of sentences by the bidirectional maximum matching method, separates the words in the sentences, and obtains a set of words.
[0071] Step S220: Screen the set of words to obtain a set of words that meet the entity components and a set of words that meet the relationship components in the set of words.
[0072] In this embodiment, the system screens the words in the set of words, classifies the words that meet the entity components into the set of macro entities, and classifies the words that meet the relationship components into the set of relationships.
[0073] Specifically, step S220 includes the following steps S221 - S223:
[0074] Step S221: Screen the set of words based on the part-of-speech combined syntactic tree in the preset word segmentation model.
[0075] Step S222: Classify the words in the set of words whose part of speech is a noun or a pronoun and whose syntactic components are the subject, object, attributive or adverbial as words that meet the entity components, and obtain the set of words that meet the entity components.
[0076] Step S223: Classify the words in the set of words whose part of speech is a verb and whose syntactic components are the predicate, adverbial, or complement as words that meet the relationship components, and obtain the set of words that meet the relationship components.
[0077] In this embodiment, the preset word segmentation model screens through the part-of-speech combined syntactic tree to determine whether the word segmentation conforms to the entity components. For word segments with parts of speech being nouns or pronouns and syntactic components being subjects, objects, attributives, or adverbials, they are grouped and placed into the macro entity set. For word segments that do not conform to the macro entity, they are delimited according to the parts of speech being verbs and sentence components being predicates, adverbials, or complements. In another embodiment, such word segments are classified as predefined relationships and stored in the predefined relationship database. The predefined relationship database can be used as the corpus for relationship classification in the triple of the later machine learning knowledge graph. After the final micro entity extraction, the EoG (Edge-oriented Graphs) algorithm is used to perform back-labeling between entities for the predefined relationships.
[0078] Step S230, store the set of words that conform to the entity components into the macro entity set, and store the set of words that conform to the relationship components into the relationship set, to obtain the macro entity set and the relationship set.
[0079] Step S300, screen the macro entity set to determine the micro entity set and the attribute set;
[0080] In this embodiment, the system determines the micro entity set and the attribute set based on the macro entity set. Among them, the micro entity set is the entity set finally used to construct the knowledge graph outline, which is a set of entities with low granularity. The micro entity set includes a standard entity set and a similar entity set. Since the macro entity set contains not only the micro entity set but also components of attributes and other elements, the attribute set is the attributes of the micro entities. The macro entity set is screened and purified to obtain the micro entity set and the attribute set. For example, the three words "windshield wiper", "wiper blade", and "wiper" are micro entities, among which "windshield wiper" and "wiper blade" are similar entities, and "wiper" is determined as the standard entity.
[0081] Specifically, the step S300 includes the following steps S310 - S320:
[0082] Step S310, calculate the similarity of the entities in the macro entity set to determine the micro entity set, where the micro entity set includes a similar entity set and a standard entity set;
[0083] In this embodiment, calculate the similarity between the entities in the macro entity set, and use the entities with similarity exceeding the threshold as similar entities to obtain the similar entity set, and then determine the standard entity set according to the similar entities.
[0084] Specifically, the step S310 includes the following steps S311 - S314:
[0085] Step S311: Calculate the similarity coefficient and distance for the set of macro entities to obtain the set of similar entities.
[0086] In this embodiment, according to the Jaccard similarity coefficient (calculate the similarity between A and B, i.e., J(A, B)), and the Jaccard distance (calculate the distance between A and B, i.e., Jδ(A, B)), the relevant formulas for calculating text similarity are as follows:
[0087]
[0088]
[0089] Where A and B are different entities, J(A, B) is the similarity coefficient, and J δ (A, B) is the distance.
[0090] Step S312: Determine the standard entity in the set of similar entities.
[0091] In this embodiment, based on the one with the highest frequency of occurrence throughout the text as the standard, determine the standard entity in the set of similar entities, and perform standardized settings. The remaining entities of the same type are regarded as similar entities, and the similar entities need to be aligned consistently with the standard entity. For example, when describing the same entity, different characters such as "windshield wiper", "wiper", and "wiper blade" may appear. Among the three characters, "windshield wiper" has the highest frequency of occurrence throughout the text, and "windshield wiper" is the standard entity.
[0092] Step S313: Based on the standard entity, align the set of similar entities to obtain the set of standard entities.
[0093] In this embodiment, the system aligns the set of similar entities and adjusts parameters to obtain the set of standard entities.
[0094] Step S314: Based on the set of similar entities and the set of standard entities, obtain the set of micro entities.
[0095] In this embodiment, the system combines the set of similar entities and the set of standard entities to obtain the set of micro entities.
[0096] Step S320: Determine the attributes of the set of similar entities and the set of standard entities in the set of statements to obtain the set of attributes.
[0097] In this embodiment, the system determines the attributes of the set of similar entities and the set of standard entities in the set of statements to obtain the set of attributes.
[0098] Specifically, step S320 includes the following steps S321:
[0099] Step S321: Traverse the preset statement positions of the similar entity set and the standard entity set in the statement set to determine the attribute set.
[0100] In this embodiment, the preset statement position can be the sentence above, the current sentence, or the sentence below the source of the entity in the statement set, or other set sentence positions. In this embodiment, for each occurrence of an entity in the similar entity set and the standard entity set in the statement set, the system traverses the original texts of the sentence above, the current sentence, and the sentence below the entity's source to retrieve whether there are attributes semantically related to the entity, and determines the attribute set. For example, when the system retrieves the original texts of the sentence above, the current sentence, and the sentence below the windshield wiper in the user manual, the obtained attribute is driving, that is, the attribute of the entity windshield wiper is driving.
[0101] Step S400: Based on the relationship set, perform relabeling on the micro-entity set to obtain a micro-entity set with relationships connected, and construct a knowledge graph outline based on the micro-entity set with relationships connected and the attribute set.
[0102] In this embodiment, after obtaining the standard entity set and the attribute set, the relationships between the standard entities can be relabeled through methods such as document-level relationship extraction in relationship set extraction, and the relabeling method can use multi-instance learning for remote supervision relationship classification.
[0103] Specifically, step S400 includes the following steps S410 - S420:
[0104] Step S410: Input the relationship set and the micro-entity set into a preset relabeling model;
[0105] Step S420: Based on the preset relabeling model, perform relabeling processing on the relationships between the standard entities to obtain the micro-entity set with relationships connected.
[0106] In this embodiment, first input the relationship set, the standard entity set, and the attribute set into a preset relabeling model. The steps for the preset relabeling model to perform relabeling processing on the relationships between the standard entities are as follows:
[0107] 1. First, initialize θ and set the size of each set to be classified as bs;
[0108] 2. Randomly select groups of sets to be classified and send them into the neural network;
[0109] 3. Find the j-th example mji (1 ≤ i ≤ bs) in each group of classification sets according to the calculation formula of j*;
[0110] 4. Update θ according to the gradient of mji;
[0111] 5. Repeat steps 2 - 4 until convergence reaches the maximum number of times.
[0112]
[0113]
[0114] Based on all elements: entities, attributes, relationships, the knowledge graph outline is constructed.
[0115] A method, system, device, and storage medium for constructing a vehicle knowledge graph outline provided by this application. Compared with the prior art where the construction of the knowledge graph outline has high labor costs and low efficiency, in this application, a set of statements is obtained; the set of statements is input into a preset word segmentation model, and based on the preset word segmentation model, the set of statements is segmented to determine a set of macro - entities and a set of relationships; the set of macro - entities is screened to determine a set of micro - entities and a set of attributes; based on the set of relationships, the set of micro - entities is labeled to obtain a set of micro - entities connected by relationships, and based on the set of micro - entities connected by relationships and the set of attributes, the knowledge graph outline is constructed. That is, in this application, based on the preset word segmentation model, the unstructured set of statements is sorted out to obtain the entity and attribute information of standard entities, and then the knowledge of triples is constructed through the set of relationships to establish the knowledge graph outline without additional manual intervention, and the establishment of the knowledge graph outline is achieved through an objective method.
[0116] In another embodiment, referring to Figure 3, first, through a pre-trained word segmentation model, after word segmentation according to the bidirectional maximum matching method in the model, screening is performed through the part-of-speech combined syntactic tree to determine whether the word segmentation conforms to the entity components. For word segments with part-of-speech being nouns or pronouns and syntactic components being subjects, objects, attributives, or adverbials, they are merged and placed into the macro entity set. By the same principle, for word segments that do not conform to the macro entity, they are delimited according to the part-of-speech being verbs and the sentence components being predicates, adverbials, or complements, and these word segments are classified as predefined relationships and stored in the predefined relationship database. The predefined relationship database can be used as the corpus for relationship classification in the triple of the machine learning knowledge graph in the later stage. After the final micro entity extraction, the EoG algorithm is used to perform back-labeling between entities for the predefined relationships. Calculate the text similarity of the macro entity set, and achieve entity alignment by adjusting parameters. All entities participating in the alignment are set standardly based on the one with the most occurrences in the full text, and the remaining similar entities are regarded as similar entities, and the similar entities need to be aligned consistently with the standard entity. For the entities in the similar entity set and the standard entity set, at each occurrence in the sentence set, traverse the original text of the previous sentence, this sentence, and the next sentence of the entity occurrence to retrieve whether there are attributes semantically related to the entity, and determine the attribute set. After obtaining the standard entity set and the attribute set, the relationships between the standard entities can be back-labeled through methods such as document-level relationship extraction in the relationship set extraction, and the back-labeling method can use multi-instance learning for distant supervision relationship classification.
[0117] This application also provides a construction system for a vehicle knowledge graph outline, and the construction system for the vehicle knowledge graph outline includes:
[0118] An acquisition module, used to acquire a sentence set;
[0119] A word segmentation module, used to input the sentence set into a preset word segmentation model, and based on the preset word segmentation model, perform word segmentation on the sentence set to determine a macro entity set and a relationship set;
[0120] A screening module, used to screen the macro entity set to determine a micro entity set and an attribute set;
[0121] A back-labeling module, used to perform back-labeling on the micro entity set based on the relationship set to obtain a micro entity set with relationships connected, and based on the micro entity set with relationships connected and the attribute set, construct and obtain a knowledge graph outline.
[0122] Optionally, the screening module includes:
[0123] A calculation module, configured to calculate the similarity of the set of macro entities to determine the set of micro entities, where the set of micro entities includes a set of similar entities and a set of standard entities;
[0124] An attribute determination module, configured to determine the attributes of the set of similar entities and the set of standard entities in the set of statements to obtain the set of attributes.
[0125] Optionally, the calculation module includes:
[0126] A similarity calculation module, configured to calculate the similarity coefficient and distance of the set of macro entities to obtain the set of similar entities;
[0127] A standard entity determination module, configured to determine the standard entities in the set of similar entities;
[0128] An alignment module, configured to align the set of similar entities based on the standard entities to obtain the set of standard entities;
[0129] A merging module, configured to obtain the set of micro entities based on the set of similar entities and the set of standard entities.
[0130] Optionally, the attribute determination module includes:
[0131] A traversal module, configured to traverse the preset statement positions of the set of similar entities and the set of standard entities in the set of statements to determine the set of attributes.
[0132] Optionally, the word segmentation module includes:
[0133] A word segmentation module, configured to segment the set of statements based on the preset word segmentation model to obtain a set of words;
[0134] A screening module, configured to screen the set of words to obtain a set of words that meet the entity components and a set of words that meet the relationship components in the set of words;
[0135] A storage module, configured to store the set of words that meet the entity components into the set of macro entities and store the set of words that meet the relationship components into the set of relationships to obtain the set of macro entities and the set of relationships.
[0136] Optionally, the screening module includes:
[0137] A part-of-speech screening module, configured to screen the set of words based on the part-of-speech combined syntactic tree in the preset word segmentation model;
[0138] An entity screening module, configured to divide the words in the word set that are nouns or pronouns in terms of part of speech and are subject, object, attributive or adverbial in terms of syntactic components into words that conform to entity components, so as to obtain the word set that conforms to entity components;
[0139] A relationship screening module, configured to divide the words in the word set that are verbs in terms of part of speech and are predicate, adverbial or complement in terms of syntactic components into words that conform to relationship components, so as to obtain the word set that conforms to relationship components.
[0140] Optionally, the back-labeling module includes:
[0141] An input module, configured to input the relationship set and the micro-entity set into a preset back-labeling model;
[0142] A relationship back-labeling module, configured to perform back-labeling processing on the relationships between the standard entities based on the preset back-labeling model, so as to obtain the set of micro-entities connected by the relationships.
[0143] The specific implementation manner of the vehicle knowledge graph outline construction system of the present application is basically the same as each embodiment of the above vehicle knowledge graph outline construction method, and will not be elaborated here.
[0144] The present application further provides a vehicle knowledge graph outline construction device, where the vehicle knowledge graph outline construction device includes: a memory, a processor, and a program stored on the memory for implementing the vehicle knowledge graph outline construction method,
[0145] The memory is used to store a program for implementing the vehicle knowledge graph outline construction method;
[0146] The processor is configured to execute the program for implementing the vehicle knowledge graph outline construction method to implement the steps of the vehicle knowledge graph outline construction method.
[0147] The specific implementation manner of the vehicle knowledge graph outline construction device of the present application is basically the same as each embodiment of the above vehicle knowledge graph outline construction method, and will not be elaborated here.
[0148] The present application further provides a storage medium, on which a program for implementing the vehicle knowledge graph outline construction method is stored, and the program for implementing the vehicle knowledge graph outline construction method is executed by a processor to implement the steps of the vehicle knowledge graph outline construction method.
[0149] The specific implementation manner of the storage medium of the present application is basically the same as each embodiment of the above vehicle knowledge graph outline construction method, and will not be elaborated here.
[0150] It should be noted that in this document, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including such element.
[0151] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.
[0152] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0153] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A method for constructing a vehicle knowledge graph outline, characterized in that, The method for constructing the vehicle knowledge graph outline includes: Obtain a set of statements; Input the set of statements into a preset word segmentation model, and based on the preset word segmentation model, perform word segmentation on the set of statements to determine a set of macro entities and a set of relationships; Filter the set of macro entities to determine a set of micro entities and a set of attributes; The step of filtering the set of macro entities to determine a set of micro entities and a set of attributes includes: Calculate the similarity of the set of macro entities to determine the set of micro entities, where the set of micro entities includes a set of similar entities and a set of standard entities; Determine the attributes of the set of similar entities and the set of standard entities in the set of statements to obtain the set of attributes; Based on the set of relationships, perform back-labeling on the set of micro entities to obtain a set of micro entities connected by relationships, and based on the set of micro entities connected by relationships and the set of attributes, construct the knowledge graph outline.
2. The method for constructing a vehicle knowledge graph outline according to claim 1, wherein, The step of calculating the similarity of the set of macro entities to determine the set of micro entities includes: Calculate the similarity coefficient and distance of the set of macro entities to obtain the set of similar entities; Determine the standard entities in the set of similar entities; Align the set of similar entities based on the standard entities to obtain the set of standard entities; Based on the set of similar entities and the set of standard entities, obtain the set of micro entities.
3. The method for constructing a vehicle knowledge graph outline according to claim 1, wherein The step of determining the attributes of the set of similar entities and the set of standard entities in the set of statements to obtain the set of attributes includes: Traverse the preset statement positions of the set of similar entities and the set of standard entities in the set of statements to determine the set of attributes.
4. The method for constructing a vehicle knowledge graph outline according to claim 1, wherein The step of inputting the set of statements into a preset word segmentation model, and based on the preset word segmentation model, perform word segmentation on the set of statements to determine a set of macro entities and a set of relationships includes: Based on the preset word segmentation model, perform word segmentation on the set of statements to obtain a set of words; Filter the set of words to obtain a set of words that meet the entity components and a set of words that meet the relationship components in the set of words; Store the set of words that meet the entity components into the set of macro entities, and store the set of words that meet the relationship components into the set of relationships to obtain the set of macro entities and the set of relationships.
5. The method for constructing a vehicle knowledge graph outline according to claim 4, wherein, The step of filtering the set of words to obtain a set of words that meet the entity components and a set of words that meet the relationship components in the set of words includes: Filter the set of words based on the part-of-speech combined syntactic tree in the preset word segmentation model; Classify the words in the set of words with the part of speech of noun or pronoun and the syntactic components of subject, object, attributive or adverbial as words that meet the entity components to obtain the set of words that meet the entity components; Classify the words in the set of words with the part of speech of verb and the syntactic components of predicate, adverbial or complement as words that meet the relationship components to obtain the set of words that meet the relationship components.
6. The method for constructing a vehicle knowledge graph outline according to claim 1, wherein The step of performing back-labeling on the set of micro entities based on the set of relationships to obtain a set of micro entities connected by relationships includes: Input the set of relationships and the set of micro-entities into a preset relabeling model; Based on the preset relabeling model, perform relabeling processing on the relationships between standard entities to obtain a set of micro-entities connected by relationships.
7. A construction system for a vehicle knowledge graph outline, characterized in that, The construction system of the vehicle knowledge graph outline includes: An acquisition module for acquiring a set of statements; A word segmentation module for inputting the set of statements into a preset word segmentation model, and based on the preset word segmentation model, performing word segmentation on the set of statements to determine a set of macro-entities and a set of relationships; A screening module for screening the set of macro-entities to determine a set of micro-entities and a set of attributes; The screening module includes: A calculation module for calculating the similarity of the set of macro-entities to determine the set of micro-entities, where the set of micro-entities includes a set of similar entities and a set of standard entities; An attribute determination module for determining the attributes of the set of similar entities and the set of standard entities in the set of statements to obtain the set of attributes; A relabeling module for relabeling the set of micro-entities based on the set of relationships to obtain a set of micro-entities connected by relationships, and constructing a knowledge graph outline based on the set of micro-entities connected by relationships and the set of attributes.
8. An apparatus for constructing a vehicle knowledge graph outline, characterized in that The construction device of the vehicle knowledge graph outline includes: a memory, a processor, and a program stored on the memory for implementing the construction method of the vehicle knowledge graph outline, The memory is used to store a program for implementing the construction method of the vehicle knowledge graph outline; The processor is used to execute the program for implementing the construction method of the vehicle knowledge graph outline to implement the steps of the construction method of the vehicle knowledge graph outline as described in any one of claims 1 to 6.
9. A storage medium, characterized in that, A program for implementing the construction method of the vehicle knowledge graph outline is stored on the storage medium, and the program for implementing the construction method of the vehicle knowledge graph outline is executed by the processor to implement the steps of the construction method of the vehicle knowledge graph outline as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Knowledge graph construction method and device and electronic device
CN109885698A