Professional knowledge response method and system, electronic equipment and storage medium
By vectorizing the description text of automobile parts and building an automotive knowledge graph, and combining large language models to process user problems, the problem of low professional knowledge retrieval efficiency in the automobile manufacturing industry is solved, and automatic response, accuracy and real-time improvements are achieved.
Patent Information
- Application Number
- CN202510057853.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
AI Technical Summary
In the automobile manufacturing industry, when technicians review large amounts of documents, they need to manually enter keywords into multiple data sources for searching, resulting in inefficient search and inability to ensure that they can retrieve the desired professional information.
By vectorizing the description text of the automotive component, it is stored in a vector library, and parsing this information is structured in the form of a JSON object. Use the Neo4j graph database to build an automotive knowledge graph, and use the large language model to extract key tags and attribute values from user problems, and search and generate response content through the graph database.
It realizes automatic response to user questions, ensures the accuracy and real-time nature of knowledge response, reduces manual processing workload, and improves data retrieval efficiency and reliability.
Smart Images

Figure CN119988547A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence, and in particular relates to a professional knowledge answering method, system electronic equipment and storage medium. Background Art
[0002] Technical support and troubleshooting in the automotive manufacturing industry mainly rely on a large number of documents and manuals. These documents are usually stored in Excel spreadsheet format, covering the status and properties of various components of the car under different working conditions.
[0003] Currently, when technicians consult these documents, they need to manually enter keywords in multiple data sources for retrieval, which is inefficient and cannot guarantee that the desired professional information can be retrieved. This impact is particularly obvious when technicians are unable to accurately set keywords or handle complex technical issues. Summary of the invention
[0004] In view of this, an embodiment of the present invention provides a professional knowledge answering method, a system electronic device and a storage medium, which are used to solve the current problems of low efficiency and low accuracy in professional knowledge retrieval.
[0005] In a first aspect of an embodiment of the present invention, a method for answering professional knowledge is provided, including: After the automobile parts description text is vectorized through the vector model, it is stored in the vector library; Parsing the automobile parts information table in the vector library, extracting automobile parts information in each Sheet and converting it into a JSON object; Extract the names and attributes of automobile parts in the JSON object, combine the names and attributes of automobile parts into a DataFrame structure and store them in a structured manner; Import the data in the DataFrame into the Neo4j graph database and define the corresponding entities and relationships; Use the large language model to extract key tags and attribute values from the user's input questions, and convert the key tags and attribute values into regularized tags in JSON format; If the regularized label contains a component name, the attribute information related to the component is retrieved through the Neo4j graph database, and the search results are input into the large language model to generate the response content.
[0006] In a second aspect of an embodiment of the present invention, a professional knowledge answering system is provided, including: A vectorization module is used to vectorize the description text of automobile parts through a vector model and store it in a vector library; A parsing and conversion module, used for parsing the automobile parts information table in the vector library, extracting the automobile parts information in each Sheet and converting it into a JSON object; The structuring module is used to extract the names and attributes of automobile parts in the JSON object, combine the names and attributes of automobile parts into a DataFrame structure and store them in a structured manner; The graph building module is used to import the data in the DataFrame into the Neo4j graph database and define the corresponding entities and relationships; The question processing module is used to extract key tags and attribute values from the questions input by the user using a large language model, and convert the key tags and attribute values into regularized tags in JSON format; The retrieval response module is used to retrieve the attribute information related to the component through the Neo4j graph database if the regularized label contains the component name, and input the retrieval result into the large language model to generate the response content.
[0007] In a third aspect of an embodiment of the present invention, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method described in the first aspect of the embodiment of the present invention when executing the computer program.
[0008] In a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method provided in the first aspect of the embodiment of the present invention are implemented.
[0009] In an embodiment of the present invention, a vector library is constructed by vectorizing the description text of automobile parts, parsing the vector library, combining automobile parts and attributes into a DataFrame structure, and constructing an automobile knowledge graph based on a graph database. User questions are understood through a large language model, and user response content is generated by searching the graph database. This not only enables automatic responses to user questions, but also ensures the accuracy and real-time nature of knowledge responses. The labor cost of knowledge extraction is low, and the data retrieval efficiency and reliability are high. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0011] Figure 1 A flowchart of a professional knowledge answering method provided by an embodiment of the present invention; Figure 2 A schematic diagram of the structure of a professional knowledge response system provided by an embodiment of the present invention; Figure 3 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0012] In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0013] It should be understood that the term "including" and other similar expressions in the specification or claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusions, such as a process, method, system, or device including a series of steps or units is not limited to the listed steps or units. In addition, "first" and "second" are used to distinguish different objects, not to describe a specific order.
[0014] See also Figure 1 , a flowchart of a professional knowledge answering method provided by an embodiment of the present invention includes: S101, vectorizing the automobile component description text through a vector model and storing it in a vector library; The vector model is used to convert high-dimensional discrete data such as words, sentences or image features into low-dimensional continuous vectors, thereby converting text data into a numerical vector form that can be processed by a computer. The vector model can be a mainstream vector model such as the BGE (BAAI General Embedding) model.
[0015] The automobile parts description text in Markdown format is vectorized through a vector model, and the obtained vectors are stored in a vector library.
[0016] Among them, at least the component name, attribute description and applicable conditions in the text are vectorized.
[0017] S102, parsing the automobile parts information table in the vector library, extracting automobile parts information in each Sheet and converting it into a JSON object; The automobile parts information in the vector library exists in a table form, that is, an automobile parts information table, and the automobile parts information of each Sheet in the table will be converted into a json object for structured processing.
[0018] S103, extracting automobile component names and attributes from the JSON object, combining the automobile component names and attributes into a DataFrame structure, and performing structured storage; DataFrame is a core data structure in the Pandas library, which is used to represent a two-dimensional tabular data structure. It is similar to an Excel table or SQL table, with row indexes and column indexes. It supports irregular data tables and has a more flexible data structure and more convenient data operations.
[0019] The fields in the DataFrame structure include: component name, attribute name, attribute value, applicable conditions, etc.
[0020] By converting the automobile parts data in the JSON object into a DataFrame structure, it can be directly imported into the graph database to build a knowledge graph.
[0021] S104, import the data in the DataFrame into the Neo4j graph database, and define corresponding entities and relationships; Neo4j graph database is a NoSQL graph database that can store structured data on the Internet and use graph models to represent data and the relationships between data, making the relationships between data intuitive and easy to query.
[0022] Among them, the parts are taken as entity nodes, and the part attribute values are taken as the relationship between nodes to form an automobile knowledge graph. The entities are generally automobile parts, and the relationship between entities can be represented by attributes, thereby obtaining the corresponding knowledge graph.
[0023] A knowledge graph is a knowledge base that represents entities, relationships, attributes, and other knowledge in a graphical form. By representing knowledge in a structured way, it enables computers to better understand and process human language. S105, extracting key tags and attribute values from the question input by the user using the large language model, and converting the key tags and attribute values into regularized tags in JSON format; The big language model is a model based on machine learning and natural language processing technology. It learns language understanding and generation capabilities by training on large amounts of text data. The big language model can process user input questions, extract keywords and attribute values related to components and attributes in the questions, and construct regularized labels for the extracted key tags and attribute values for data retrieval.
[0024] S106: If the regularized tag contains a component name, attribute information related to the component is retrieved through the Neo4j graph database, and the retrieval result is input into the large language model to generate a response content.
[0025] When the user's question directly contains the name of the component, the attribute information related to the component can be directly retrieved. The large language model can generate the answer content based on the component name, attribute information, etc. and display it to the user.
[0026] Optionally, if the regularized label does not contain the component name, the corresponding component name is searched in the Neo4j graph database through reverse reasoning based on the known attribute information; the search is performed according to the component name, and the search result is input into the large language model to generate a response.
[0027] When the user's question does not contain the name of the component, it is necessary to use reverse reasoning to find the corresponding component name in the graph database using the known attribute value, and then perform forward reasoning to retrieve the attribute-related information to generate the answer content. Especially when the attribute information is incomplete or missing, the system can complete the component location through reverse reasoning, which can improve the retrieval efficiency.
[0028] In this embodiment, by vectorizing the description text of automobile parts, converting the automobile parts information into JSON objects, combining the names and attributes of automobile parts into DataFrame, building a knowledge graph based on the graph database, processing user questions through a large language model, and generating answer content based on the search results of the graph database, thereby realizing automatic extraction of automobile professional knowledge and constructing a corresponding graph knowledge base, which can greatly reduce the workload of manual processing. Combining a large language model with a graph database can facilitate efficient retrieval of user questions and ensure retrieval accuracy. Using the question-and-answer mode for technical support can significantly improve the work efficiency of technicians.
[0029] It should be understood that the serial numbers of the steps in the above embodiments do not imply a sequence of execution, and the execution sequence of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0030] Figure 2 A schematic diagram of the structure of a professional knowledge answering system provided by an embodiment of the present invention, the system includes: A vectorization module 210 is used to vectorize the automobile component description text through a vector model and store it in a vector library; Among them, at least the component name, attribute description and applicable conditions in the text are vectorized.
[0031] A parsing and conversion module 220 is used to parse the automobile component information table in the vector library, extract the automobile component information in each Sheet and convert it into a JSON object; The structuring module 230 is used to extract the names and attributes of automobile parts in the JSON object, combine the names and attributes of automobile parts into a DataFrame structure and perform structured storage; A graph construction module 240 is used to import the data in the DataFrame into a Neo4j graph database and define corresponding entities and relationships; Among them, the components are regarded as entity nodes, and the component attribute values are regarded as the relationship between nodes to form an automobile knowledge graph.
[0032] The question processing module 250 is used to extract key tags and attribute values in the question input by the user by using the large language model, and convert the key tags and attribute values into regularized tags in JSON format; The retrieval response module 260 is used to retrieve the attribute information related to the component through the Neo4j graph database if the regularized label contains the component name, and input the retrieval result into the large language model to generate the response content.
[0033] Optionally, if the regularized label does not contain the component name, the corresponding component name is searched in the Neo4j graph database through reverse reasoning based on the known attribute information; Search by component name and input the search results into a large language model to generate response content.
[0034] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems and modules can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here.
[0035] Figure 3 1 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. The electronic device is used for answering automobile professional knowledge. Figure 3 As shown, the electronic device 3 of this embodiment includes: a memory 310, a processor 320 and a system bus 330, wherein the memory 310 includes an executable program 3101 stored thereon, and those skilled in the art can understand that Figure 3 The electronic device structure shown in the figure does not constitute a limitation of the electronic device, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.
[0036] Combine the following Figure 3 A detailed introduction to the various components of electronic equipment: The memory 310 can be used to store software programs and modules. The processor 320 executes various functional applications and data processing of the electronic device by running the software programs and modules stored in the memory 310. The memory 310 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the electronic device (such as cache data), etc. In addition, the memory 310 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0037] The memory 310 includes an executable program 3101 of a network request method, and the executable program 3101 can be divided into one or more modules / units, which are stored in the memory 310 and executed by the processor 320 to implement professional knowledge questions and answers, etc. The one or more modules / units can be a series of computer program instruction segments that can complete specific functions, and the instruction segments are used to describe the execution process of the computer program 3101 in the electronic device 3. For example, the computer program 3101 can be divided into functional modules such as a vectorization module, a parsing conversion module, a structuring module, a graph construction module, a question processing module, and a retrieval response module.
[0038] The processor 320 is the control center of the electronic device. It uses various interfaces and lines to connect various parts of the entire electronic device. By running or executing software programs and / or modules stored in the memory 310, and calling data stored in the memory 310, it performs various functions of the electronic device and processes data, thereby monitoring the overall status of the electronic device. Optionally, the processor 320 may include one or more processing units; preferably, the processor 320 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, application programs, etc., and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 320.
[0039] The system bus 330 is used to connect the various functional components inside the computer, and can transmit data information, address information, and control information. Its types can be, for example, PCI bus, ISA bus, CAN bus, etc. The instructions of the processor 320 are transmitted to the memory 310 through the bus, and the memory 310 feeds back data to the processor 320. The system bus 330 is responsible for the data and instruction exchange between the processor 320 and the memory 310. Of course, the system bus 330 can also be connected to other devices, such as network interfaces, display devices, etc.
[0040] In the embodiment of the present invention, the executable program executed by the processor 320 included in the electronic device includes: After the automobile parts description text is vectorized through the vector model, it is stored in the vector library; Parsing the automobile parts information table in the vector library, extracting automobile parts information in each Sheet and converting it into a JSON object; Extract the names and attributes of automobile parts in the JSON object, combine the names and attributes of automobile parts into a DataFrame structure and store them in a structured manner; Import the data in the DataFrame into the Neo4j graph database and define the corresponding entities and relationships; Use the large language model to extract key tags and attribute values from the user's input questions, and convert the key tags and attribute values into regularized tags in JSON format; If the regularized label contains a component name, the attribute information related to the component is retrieved through the Neo4j graph database, and the search results are input into the large language model to generate the response content.
[0041] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and modules described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0042] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0043] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A professional knowledge answering method, characterized in that: include: After the automobile parts description text is vectorized through the vector model, it is stored in the vector library; Parsing the automobile parts information table in the vector library, extracting automobile parts information in each Sheet and converting it into a JSON object; Extract the names and attributes of automobile parts in the JSON object, combine the names and attributes of automobile parts into a DataFrame structure and store them in a structured manner; Import the data in the DataFrame into the Neo4j graph database and define the corresponding entities and relationships; Use the large language model to extract key tags and attribute values from the user's input questions, and convert the key tags and attribute values into regularized tags in JSON format; If the regularized label contains a component name, the attribute information related to the component is retrieved through the Neo4j graph database, and the search results are input into the large language model to generate the response content.
2. The method according to claim 1, characterized in that The vectorization processing of the automobile component description text by using a vector model comprises: At least the component name, attribute description and applicable conditions in the text are vectorized.
3. The method according to claim 1, characterized in that Importing the data in the DataFrame into the Neo4j graph database and defining the corresponding entities and relationships includes: The components are regarded as entity nodes and the component attribute values are regarded as the relationships between nodes to form an automobile knowledge graph.
4. The method according to claim 1, characterized in that If the regularized tag contains a component name, then retrieving the component-related attribute information through the Neo4j graph database, and inputting the retrieval result into the large language model to generate a response also includes: If the regularized label does not contain the component name, the corresponding component name is searched in the Neo4j graph database through reverse reasoning based on the known attribute information; Search by component name and input the search results into a large language model to generate response content.
5. A professional knowledge answering system, characterized in that: include: A vectorization module is used to vectorize the description text of automobile parts through a vector model and store it in a vector library; A parsing and conversion module, used for parsing the automobile parts information table in the vector library, extracting the automobile parts information in each Sheet and converting it into a JSON object; The structuring module is used to extract the names and attributes of automobile parts in the JSON object, combine the names and attributes of automobile parts into a DataFrame structure and store them in a structured manner; The graph building module is used to import the data in the DataFrame into the Neo4j graph database and define the corresponding entities and relationships; The question processing module is used to extract key tags and attribute values from the questions input by the user using a large language model, and convert the key tags and attribute values into regularized tags in JSON format; The retrieval response module is used to retrieve the attribute information related to the component through the Neo4j graph database if the regularized label contains the component name, and input the retrieval result into the large language model to generate the response content.
6. The system according to claim 5, characterized in that The vectorization processing of the automobile component description text by using a vector model comprises: At least the component name, attribute description and applicable conditions in the text are vectorized.
7. The system according to claim 5, characterized in that Importing the data in the DataFrame into the Neo4j graph database and defining the corresponding entities and relationships includes: The components are regarded as entity nodes and the component attribute values are regarded as the relationships between nodes to form an automobile knowledge graph.
8. The system according to claim 5, characterized in that If the regularized tag contains a component name, then retrieving the component-related attribute information through the Neo4j graph database, and inputting the retrieval result into the large language model to generate a response also includes: If the regularized label does not contain the component name, the corresponding component name is searched in the Neo4j graph database through reverse reasoning based on the known attribute information; Search by component name and input the search results into a large language model to generate response content.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the professional knowledge answering method as described in any one of claims 1 to 4 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed, the steps of the professional knowledge answering method as claimed in any one of claims 1 to 4 are implemented.