Automobile specification analysis, voice broadcast and instruction control method based on large model
Through large-scale model analysis and vectorized storage technology, combined with the large language model and the automatic call of vehicle computer control commands, the automatic sorting and intelligent interaction problems of car book information are solved, and efficient and accurate information retrieval and safe vehicle operation are achieved.
Patent Information
- Application Number
- CN202510344314.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-11
AI Technical Summary
The existing technology has problems such as large workload, error-prone, inaccurate search and insufficient intelligence in the automatic sorting, accurate retrieval and intelligent interaction of car book information, which affect driving safety.
Using a large model-based method, automatic preprocessing, structured analysis and high-dimensional vectorized storage of automobile manual documents is performed, vector library is used for retrieval, and natural language answers are generated in combination with large language models to realize automatic call and safe verification of vehicle computer control commands.
It improves information update efficiency and retrieval accuracy, ensures driving safety, and provides efficient and accurate information retrieval and intelligent interactive experience.
Smart Images

Figure CN120296047A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of voice interaction technology, and in particular to a method for analyzing automobile manuals, voice broadcasting and command control based on a large model. Background Art
[0002] The car manual ("car manual" for short) is a detailed document provided by the car manufacturer, which aims to provide car owners and users with comprehensive vehicle information, operating instructions and maintenance suggestions. With the continuous improvement of the level of automobile intelligence and electrification, traditional paper car manuals are gradually being replaced by electronic car manuals due to problems such as inconvenience in reading and delayed updates. Electronic car manuals are easy to store in the car system, which is convenient for users to read at any time. However, viewing relevant information through the car screen while driving may have an adverse effect on driving safety.
[0003] In view of the above situation, the existing technology proposes a solution based on voice interaction: when the car computer is equipped with a voice device, the system can collect the driver's voice information through the voice device, extract keywords, search for corresponding content in the car book based on the keywords, and then pass the search results to the driver through voice broadcast; when the car computer is not equipped with a voice device, the system sends a voice interaction activation request to the mobile device connected to the car computer, collects user voice information through a preset voice interaction application, extracts keywords, and searches for relevant content in the car book based on the keywords, and finally voice broadcasts the results through the mobile device. This solution effectively reduces the risk of the driver operating the screen while driving.
[0004] Although the existing technology has made some progress in realizing voice interaction of car book information, there are still several shortcomings: First, the organization and update of the text content of the car book mainly rely on manual processing, which is labor-intensive and prone to errors; second, the traditional keyword-based search method is prone to inaccurate search results or incomplete information due to synonyms, similar characters and other problems; third, the existing solution only realizes the retrieval and broadcast of text content, and cannot attach and execute car system commands, and the degree of intelligence is insufficient, which may increase driving safety risks. In summary, the main technical problems faced by the existing technology are: how to realize the automatic organization, accurate retrieval and intelligent interaction of car book information, so as to effectively improve the efficiency of information update and retrieval accuracy, thereby ensuring driving safety. Summary of the invention
[0005] This application provides a method for analyzing, voice broadcasting and command control of automobile manuals based on a large model, which can realize automatic sorting, accurate retrieval and intelligent interaction of automobile manual information. This application provides the following technical solutions:
[0006] In a first aspect, the present application provides a method for parsing, voice broadcasting and command control of automobile manuals based on a large model, the method comprising:
[0007] Upload the car instruction manual document, parse it, convert the parsed car instruction manual document into a high-dimensional vector representation, and store it in a vector database;
[0008] In response to a user's query request, perform vector retrieval in the vector database and generate a retrieval result;
[0009] Input the retrieval result into a large language model to generate a natural language answer that meets the user's needs;
[0010] Call the text-to-speech technology to play the natural language answer, identify whether there is a car machine control command. If not, play it directly. If so, determine whether to execute the car machine control command in response to the user's feedback.
[0011] In a specific feasible implementation, the uploading the car instruction manual document, parsing it, converting the parsed car instruction manual document into a high-dimensional vector representation, and storing it in a vector database includes:
[0012] Upload the car instruction manual document and preprocess it;
[0013] Identify the page numbers and structure of the car instruction manual document, process the headers and footers of each page, divide and identify each chapter;
[0014] Extract the text, images, and tables from the divided car instruction manual document, and configure corresponding car machine control commands for the parsed chapters;
[0015] Create independent data objects for each chapter and store them in a structured format to generate the structured data of the car instruction manual document;
[0016] Perform post-processing and verification on the structured data of the car instruction manual document, convert the structured data into a high-dimensional vector representation, and store it.
[0017] In a specific feasible implementation, the uploading the car instruction manual document and preprocessing it includes:
[0018] Comprehensively identify the document type based on the file suffix, magic code, and media type of the MIME standard;
[0019] For the PDF-format car instruction manual document, call the PDF parsing library to load the file and perform lossless compression on the car instruction manual document;
[0020] For scanned PDF or pages containing embedded picture text, apply optical character recognition technology to convert the image text into editable text;
[0021] Unify the conversion of PDF pages into a standardized format.
[0022] In a specific feasible implementation, identifying the page numbers and structure of the automobile instruction manual document, processing the headers and footers of each page, and dividing and identifying each chapter include:
[0023] Identifying the first page of the automobile instruction manual document as the cover, marking the range of the table of contents page through the alignment feature of the chapter title list and page numbers, and positioning the pages after the table of contents page as the main text;
[0024] Stripping the headers and footers of each page, dividing the chapter boundaries based on predefined format rules and marking the chapter titles, and dividing the starting and ending page numbers of each chapter based on the positions of the chapter titles;
[0025] Further detecting the titles of sub-chapters or subsections within each chapter.
[0026] In a specific feasible implementation, extracting the text, images, and tables in the divided automobile instruction manual document and configuring corresponding vehicle infotainment system control commands for the parsed chapters include:
[0027] Using an image segmentation algorithm based on histogram analysis to distinguish text and image regions in the multi-column document of the divided automobile instruction manual document, extracting features from each distinguished text block, determining the semantic level through weighted analysis, and subdividing the text into paragraphs and sentences;
[0028] For the distinguished image regions, obtaining image objects through the image extraction function of the PDF parsing library, saving the extracted image objects to a predefined common storage location, generating a unique address, and inserting them into the text content of the corresponding chapter;
[0029] For the tables in the automobile instruction manual document, using the table detection algorithm of the table recognition tool to distinguish different structural elements such as titles, rows, and columns, detecting and extracting the table structures on the page, classifying the detected table types, parsing the table content row by row, and integrating the parsed table content with the text content of the corresponding chapter;
[0030] Configuring corresponding vehicle infotainment system control commands for the parsed chapters.
[0031] In a specific feasible implementation, performing vector retrieval in the vector database and generating retrieval results include:
[0032] Invoking the same industry pre-trained embedding model as in the vectorization stage of the automobile instruction manual document to convert the query text into a query vector with the same dimension as the stored data;
[0033] Perform approximate nearest neighbor retrieval based on cosine similarity in a vector database. By calculating the similarity scores between the query vector and the high-dimensional vector representations of the stored car instruction manuals, the top K relevant data with the highest similarity are filtered out;
[0034] The relevant data are sorted by similarity and output as the retrieval result.
[0035] In a specific implementable embodiment, the inputting the retrieval result into a large language model to generate a natural language answer that meets the user's needs includes:
[0036] Input the generated retrieval result, that is, the relevant data, into a pre-trained industry-specific large language model, and then perform semantic understanding and context association in combination with the user's query text. Based on the input context content and the user's query intention, the large language model generates a natural language answer that meets the user's needs through deep semantic matching and logical reasoning.
[0037] In a second aspect, the present application provides a large model-based car instruction manual parsing, voice broadcasting, and instruction control system, adopting the following technical solutions:
[0038] A large model-based car instruction manual parsing, voice broadcasting, and instruction control system includes:
[0039] An instruction manual document parsing module for uploading a car instruction manual document and parsing it, converting the parsed car instruction manual document into a high-dimensional vector representation and storing it in a vector database;
[0040] A vector database retrieval module for performing vector retrieval in the vector database in response to a user's query request and generating a retrieval result;
[0041] A retrieval result optimization module for inputting the retrieval result into a large language model to generate a natural language answer that meets the user's needs;
[0042] A retrieval result feedback module for calling voice synthesis technology to play the natural language answer, identifying whether there is a car machine control command. If not, it plays directly. If so, it determines whether to execute the car machine control command in response to the user's feedback.
[0043] In a third aspect, the present application provides an electronic device, the device includes a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement a large model-based car instruction manual parsing, voice broadcasting, and instruction control method as described in the first aspect.
[0044] In a fourth aspect, the present application provides a computer-readable storage medium, in which a program is stored. When the program is executed by a processor, it is used to implement a large model-based automobile manual parsing, voice broadcasting and command control method as described in the first aspect.
[0045] In summary, the beneficial effects of this application include at least:
[0046] 1) Through customized algorithms and automated parsing functions (covering system-customized algorithm parsing documents and system-automatic parsing documents), this solution can perform full-process preprocessing, structure recognition, content extraction, and instruction binding for automobile manual documents. This process effectively replaces the tedious work of traditional manual sorting, significantly reduces manpower input, improves processing efficiency and human effectiveness, and ensures the consistency and accuracy of parsing, thereby providing a high-quality data foundation for subsequent intelligent retrieval and interaction.
[0047] 2) Using the retrieval technology based on the vector library, documents and user queries are converted into high-dimensional vectors, and cosine similarity is used for approximate nearest neighbor retrieval, which can quickly locate the content most relevant to the query intent in the massive data. By optimizing the retrieval process through multi-threaded parallel computing technology, this solution significantly improves the question retrieval capability, effectively solves the problem of incomplete or incorrect retrieval caused by synonyms and similar characters in traditional keyword matching, and provides users with more accurate and efficient information retrieval services.
[0048] 3) Based on the retrieval, this solution further introduces the inference answering technology based on the large model, and uses the industry-specific large language model to semantically integrate and contextualize the retrieval results, so as to generate professional and accurate natural language answers. At the same time, through the reasonable organization of command data, it is possible to bind specific chapters to the vehicle control commands, and synchronously call the vehicle instructions during the answer generation process, so that the system has the command scheduling capability. This effect not only improves the accuracy of question answers, but also greatly improves the overall intelligence level of the system and the user interaction experience, ensuring safe and reliable vehicle operation while providing intelligent answers.
[0049] Through the automated preprocessing, structured parsing and high-dimensional vector storage of automobile manual documents, the goal of converting the manual contents that traditionally rely on manual organization, are prone to errors and have weak retrieval capabilities into accurate and structured data is achieved; in implementation, the solution first uses file type recognition, lossless compression, OCR technology and unified page conversion to automatically parse PDF manuals, and removes interference information through accurate recognition of page numbers, directories, chapters and sub-chapters, thereby extracting text, images and table data, and configuring corresponding vehicle control commands for key chapters, and finally storing and converting them into high-dimensional vectors in a unified structured format such as JSON. Considering that traditional technologies mainly rely on manual operations and keyword matching, and are easily affected by synonyms or format differences, resulting in incomplete retrieval and errors, the solution further converts user query text into query vectors through the same embedding model as document vectors, and performs multi-threaded parallel retrieval based on cosine similarity in the vector database, thereby accurately matching relevant data and providing a basis for subsequent intelligent interaction. Ultimately, by integrating industry-specific large language models to generate professional natural language answers, and combining them with automatic calling and security verification of vehicle commands, the solution effectively solved key technical problems such as manual organization, weak retrieval capabilities, and insufficient intelligence, ensuring the efficiency of information updates, the accuracy of retrieval results, and the safety of vehicle operations.
[0050] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application and implement it according to the contents of the specification, the following is a detailed description of the preferred embodiments of the present application in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a schematic diagram of the overall process of the automobile manual parsing, voice broadcasting and command control method based on a large model in the embodiment of the present application.
[0052] Figure 2 It is a schematic diagram of the process of parsing the automobile manual document in the embodiment of the present application.
[0053] Figure 3 It is a structural block diagram of the automobile manual parsing, voice broadcasting and command control system based on a large model in the embodiment of the present application.
[0054] Figure 4 It is a block diagram of an electronic device for automobile manual parsing, voice broadcasting and command control based on a large model in an embodiment of the present application. DETAILED DESCRIPTION
[0055] The specific implementation methods of the present application are further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present application but are not intended to limit the scope of the present application.
[0056] Optionally, the present application uses the large model-based automobile manual parsing, voice broadcasting and command control methods provided in various embodiments as an example for use in an electronic device, where the electronic device is a terminal or a server, and the terminal can be a mobile phone, a computer, a tablet computer, etc. This embodiment does not limit the type of electronic device.
[0057] Reference Figure 1 , is a schematic diagram of the overall flow of a method for parsing, voice broadcasting and command control of automobile manuals based on a large model provided by an embodiment of the present application, the method comprising at least the following steps:
[0058] Step S101, upload and parse the automobile manual document, convert the parsed automobile manual document into a high-dimensional vector representation and store it in a vector database.
[0059] In step S101, the automobile manual document uploaded by the user is converted into searchable vectorized data through intelligent parsing and structured processing, which specifically includes sub-steps such as document preprocessing, structure recognition, content extraction, instruction binding and storage verification. Figure 2 The flowchart of parsing the automobile manual document in the embodiment of the present application includes the following sub-steps:
[0060] Step S1011, upload the automobile manual document and pre-process it.
[0061] In the implementation, the user first uploads the car manual document through the system interface and selects different parsing strategies and segmentation strategies. The uploaded car manual document is then pre-processed, and the document type is first comprehensively identified based on the file suffix, magic byte and MIME standard media type. This application mainly focuses on and processes car manuals in PDF format. For car manual documents in PDF format, the following operations are performed: first call the PDF parsing library to load the file. The PDF parsing library supports large files and complex layout processing, and losslessly compresses the car manual document; then for scanned PDFs or pages containing embedded image text, apply optical character recognition (OCR) technology to convert image text into editable text; finally, convert PDF pages into a standardized format to ensure the consistency of subsequent parsing processes.
[0062] Step S1012: Identify the page number and structure of the automobile manual document, process the header and footer of each page, and divide and identify each chapter.
[0063] Specifically, the document structure is first identified: the first page of the automobile manual document is identified as the cover, which usually contains the company logo and product name, and then the range of the directory page is marked through the chapter title list and page number alignment features, and the page after the directory page is positioned as the main text, which contains detailed information for each chapter. Then the header and footer of each page are stripped off. The header contains the chapter title and company logo, and the footer contains the page number, so as to ensure the purity of the main text content. Then based on predefined formatting rules, such as title font size, bold, numbering, etc., the chapter boundaries are divided and the chapter titles are marked. Based on the position of the chapter title, the start and end page numbers of each chapter are divided. At the same time, the sub-chapter or subsection title is further detected within each chapter to achieve a more fine-grained content division.
[0064] Step S1013: extract text, images and tables from the divided automobile manual document, and configure corresponding vehicle control commands for the parsed chapters.
[0065] In the implementation, we first use an image segmentation algorithm based on histogram analysis to distinguish the text and image areas in the column documents of the divided automobile manual documents, extract features from each text block after distinction, including font, size, color, position and other parameters, and then determine the semantic level through weighted analysis, such as title, text, and comments. Then, we integrate the text content under the same small chapter to ensure logical coherence, and subdivide the text into paragraphs and sentences according to the format and punctuation.
[0066] For the distinguished image areas, the image objects are obtained through the image extraction function of the PDF parsing library, and the extracted image objects are saved to a predefined public storage location, such as a server directory or cloud storage. After generating a unique address, it is inserted into the text content of the corresponding chapter to ensure the relevance of the text and the image.
[0067] For the tables in the car manual documents, the table detection algorithm of the table recognition tool is used to distinguish different structural elements such as titles, rows, and columns to detect and extract the table structure in the page, and the detected table types are classified into two types: one is a two-line description table, that is, a table with only two lines, the first line is an important description, and the second line is the description content. The other is a two-dimensional function table, that is, a table with multiple rows and columns, the first column is the main function, and the subsequent columns are the detailed content of the function. Finally, the table content is parsed by row, each row represents a small function, and the logical association of the data in the row is maintained. The parsed table content is integrated with the text content of the corresponding chapter to ensure the integrity of the information.
[0068] Finally, configure corresponding in-vehicle control commands for the parsed chapters, such as the windshield wiper operation, air conditioner control, etc. For example, the command bound to the chapter for turning off the windshield wiper is WiperOffCommand, and the command bound to the chapter for turning on the air conditioner is AirConditionerOnCommand. And store the configuration data as structured metadata in the system's vector database or related metadata warehouse to ensure that the corresponding in-vehicle commands can be extracted simultaneously during the vector retrieval process.
[0069] Step S1014: Create independent data objects for each chapter and store them in a structured format to generate the structured data of the vehicle instruction manual document.
[0070] In implementation, create independent data objects for each chapter, including the title, text content, picture address, table content, and bound commands. In the data object of each chapter, establish the association between the text paragraphs and the relevant picture addresses, and embed the parsed table content to maintain the relevance with the text. All chapter objects are stored in the database in a unified structured format, such as JSON. The table data adopts a multi-dimensional array storage model to ensure data accuracy and compatibility with docking requirements. Finally, generate the structured data of the vehicle instruction manual document to meet the requirements in multiple scenarios, such as vehicle function query, after-sales service system call, etc.
[0071] Step S1015: Perform post-processing and verification on the structured data of the vehicle instruction manual document, convert the structured data into a high-dimensional vector representation, and store it.
[0072] Specifically, first check the integrity of the chapter data, such as whether the text, pictures, and tables are missing, and automatically correct or manually intervene in the errors that occur during parsing, such as unrecognized table structures. Subsequently, call the vector database interface, such as Milvus, to convert the structured data into a high-dimensional vector representation, and improve the efficiency through multi-threaded parallel processing. Finally, store the vectorized data to support subsequent retrieval.
[0073] Step S102: In response to the user's query request, perform vector retrieval in the vector database and generate the retrieval results.
[0074] In implementation, first, in response to the user's query request, the query text input by the user is received. The query text includes directly input text or text converted through speech recognition. Subsequently, the same industry pre-trained embedding model as in the vectorization stage of the car instruction manual document is called to convert the query text into a query vector with the same dimension as the stored data. Finally, approximate nearest neighbor retrieval based on cosine similarity is performed in the vector database. By calculating the similarity scores between the query vector and the high-dimensional vector representations of the stored car instruction manual documents, the top K relevant data with the highest similarity are filtered out. The relevant data includes text fragments, technical parameters, bound in-vehicle control commands, etc. Finally, the retrieval results are sorted by similarity and output as the data basis for subsequent response generation and instruction invocation.
[0075] In addition, preferably, the retrieval process adopts multi-thread parallel computing technology to optimize the matching efficiency of large-scale vector data, support processing more than a thousand query requests per second, and at the same time improve the retrieval throughput through a distributed architecture to ensure the response speed and stability in high-concurrency scenarios.
[0076] Step S103: Input the retrieval results into a large language model to generate a natural language answer that meets the user's needs.
[0077] In implementation, first, the generated retrieval results, that is, the relevant data, are input into a pre-trained industry-specific large language model, and then semantic understanding and context association are combined with the user's query text. Based on the input context content and the user's query intention, the large language model generates a natural language answer that meets the user's needs through deep semantic matching and logical reasoning.
[0078] It should be noted that the industry-specific large language model in this application is centered around the Sibyl Dongfeng large model, and at the same time supports extended integration of external large models such as GPT, Doubao, Kimi, etc. In addition, it is optimized through fine-tuning of a large amount of text in the automotive field to ensure the professionalism and accuracy of the answers.
[0079] Step S104: Invoke speech synthesis technology to play the natural language answer, identify whether there is an in-vehicle control command. If not, play it directly. If so, determine whether to execute the in-vehicle control command in response to the user's feedback.
[0080] In implementation, first, the speech synthesis technology is called to convert the natural language answer generated in step S103 into a voice signal and broadcast it to the user. At the same time, it is identified whether there is a vehicle-mounted control instruction bound to the chapter in the natural language answer. If not, the natural language answer is directly played. If so, the vehicle-mounted command execution process is synchronously triggered. Specifically, the instruction is sent to the vehicle control system through the vehicle-mounted interface to drive the vehicle to perform the corresponding operation. At the same time, it is ensured that the generation of the voice response is synchronized with the instruction call to avoid interaction interruption. When broadcasting a response involving vehicle-mounted operations, such as "Now turning off the windshield wipers for you, please confirm", the user confirmation feedback mechanism is started, and the system waits for the user to input a confirmation signal through a voice instruction or the vehicle-mounted touch screen. If the user confirms, the execution of the vehicle-mounted control instruction is completed and the operation log is recorded. If the user cancels or fails to respond within the time limit, the instruction is rolled back or the execution is paused.
[0081] In addition, preferably, the system integrates a security verification mechanism to judge whether the instruction is allowed to be executed according to the predefined vehicle state (such as vehicle speed, gear) (for example, only when the vehicle is stationary is it allowed to "open the door"), and the redundant control module automatically retries or feeds back error information (such as "Failed to turn off the windshield wipers, please try again") when the instruction execution fails, ensuring operation safety and system robustness.
[0082] To sum up, through the automated preprocessing, structured parsing, and high-dimensional vector storage of the automobile instruction manual document, the goal of converting the content of the instruction manual, which was traditionally dependent on manual collation, error-prone, and had weak retrieval capabilities, into accurate and structured data is achieved; in implementation, the solution first automatically parses the PDF-format instruction manual by means of file type recognition, lossless compression, OCR technology, and unified page conversion, and strips the interference information through the accurate recognition of page numbers, tables of contents, chapters, and sub-chapters, so as to extract text, images, and table data, and configure corresponding vehicle-mounted control commands for key chapters, and finally store them in a unified structured format such as JSON and convert them into high-dimensional vectors. Considering that traditional technologies mainly rely on manual operations and keyword matching, and are vulnerable to synonyms or format differences, resulting in incomplete and incorrect retrievals, the solution further converts the user's query text into a query vector through the same embedding model as the document vector, and performs multi-threaded parallel retrieval in the vector database based on cosine similarity, so as to accurately match the relevant data and provide a basis for subsequent intelligent interaction. Finally, by integrating an industry-specific large language model to generate professional natural language answers, and combining the automatic call and security verification of vehicle-mounted commands, this solution effectively solves key technical problems such as manual collation, weak retrieval capabilities, and insufficient intelligence, ensuring the high efficiency of information update, the accuracy of retrieval results, and the safety of vehicle operations.
[0083] Figure 3This is a structural block diagram of a large-model-based automobile manual parsing, voice broadcasting and command control system provided by an embodiment of the present application. The system includes at least the following modules:
[0084] The manual document parsing module is used to upload the automobile manual document and parse it, convert the parsed automobile manual document into a high-dimensional vector representation and store it in the vector database;
[0085] A vector database retrieval module, used to perform vector retrieval in the vector database and generate retrieval results in response to a user's query request;
[0086] The search result optimization module is used to input the search results into the large language model to generate natural language answers that meet user needs;
[0087] The retrieval result feedback module is used to call the speech synthesis technology to play the natural language answer and identify whether there is a vehicle control command. If not, it will be played directly. If it exists, it will determine whether to execute the vehicle control command in response to the user's feedback.
[0088] For relevant details, refer to the above method embodiment.
[0089] Figure 4 4 is a block diagram of an electronic device provided by an embodiment of the present application. The device at least includes a processor 401 and a memory 402.
[0090] The processor 401 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 401 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 401 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 401 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 401 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
[0091] The memory 402 may include one or more computer-readable storage media, which may be non-transitory. The memory 402 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 402 is used to store at least one instruction, which is used to be executed by the processor 401 to implement the large model-based automobile manual parsing, voice broadcasting and instruction control method provided in the method embodiment of the present application.
[0092] In some embodiments, the electronic device may further optionally include: a peripheral device interface and at least one peripheral device. The processor 401, the memory 402 and the peripheral device interface may be connected via a bus or a signal line. Each peripheral device may be connected to the peripheral device interface via a bus, a signal line or a circuit board. Schematically, the peripheral devices include but are not limited to: a radio frequency circuit, a touch display screen, an audio circuit, and a power supply.
[0093] Of course, the electronic device may also include fewer or more components, which is not limited in this embodiment.
[0094] Optionally, the present application also provides a computer-readable storage medium, in which a program is stored, and the program is loaded and executed by a processor to implement the large model-based automobile manual parsing, voice broadcasting and command control method of the above method embodiment.
[0095] Optionally, the present application also provides a computer product, which includes a computer-readable storage medium, in which a program is stored, and the program is loaded and executed by a processor to implement the large model-based automobile manual parsing, voice broadcasting and command control method of the above-mentioned method embodiment.
[0096] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0097] The above embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A method for analyzing, voice broadcasting and command control of automobile manuals based on a large model, characterized in that: The method includes: Uploading a vehicle instruction manual document, parsing it, converting the parsed vehicle instruction manual document into a high-dimensional vector representation, and storing it in a vector database; In response to a user's query request, performing vector retrieval in the vector database and generating a retrieval result; Inputting the retrieval result into a large language model to generate a natural language answer that meets the user's needs; Invoking speech synthesis technology to play the natural language answer, identifying whether there is a vehicle machine control command. If not, play it directly. If so, determine whether to execute the vehicle machine control command in response to the user's feedback.
2. The method for analyzing, voice broadcasting and command control of automobile manuals based on a large model according to claim 1 is characterized in that: The uploading the vehicle instruction manual document, parsing it, converting the parsed vehicle instruction manual document into a high-dimensional vector representation, and storing it in a vector database includes: Uploading a vehicle instruction manual document and performing preprocessing on it; Identifying the page numbers and structure of the vehicle instruction manual document, processing the header and footer of each page, dividing and identifying each chapter; Extracting text, images, and tables from the divided vehicle instruction manual document, and configuring corresponding vehicle machine control commands for the parsed chapters; Creating independent data objects for each chapter and storing them in a structured format to generate structured data of the vehicle instruction manual document; Performing post-processing and verification on the structured data of the vehicle instruction manual document, converting the structured data into a high-dimensional vector representation, and storing it.
3. The method for analyzing, voice broadcasting and command control of automobile manuals based on a large model according to claim 2 is characterized in that: The uploading the vehicle instruction manual document and performing preprocessing on it includes: Comprehensively identifying the document type based on the file suffix name, magic code, and media type of the MIME standard; For a vehicle instruction manual document in PDF format, invoking a PDF parsing library to load the file and performing lossless compression on the vehicle instruction manual document; For scanned PDF or pages containing embedded picture text, applying optical character recognition technology to convert image text into editable text; Uniformly converting PDF pages into a standardized format.
4. The method for analyzing, voice broadcasting and command control of automobile manuals based on a large model according to claim 2 is characterized in that: The identifying the page numbers and structure of the vehicle instruction manual document, processing the header and footer of each page, dividing and identifying each chapter includes: Identifying the first page of the vehicle instruction manual document as the cover, marking the range of the table of contents page through the alignment feature of the chapter title list and page numbers, and positioning the pages after the table of contents page as the main text; Stripping the header and footer of each page, dividing the chapter boundaries based on predefined format rules and marking the chapter titles, and dividing the start and end page numbers of each chapter based on the positions of the chapter titles; Further detecting small chapter or subsection titles within each chapter.
5. The method for analyzing, voice broadcasting and command control of automobile manuals based on a large model according to claim 2 is characterized in that: The extracting text, images, and tables from the divided vehicle instruction manual document, and configuring corresponding vehicle machine control commands for the parsed chapters includes: Using an image segmentation algorithm based on histogram analysis to distinguish text and image regions in the divided multi-column document of the vehicle instruction manual document, extracting features from each distinguished text block, determining the semantic level through weighted analysis, and subdividing the text into paragraphs and sentences; For the distinguished image regions, obtaining image objects through the image extraction function of the PDF parsing library, saving the extracted image objects to a predefined common storage location, generating a unique address, and inserting it into the text content of the corresponding chapter; For tables in automobile manual documents, the table detection algorithm of the table recognition tool is used to distinguish different structural elements such as titles, rows, and columns to detect and extract the table structure in the page, classify the detected table types, parse the table content by row, and integrate the parsed table content with the text content of the corresponding chapter; Configure the corresponding vehicle control commands for the parsed chapters.
6. The method for analyzing, voice broadcasting and command control of automobile manuals based on a large model according to claim 1 is characterized in that: The performing vector search in the vector database and generating search results comprises: Call the same industry pre-trained embedding model used in the vectorization phase of the car manual document to convert the query text into a query vector consistent with the dimension of the stored data; Perform an approximate nearest neighbor search based on cosine similarity in the vector database, and filter out the top K relevant data with the highest similarity by calculating the similarity score between the query vector and the high-dimensional vector representation of the stored car manual document; The relevant data are sorted by similarity and output as retrieval results.
7. The method for analyzing, voice broadcasting and command control of automobile manuals based on a large model according to claim 1 is characterized in that: Inputting the search results into a large language model to generate a natural language answer that meets the user's needs includes: The generated retrieval results, i.e. the relevant data, are input into the pre-trained industry-specific large language model, and then combined with the user's query text for semantic understanding and context association. The large language model generates natural language answers that meet user needs based on the input context content and user query intent through deep semantic matching and logical reasoning.
8. An automobile instruction manual parsing, voice broadcast and instruction control system based on a large model, characterized in that, include: The manual document parsing module is used to upload the automobile manual document and parse it, convert the parsed automobile manual document into a high-dimensional vector representation and store it in a vector database; A vector database search module, used to perform vector search in the vector database and generate search results in response to a user's query request; A search result optimization module, used to input the search results into a large language model to generate a natural language answer that meets the user's needs; The retrieval result feedback module is used to call the speech synthesis technology to play the natural language answer, identify whether there is a vehicle control command, and directly play it if not. If so, determine whether to execute the vehicle control command in response to user feedback.
9. An electronic device, characterized in that, The device includes a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement a large model-based automobile manual parsing, voice broadcasting and command control method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a program, which, when executed by a processor, is used to implement a large-model-based automobile manual parsing, voice broadcasting and command control method as described in any one of claims 1 to 7.
Citation Information
Cited By
Multivariate data structured processing method, system and program product in automobile field
CN121188140A
Intelligent car book system and method based on large model
CN121327066A