Intelligent interaction method, device and equipment of vehicle and storage medium
By acquiring the semantic information of the query intent of user interaction commands, identifying knowledge categories and retrieving driving behavior and rule data, and performing fusion analysis and voice conversion, the problem of complex interaction methods in existing platforms has been solved, realizing an intelligent vehicle safety operation platform that improves user experience and safety.
Patent Information
- Application Number
- CN202511337041.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2025-12-30
AI Technical Summary
Existing automotive safety operation platforms have a limited interaction method, requiring users to be familiar with data structures and indicator meanings, resulting in a high barrier to entry. The feedback is not intuitive, and the query results are mainly in tables, which are costly to understand. Furthermore, users cannot ask questions directly in natural language, and the question-and-answer results are inaccurate.
This paper provides a vehicle intelligent interaction method that obtains the query intent semantic information of user interaction commands, identifies knowledge categories, retrieves driving behavior data and rule data, performs fusion semantic analysis, generates speech conversion results and displays them, and supports multimodal presentation.
It improves driving safety and user experience, reduces operational complexity, enables joint analysis of vehicle behavior data and rule standards, ensures the accuracy and compliance of query results, and enhances interaction efficiency.
Smart Images

Figure CN121233748A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, specifically to an intelligent interaction method, device, equipment, and storage medium for vehicles. Background Technology
[0002] With the rapid development of new energy vehicles and intelligent connected vehicles, the scale and complexity of vehicle operation data are constantly increasing. In order to improve the level of safety supervision of operating vehicles, various regions have successively built vehicle safety operation platforms to collect, store and analyze vehicle operating status in real time.
[0003] However, existing automotive safety operation platforms still have significant shortcomings in terms of functionality and user experience: most platforms use dropdown menus, tables, or keyword searches, requiring users to understand the data structure and indicator meanings beforehand, resulting in a high barrier to entry and hindering non-professional users from quickly obtaining the information they need. Existing platforms mostly display query results in tabular form, lacking multimodal presentation methods, which increases the cost for users to understand and analyze. Summary of the Invention
[0004] This application provides an image data processing method, apparatus, and device for intelligent driving, which can significantly improve the naturalness and convenience of human-computer interaction during vehicle driving, reduce the complexity of user operation, and thus enhance driving safety and user experience.
[0005] The technical solution of this application is as follows: On the one hand, this application provides an intelligent interaction method for a vehicle, the method comprising: In response to an interactive command carrying the user's vehicle and driver query information, obtain the query intent semantic information corresponding to the vehicle and driver query information; Based on the semantic information of the query intent, the query knowledge category is identified to obtain the target category, which is used to represent the knowledge data category corresponding to the semantic information of the query intent. When the target category includes first category information and second category information, target driving behavior data matching the query intent semantic information is retrieved from the first database corresponding to the first category information and target rule data matching the query intent semantic information is retrieved from the second database corresponding to the second category information, respectively. The target rule data is used to indicate the behavior standard corresponding to the target driving behavior data. The target driving behavior data and the target rule data are subjected to fusion semantic analysis to obtain fusion query results; The fusion query result is subjected to speech conversion processing to obtain the target speech information corresponding to the fusion query result; Display the target speech information.
[0006] In a possible implementation, the method further includes: If the fusion query result includes data with visual features, display data is generated from the data with visual features to obtain the target chart information corresponding to the fusion query result; Display the target chart information.
[0007] In a possible implementation, the step of performing speech conversion processing on the fused query result to obtain the target speech information corresponding to the fused query result includes: Semantic analysis is performed on the fused query results to obtain the target semantic category, which is used to characterize the content category corresponding to the fused query results; Based on a preset mapping relationship, the target tone category corresponding to the target semantic category is determined. The preset tone mapping relationship is used to indicate the mapping relationship between multiple semantic categories and multiple tone categories. The fusion query result is input into the speech synthesis model for speech conversion to obtain the target speech information; The display of the target voice information includes: The target speech information is played in the speech tone corresponding to the target tone category.
[0008] In a possible implementation, displaying the target voice information includes: Speech feature analysis is performed on the target speech information to obtain the target phoneme sequence corresponding to the target speech information and the target prosodic parameters corresponding to each phoneme in the target phoneme sequence. The target tone parameters are used to characterize the prosodic features of the phonemes. Visual transformation is performed on the target prosodic parameters corresponding to each phoneme to obtain the mouth shape feature sequence and facial expression feature sequence corresponding to the target speech information; The digital human broadcasts the target speech information using lip movements corresponding to the lip movement feature sequence and facial expressions corresponding to the facial expression feature sequence.
[0009] In a possible implementation, when the target category includes first category information and second category information, retrieving target driving behavior data matching the query intent semantic information from a first database corresponding to the first category information and target rule data matching the query intent semantic information from a second database corresponding to the second category information, respectively, includes: Retrieve target driving behavior data that matches the semantic information of the query intent from the first database corresponding to the first category information; The target driving behavior data is parsed to obtain driving category information and behavior feature parameters corresponding to the target driving behavior. The driving category information is used to characterize the behavior category to which the target driving behavior belongs, and the behavior feature parameters are used to characterize the driving feature value of the target driving behavior. Retrieve target rule data that matches the category label and the behavioral feature parameters from the second database corresponding to the second category information.
[0010] In a possible implementation, the step of performing fusion semantic analysis on the target driving behavior data and the target rule data to obtain the fusion query result includes: Based on the target rule data, compliance detection is performed on the target driving behavior data to obtain target detection results. The target detection results are used to characterize the matching between the target driving behavior data and the behavioral standards indicated by the target rule data. The target driving behavior data, the target rule data, and the target detection results are semantically fused to obtain the fused query result.
[0011] In a possible implementation, the vehicle query information is the user's original natural language; The step of responding to the interaction command carrying the user's vehicle and driver query information and obtaining the query intent semantic information corresponding to the vehicle and driver query information includes: The vehicle and driver query information is preprocessed to obtain a target text sequence, which includes standardized text representing the vehicle and driver query information. The target text sequence is input into the intent recognition model for intent recognition to obtain the target intent category, which is used to characterize the query type to which the vehicle query information belongs; Extract target slot information corresponding to the target intent category from the target text sequence; the target slot information is used to characterize the key query information corresponding to the query type. The target intent category and the target slot information are structured and encapsulated to obtain the query intent semantic information.
[0012] On the other hand, this application also provides an image data processing device for intelligent driving, the device comprising: The query intent acquisition module is used to respond to the interactive command carrying the user's vehicle and driver query information and acquire the query intent semantic information corresponding to the vehicle and driver query information; The target category determination module is used to identify the query knowledge category based on the query intent semantic information to obtain the target category, wherein the target category is used to characterize the knowledge data category corresponding to the query intent semantic information. The retrieval module is configured to, when the target category includes first category information and second category information, retrieve target driving behavior data matching the semantic information of the query intent from the first database corresponding to the first category information and target rule data matching the semantic information of the query intent from the second database corresponding to the second category information, respectively, wherein the target rule data is used to indicate the behavioral standard corresponding to the target driving behavior data; The fusion module is used to perform fusion semantic analysis on the target driving behavior data and the target rule data to obtain fusion query results; The speech conversion module performs speech conversion processing on the fused query results to obtain the target speech information corresponding to the fused query results; The display module is used to display the target voice information.
[0013] On the other hand, this application also provides a vehicle intelligent interaction device, the device including a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the vehicle intelligent interaction method as described in any of the above embodiments.
[0014] On the other hand, this application also provides a computer-readable storage medium storing at least one instruction or at least one program, which is loaded and executed by a processor to implement the intelligent interaction method of the vehicle as described above.
[0015] The intelligent interaction method, device, equipment, and storage medium for vehicles provided in this application have the following technical effects: Using the technical solution provided in this application, the system first acquires an interaction command carrying user vehicle and driver query information and performs semantic parsing to obtain query intent semantic information. Then, based on the query intent semantic information, it identifies knowledge categories to determine the target category, which represents the scope of knowledge data corresponding to the query. When the target category includes driving behavior information and rule information, it retrieves target driving behavior data and target rule data from the corresponding databases, where the rule data indicates the compliance standards of the driving behavior data. Next, it performs fusion semantic analysis on the driving behavior data and rule data to obtain a fusion query result. Finally, it performs speech conversion on the fusion query result to obtain target speech information and displays it. Through the above process, this application can achieve joint analysis of vehicle and driver behavior data and rule standards, ensuring the accuracy and compliance of the query results. Simultaneously, it reduces driving interference through voice output, thereby improving interaction efficiency and driving safety, and enhancing the overall intelligence and user experience of the system. Attached Figure Description
[0016] To more clearly illustrate the technical solutions and advantages in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a schematic flowchart of an intelligent interaction method for a vehicle provided in an embodiment of this application; Figure 2 This is a schematic flowchart of a vehicle intelligent interaction method provided in an embodiment of this application; Figure 3 This is a schematic flowchart of a vehicle intelligent interaction method provided in an embodiment of this application; Figure 4 This is a schematic diagram of a vehicle intelligent interaction method device provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of a vehicle intelligent interaction method device provided in an embodiment of this application. Detailed Implementation
[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0019] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0020] Understandably, with the rapid development of new energy vehicles, various regions have successively built vehicle safety operation platforms to address increasingly complex vehicle safety issues. These platforms are used to collect, store, and analyze vehicle operation data in real time to strengthen operation management and risk control. However, existing vehicle safety operation platforms still have many shortcomings, such as limited interaction methods, relying heavily on drop-down menus, tables, or keyword searches, requiring users to be familiar with data structures and indicator meanings, resulting in a high barrier to entry; unintuitive feedback, with query results primarily in tables, leading to high comprehension costs; and the inability for users to directly ask questions or conduct multi-round exchanges in natural language, resulting in inaccurate question-and-answer results.
[0021] The following describes a vehicle intelligent interaction method provided by an embodiment of this application. Figure 1 This is a flowchart illustrating a vehicle intelligent interaction method provided in an embodiment of this application. It should be noted that while this specification provides method operation steps as shown in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-inventive labor. The order of steps listed in the embodiments is merely one possible execution order among many and does not represent the only execution order. In actual system or product execution, the method can be executed sequentially according to the embodiments or accompanying drawings, or in parallel (e.g., in a parallel processor or multi-threaded processing environment). Specifically, as shown... Figure 1 As shown, the method may include: S101, in response to the interactive command carrying the user's vehicle and driver query information, obtain the query intent semantic information corresponding to the vehicle and driver query information.
[0022] Specifically, interactive commands are operation requests issued by users to the system during vehicle and driver information queries. Interactive commands can be natural language input, text input, or image input. For example, an interactive command might be a voice inquiry, such as "Did I just speed?"; a text input, such as typing "Check if my driving today is compliant" on the vehicle's infotainment system or mobile terminal; or an image input, such as uploading a screenshot of the vehicle's operation or footage from a driving recorder. The query intent semantic information is a structured semantic representation obtained after semantic parsing the original vehicle and driver query information in the interactive command. The query intent semantic information includes at least the query action, the query object, and the query conditions.
[0023] In one embodiment, the user's voice input is recognized and parsed by an Automatic Speech Recognition (ASR) module and converted into text data corresponding to the voice input.
[0024] In one embodiment, the vehicle query information is the user's original natural language; step S101 includes: S1011, Preprocess the vehicle and driver query information to obtain the target text sequence, which includes standardized text representing the vehicle and driver query information.
[0025] Specifically, the preprocessing steps include word segmentation of the user's natural language input, removal of stop words, spelling or speech recognition error correction, and synonym normalization. The target text sequence is the text sequence obtained after standardization processing such as word segmentation, noise reduction, and error correction. The target text sequence may include at least the word segmentation sequence, standardized phrases, keywords after stop word removal, and synonym normalization results.
[0026] S1012, Input the target text sequence into the intent recognition model to perform intent recognition and obtain the target intent category. The target intent category is used to characterize the query type to which the vehicle query information belongs.
[0027] Specifically, the target intent category is the semantic classification label identified by the intent recognition model, which drives the extraction of target slot information. The intent recognition model is an algorithm or network that maps text to intent categories, determining the functional type of the query and limiting the retrieval scope. The intent recognition model can be BERT, RoBERTa, ERNIE, BiLSTM-CRF, lightweight CNN, or a distilled small model. The target intent category is the functional classification label of the vehicle / driving query information. Target intent categories can include speeding queries, fatigue driving, rapid acceleration statistics, speed limit rules, compliance judgments, accident replays, etc.
[0028] S1013, extract target slot information corresponding to the target intent category from the target text sequence. The target slot information is used to represent the key query information corresponding to the query type.
[0029] Specifically, based on the target intent category, the parameters required to execute the target intent category are extracted from the target text sequence by a Named Entity Recognition (NER) model. Target slot information is a set of key-value pairs extracted from the target text sequence; for example, target slot information includes "last week", "speeding", and "G5 Expressway".
[0030] S1014, the target intent category and target slot information are encapsulated in a structured manner to obtain the query intent semantic information.
[0031] Specifically, the target intent category is combined with the target slot information and encapsulated according to a predefined data structure to form standardized query intent semantic information. The predefined data structure can be JSON, XML or Protobuf.
[0032] By preprocessing, intent recognition, slot extraction, and structured encapsulation of vehicle and driver query information, the user's original natural language can be accurately transformed into structured semantic information that can be retrieved and analyzed later. This effectively eliminates interference from colloquial expressions, polysemous words, and input noise, improves the accuracy of query intent recognition and the system's intelligence level, and ensures that the retrieval results of subsequent driving behavior data and rule data are more accurate.
[0033] S102, Based on the semantic information of the query intent, the query knowledge category is identified to obtain the target category. The target category is used to represent the knowledge data category corresponding to the semantic information of the query intent.
[0034] Specifically, the intent category and slot information from the query intent semantic information are input into the category recognition module. The category recognition module can determine the knowledge data category corresponding to the query intent semantic information based on a classification model or rule mapping table. The target category includes first category information, second category information, or third category information. First category information includes driving behavior data, second category information includes rule-based data, and third category information includes maintenance-related data.
[0035] S103, when the target category includes first category information and second category information, target driving behavior data matching the semantic information of the query intent is retrieved from the first database corresponding to the first category information and target rule data matching the semantic information of the query intent is retrieved from the second database corresponding to the second category information. The target rule data is used to indicate the behavior standard corresponding to the target driving behavior data.
[0036] Specifically, if the target category includes both first category information and second category information, the control system retrieves target driving behavior data from the first database based on key semantic parameters extracted from the user's input vehicle and driver query information, such as license plate number, time range, and road type. Then, it retrieves target rule data from the second database and uses the target rule data as a benchmark to indicate whether the target driving behavior data conforms to the corresponding behavioral norms.
[0037] Specifically, the first category of information is data related to driving behavior, including speed, acceleration, steering angle, number of braking actions, driving duration, trajectory data, and CAN bus signals. This first category reflects the driver's actual driving behavior at a specific time and location. The first database stores historical driving behavior data for the current vehicle, including vehicle runtime sequence tables, event data tables, trajectory data tables, and raw sensor data. The second category of information is data related to rules or standards, including road speed limits, continuous driving time limits, regional driving rules, and corporate safety management systems. This second category serves as a reference standard for determining whether driving behavior is compliant. Target driving behavior data is a subset of driving behavior data retrieved from the first database that matches the query conditions. This includes the vehicle's speed curve within a specified time, number of rapid accelerations, and fatigue driving duration. Target rule data is a subset of rules retrieved from the second database that matches the query conditions. The second database stores behavioral standard rule data corresponding to driving behavior, such as traffic regulations, road safety specifications, industry operation management regulations, vehicle operation safety standards, driving behavior scoring criteria, and risk handling procedures. Target rule data may include speed limits for a specific road, nighttime driving restrictions, continuous driving duration standards, and regulatory clause numbers. Target rule data is used to indicate the behavioral standards corresponding to target driving behavior data.
[0038] Specifically, before step S103, the process includes comparing the target category with a preset analogy to obtain a target comparison result. If the target comparison result indicates that the target category includes first category information, target driving behavior data matching the semantic information of the query intent is retrieved from the first database corresponding to the first category information. If the target comparison result indicates that the target category includes first category information, target rule data matching the semantic information of the query intent is retrieved from the second database corresponding to the second category information. The target rule data is used to indicate the behavioral standards corresponding to the target driving behavior data. The preset analogy is a set of rules used to compare the target category with the system's predefined data category mapping relationship. The preset analogy includes multiple database types and retrieval conditions corresponding to different categories. The preset analogy is used to guide the data acquisition path, and the target comparison result is used to characterize the result information after comparing the target category with the preset analogy.
[0039] In one embodiment, a first database stores unstructured data related to the automotive safety industry, including laws, regulations, industry standards, and general knowledge. This first database includes legal texts, industry standard documents, technical white papers, policy guidelines, and other knowledge base information. The first database provides users with the legal basis and general knowledge support needed to determine the compliance of driving behavior and safety management. A second database stores structured data closely related to the monitoring of the automotive safety operation platform. This second database includes automotive safety incident alarm records, operational trend data, threat intelligence, vulnerability information, and monitoring indicators related to vehicle operating status. The second database provides users with the dynamic data support needed to analyze actual driving behavior data and assess risks.
[0040] In one embodiment, such as Figure 2 As shown, step S103 includes: S202, retrieve target driving behavior data that matches the semantic information of the query intent from the first database corresponding to the first category information.
[0041] Specifically, based on the slot conditions in the semantic information of the query intent, such as license plate number, time interval, and road type, target driving behavior data is retrieved from the first database; the first database may be a time-series data table, an event data table, or a trajectory data table.
[0042] S204. Perform data parsing processing on the target driving behavior data to obtain the driving category information and behavior feature parameters corresponding to the target driving behavior. The driving category information is used to characterize the behavior category to which the target driving behavior belongs, and the behavior feature parameters are used to characterize the driving feature value of the target driving behavior.
[0043] Specifically, the retrieved target driving behavior data is parsed to extract driving category information and behavioral feature parameters. For example, driving category information can be "speeding," "rapid acceleration," "sharp turn," or "fatigued driving," while behavioral feature parameters can be average speed, maximum speed, duration, acceleration threshold, and change in direction angle. The parsing methods for the target driving behavior data can include rule engine discrimination, feature extraction algorithms, statistical calculations, or machine learning models. The driving category information is used as the category basis for matching rule data.
[0044] S206, retrieve target rule data that matches the category label and behavioral feature parameters from the second database corresponding to the second category information.
[0045] By parsing the target driving behavior data, driving category information and behavioral feature parameters are extracted, enabling the raw data to be expressed in a structured manner. Then, by combining the category labels and behavioral feature parameters, the corresponding target rule data is retrieved from the rule database. This achieves a step-by-step matching of driving behavior, feature parameters, and rule standards, improving the compliance and interpretability of the query results. As a result, more reliable data support is provided for risk warning, compliance review, and auxiliary decision-making for driving behavior.
[0046] Specifically, when driving category information is used as a category label and behavioral feature parameters are used as constraints, the corresponding target rule data is retrieved from the second database. The target rule data is a set of rules retrieved from the second database that match the category label and behavioral feature parameters. The target rule data may include road speed limits, acceleration thresholds, maximum continuous driving time, night driving restrictions, and legal clause numbers. The target rule data is used as the standard basis for determining the compliance of behavioral data.
[0047] S104. Perform semantic analysis on the target driving behavior data and target rule data to obtain the fused query results.
[0048] In one embodiment, such as Figure 3 As shown, step S104 includes: S301, based on target rule data, performs compliance detection on target driving behavior data to obtain target detection results, which are used to characterize the matching between target driving behavior data and the behavioral standards indicated by target rule data.
[0049] Specifically, the target driving behavior data is compared with target rule data. For example, speed values are compared with road speed limits, continuous driving time is compared with regulatory limits, and the number of rapid accelerations is compared with company management thresholds. The target detection results are then output to characterize whether the target driving behavior data complies with the rules and the degree of deviation. The target detection results are compliance judgment information obtained after comparing the target driving behavior data with the target rule data. The target detection results may include the judgment conclusion, deviation value (e.g., speeding exceeding 15 km / h, duration, and the triggered rule entry number). The target detection results characterize the degree of matching between driving behavior and rules, providing a basis for judging the fused query results.
[0050] S303 performs semantic fusion processing on target driving behavior data, target rule data, and target detection results to obtain fused query results.
[0051] By conducting compliance checks on target driving behavior data and target rule data, and further combining the check results with semantic fusion processing, this embodiment can establish an interpretable correspondence between factual data and normative standards. Furthermore, by integrating behavioral evidence, rule standards, and judgment conclusions through fusion analysis, it generates fusion query results, improving the interpretability and credibility of vehicle and driver information queries, and providing a reliable basis for driving behavior supervision, risk warning, and auxiliary decision-making.
[0052] Specifically, target driving behavior data, target rule data, and target detection results are fused in a unified semantic space to form semantic links. Different behavior categories, rule clauses, and detection judgment results are integrated at the same semantic level, eliminating redundancy or conflicts. A semantic fusion model or rule engine is used to associate target detection results with target driving behavior data and target rule data. The fusion query result is a comprehensive output generated based on the target driving behavior data, target rule data, and detection results. The fusion query result can include judgment conclusions, deviation indicators, rule basis, visualized data, and voice broadcast content. The fusion query result is used for driving behavior compliance feedback, risk warnings, or decision support.
[0053] In one embodiment, after step S104, the method further includes: S402, if the fusion query result includes data with visualization features, generate display data for the data with visualization features to obtain the target chart information corresponding to the fusion query result.
[0054] S404 displays target chart information.
[0055] By identifying data with visual characteristics in the fusion query results and generating corresponding chart information, abstract numerical, time-series, or categorical statistical data can be displayed intuitively in chart form, improving users' efficiency in understanding the results and their interactive experience. At the same time, in in-vehicle or mobile terminal scenarios, chart display can assist voice broadcasting, enabling drivers or managers to quickly grasp key information, thereby improving the system's usability and ease of use.
[0056] Specifically, the fused query results are examined to determine whether they contain data with visual features. This visually appealing data can be numerical sequences, time series, or categorical statistical results. If the fused query results include data with visual features, the chart generation module is invoked. This module transforms the data with visual features into target chart information. The target chart information is used to represent the charted results output by the data generation module.
[0057] Specifically, the target chart information is presented on the user terminal interface, which can be displayed on an in-vehicle screen, a mobile device app interface, or a web interface.
[0058] In one embodiment, the target chart information includes chart type, data mapping relationship, coordinate axis, and legend. Chart type may include line chart, bar chart, pie chart, heat map, trajectory chart, etc.
[0059] In one embodiment, the chart type of the target chart information is determined based on data characteristics. When the data in the fusion query result is time-series data, the chart type is a line chart; when the data in the fusion query result is categorical statistical data, such as the number of speeding incidents on different road segments, the chart type is a bar chart; when the data in the fusion query result is proportional distribution data, such as the percentage of driving time, the chart type is a pie chart or a donut chart; when the data in the fusion query result is geospatial data, such as a set of driving trajectory points, the chart type is a trajectory map or a heat map.
[0060] In another embodiment, the chart type of the target chart information is determined based on the intent category association. If the target intent category is "driving time statistics", the chart type is a bar chart or a pie chart; if the target intent category is "driving behavior trend analysis", the chart type is a line chart; if the target intent category is "trajectory playback", the chart type is a trajectory chart.
[0061] In one example, after step S104, the system further includes the following steps: upon receiving a user's query instruction, the control system writes the query intent semantic information and the fused query results into the context cache to form the context data for the first round of query; when the user subsequently inputs statements such as "continue querying," "what about last week," or "what about regulations and standards," the control system determines that the query intent semantic information is a follow-up question based on the intent recognition model, and matches the query intent semantic information with the cached previous round of query intent; for time references or result references in the user's input, the control system parses the context data to obtain the corresponding time range or target object, fuses the new follow-up question information with the previous round of context data to generate updated query intent semantic information, and initiates a supplementary query in the corresponding database or knowledge base based on the updated query intent semantic information.
[0062] In one embodiment, after step S104, the method further includes: performing semantic detection on the fused query results or the system-generated answer content to obtain the matching degree between the fused query results and preset security conditions. The preset security conditions are used to characterize whether the content contains sensitive information, illegal expressions, privacy data, or potential risk information. Based on the review results, the fused query results are divided into different levels: safe, requiring desensitization, or requiring risk warning. Safe content can be output directly, content requiring desensitization has its information hidden or replaced before output, and risk-level content is accompanied by a warning message. For answers that do not meet the specifications, the system calls a rewriting strategy to semantically rewrite the fused query results, making the fused query results meet security and compliance requirements. The system can dynamically adjust the tone and expression of the broadcast based on the emotional intensity of the answer, thereby improving user experience and interaction credibility.
[0063] S105, perform speech conversion processing on the fusion query results to obtain the target speech information corresponding to the fusion query results.
[0064] Specifically, the system obtains structured text of the fusion query results, such as conclusions, indicators, rule bases, and suggested terms. It standardizes the pronunciation of numerical units, time, road names, place names, license plates, and abbreviations. Based on semantic categories or strategy mapping, it determines tone parameters, speech rate, pauses, and stress positions. Emergency alarms can increase speech rate and add pre-voice prompts. The standardized text and control parameters are input into the speech synthesis model to obtain an audio stream, generating target speech information. This target speech information includes audio payloads and broadcast metadata, and produces optional subtitles for synchronized display.
[0065] S106, Display target voice information.
[0066] Specifically, the timing and volume limit of the broadcast are determined based on the driving status or priority. Alarms can be interrupted by background noise and a pre-warning sound can be inserted, while informational messages are broadcast in sequence. Accompanying subtitles and key values are displayed in the screen status bar or pop-up window. The text, timestamp, tone parameters, and audio fingerprint of this broadcast are written to the log, supporting historical playback and auditing.
[0067] In one embodiment, step S105 further includes: S501, Perform semantic analysis on the fusion query results to obtain the target semantic category, which is used to characterize the content category corresponding to the fusion query results; S502, Based on the preset mapping relationship, determine the target tone category corresponding to the target semantic category. The preset tone mapping relationship is used to indicate the mapping relationship between multiple semantic categories and multiple tone categories. S503 inputs the fusion query results into the speech synthesis model for speech conversion to obtain the target speech information.
[0068] Specifically, the fusion query results undergo text semantic analysis to identify the content category expressed by the fusion query results. Based on the target semantic category, the corresponding target tone category is searched in the mapping table. For example, "speeding" is a warning; "driving time today is 3 hours" is an information prompt; and "driving behavior complies with regulations" is a compliance confirmation. The control system uses preset mapping relationships, such as a rapid tone for warnings, a calm tone for information prompts, and an affirmative tone for compliance confirmation. The text content of the fusion query results, along with the determined tone category parameters, is input into the speech synthesis model. The output target speech information has corresponding tone features, such as adjustments to speech rate, pitch, and stress position.
[0069] Accordingly, step S106 includes: Play the target speech information in the tone of voice corresponding to the target tone category.
[0070] By performing semantic analysis on the fusion query results before speech conversion, the content category is obtained, and the corresponding tone category is determined based on the preset mapping relationship. Then, the speech synthesis model is used for conversion. This embodiment can not only convert the text results into audio output, but also dynamically adjust the tone and expression of the broadcast according to different semantic scenarios. Drivers can understand information more quickly and intuitively while driving, thereby improving the interactive experience and driving safety.
[0071] In one embodiment, step S106 includes: S602, perform speech feature analysis on the target speech information to obtain the target phoneme sequence corresponding to the target speech information and the target prosodic parameters corresponding to each phoneme in the target phoneme sequence. The target mood parameters are used to characterize the prosodic features of the phonemes.
[0072] Specifically, feature extraction is performed on the target speech information. A forced aligner is used to obtain the target phoneme sequence and the start and end timestamps of each phoneme. During alignment, a dictionary or G2P (grapheme-to-phoneme) can be introduced to correct homophones, connected speech, and omissions. Prosodic parameters are extracted or fitted phoneme by phoneme from the target phoneme sequence. Prosodic parameters include fundamental frequency F(t), energy E(t), duration D, pauses, stress, or intonation. Then, target mood parameters, such as mood categories like calm, excited, serious, friendly, and urgent, are mapped to prosodic shaping coefficients. The output includes the target phoneme sequence and the timestamp of each phoneme, the target prosodic parameters for each phoneme, and the target mood parameters for the segment.
[0073] S604 performs visual transformation on the target prosodic parameters corresponding to each phoneme to obtain the lip shape feature sequence and facial expression feature sequence corresponding to the target speech information. Specifically, phoneme Pi is converted into visual phoneme Vi based on a many-to-one mapping table. A co-articulation model is used to calculate the influence weight of adjacent phonemes on the lip shape of the digital human in the current frame, forming a continuous lip shape curve. The fundamental frequency F(t) and energy E(t) are used to modulate the opening and closing amplitude and speed of the digital human's lip shape. For example, a higher energy E corresponds to an increased opening amplitude; an upward increase in the fundamental frequency F(t) corresponds to a slight upward increase in the corners of the mouth or jaw of the digital human; and stress corresponds to an increase in the contrast between closing and opening at the peak moment. For plosives, a short closed-mouth keyframe is inserted at the release point; for nasal sounds, the soft palate is raised or the lips are kept closed.
[0074] S606 uses a digital human to broadcast target speech information using lip movements corresponding to lip shape feature sequences and facial expressions corresponding to facial expression feature sequences.
[0075] Specifically, the tone parameter e and prosodic contour are mapped to facial expression action units (AUs). For example, tone parameter e represents excitement, and facial expression action units (AUs) correspond to raising eyebrows, raising the corners of the mouth, and widening the eyes; tone parameter e represents seriousness, and facial expression action units (AUs) correspond to a slight increase in frowning, a decrease in raising the corners of the mouth, and a decrease in blinking frequency. This generates an expression feature sequence E(t) and aligns it with the lip-sync sequence on the same time axis. The output lip-sync feature sequence {L(tk)} includes the skeletal parameters and corresponding timestamps for each frame; the output expression feature sequence {E(tk)} includes the action parameters and the timestamps corresponding to the action parameters.
[0076] By extracting the phoneme sequence and prosodic parameters corresponding to the speech and converting them into lip-shape feature sequences and facial expression feature sequences, the digital human can broadcast with lip-shape movements and facial expressions that are highly synchronized with the speech, thereby improving the naturalness and emotional expressiveness of the speech broadcast. At the same time, the use of a unified time axis output and smooth interpolation strategy ensures the stability and real-time performance of the animation process, and has good cross-language scalability and engineering application value.
[0077] As can be seen from the technical solutions provided in the embodiments of this specification above, in the application scenario of vehicle and driver information query, the interaction command carrying the user's vehicle and driver query information is first obtained and semantically parsed to obtain query intent semantic information; then, based on the query intent semantic information, knowledge category identification is performed to determine the target category, which is used to characterize the scope of knowledge data corresponding to the query; when the target category includes driving behavior information and rule information, target driving behavior data and target rule data are retrieved from the corresponding databases respectively, wherein the rule data is used to indicate the compliance standards of driving behavior data; further, the driving behavior data and rule data are subjected to fusion semantic analysis to obtain fusion query results; then, the fusion query results are converted into speech to obtain target speech information and displayed. Through the above process, this application can realize the joint analysis of vehicle and driver behavior data and rule standards, ensuring the accuracy and compliance of query results, while reducing driving interference through voice output, thereby improving interaction efficiency and driving safety, and enhancing the overall intelligence and user experience of the system.
[0078] This application provides an intelligent interaction device for a vehicle, such as... Figure 4 As shown, the device may include: The query intent acquisition module 410 is used to respond to an interactive command carrying the user's vehicle and driver query information and acquire the query intent semantic information corresponding to the vehicle and driver query information. The target category determination module 420 is used to identify the query knowledge category based on the query intent semantic information to obtain the target category, wherein the target category is used to characterize the knowledge data category corresponding to the query intent semantic information. The retrieval module 430 is configured to, when the target category includes first category information and second category information, retrieve target driving behavior data matching the query intent semantic information from the first database corresponding to the first category information and target rule data matching the query intent semantic information from the second database corresponding to the second category information, respectively, wherein the target rule data is used to indicate the behavior standard corresponding to the target driving behavior data; The fusion module 440 is used to perform fusion semantic analysis on the target driving behavior data and the target rule data to obtain fusion query results; The speech conversion module 450 performs speech conversion processing on the fusion query result to obtain the target speech information corresponding to the fusion query result; Display module 460 is used to display the target voice information query intent acquisition module, and is used to respond to the interactive command carrying the user's vehicle query information to acquire the query intent semantic information corresponding to the vehicle query information; In one specific embodiment, the above-described apparatus may further include: The visualization module is used to generate data display data for data with visualization features if the fused query results include data with visualization features, and obtain the target chart information corresponding to the fused query results. The chart display module is used to display target chart information.
[0079] In one specific embodiment, the above-mentioned speech conversion module 450 includes: The semantic analysis unit is used to perform semantic analysis on the fused query results to obtain the target semantic category, which is used to characterize the content category corresponding to the fused query results; The tone category determination unit is used to determine the target tone category corresponding to the target semantic category based on a preset mapping relationship. The preset tone mapping relationship is used to indicate the mapping relationship between multiple semantic categories and multiple tone categories. The speech conversion unit inputs the fused query results into the speech synthesis model for speech conversion to obtain the target speech information. In one specific embodiment, the above-mentioned display module 460 includes: The feature analysis unit is used to perform speech feature analysis on the target speech information to obtain the target phoneme sequence corresponding to the target speech information and the target prosodic parameters corresponding to each phoneme in the target phoneme sequence. The target mood parameters are used to characterize the prosodic features of the phonemes. The visual conversion unit is used to perform visual conversion on the target prosodic parameters corresponding to each phoneme to obtain the mouth shape feature sequence and facial expression feature sequence corresponding to the target speech information. The broadcasting unit is used to broadcast target speech information using a digital human with lip movements corresponding to lip shape feature sequences and facial expressions corresponding to facial expression feature sequences.
[0080] In one specific embodiment, the retrieval module 430 includes: The first matching unit is used to retrieve target driving behavior data that matches the semantic information of the query intent from the first database corresponding to the first category information; The parsing unit is used to perform data parsing processing on the target driving behavior data to obtain the driving category information and behavior feature parameters corresponding to the target driving behavior. The driving category information is used to characterize the behavior category to which the target driving behavior belongs, and the behavior feature parameters are used to characterize the driving feature values of the target driving behavior. The matching unit is used to retrieve target rule data that matches the category label and behavioral feature parameters from the second database corresponding to the second category information.
[0081] In one specific embodiment, the fusion module 440 includes: The compliance detection unit is used to perform compliance detection on the target driving behavior data based on the target rule data, and obtain the target detection result. The target detection result is used to characterize the matching between the target driving behavior data and the behavior standard indicated by the target rule data. The fusion unit is used to perform semantic fusion processing on target driving behavior data, target rule data, and target detection results to obtain fused query results.
[0082] In one specific embodiment, the query intent acquisition module 410 includes: The first acquisition unit is used to respond to the interaction command carrying the user's vehicle and driver query information and acquire the query intent semantic information corresponding to the vehicle and driver query information, including: The preprocessing unit is used to preprocess the vehicle and driver query information to obtain the target text sequence, which includes standardized text representing the vehicle and driver query information. The intent recognition unit is used to input the target text sequence into the intent recognition model to recognize the intent and obtain the target intent category. The target intent category is used to characterize the query type to which the vehicle query information belongs. The slot information extraction unit is used to extract target slot information corresponding to the target intent category from the target text sequence. The target slot information is used to represent the key query information corresponding to the query type. The encapsulation unit is used to structurally encapsulate the target intent category and target slot information to obtain query intent semantic information.
[0083] This application provides a computer device including a processor and a memory. The memory stores at least one instruction or at least one program, which is loaded and executed by the processor to implement a vehicle intelligent interaction method as provided in the above method embodiments.
[0084] Figure 5 This diagram illustrates a hardware structure of a device for implementing a vehicle intelligent interaction method provided in an embodiment of this application. The device may participate in or include the apparatus or system provided in the embodiment of this application. Figure 5As shown, device 5 may include one or more processors 502 (shown as 502a, 502b, ..., 502n in the figure) 502 (processor 502 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 504 for storing data, and a transmission device 506 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 5 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, device 5 may also include a... Figure 5 The more or fewer components shown, or having the same Figure 5 The different configurations shown.
[0085] It should be noted that the aforementioned one or more processors 502 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within device 5 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0086] The memory 504 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method in this embodiment. The processor 502 executes various functional applications and data processing by running the software programs and modules stored in the memory 504, thereby realizing the above-described intelligent interaction method for a vehicle. The memory 504 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 504 may further include memory remotely located relative to the processor 502, and these remote memories can be connected to the device 5 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0087] The transmission device 506 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of device 5. In one example, the transmission device 506 includes a network interface controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 506 may be a radio frequency (RF) module, used for wireless communication with the Internet.
[0088] The display can be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of device 5 (or mobile device).
[0089] This application embodiment also provides a computer-readable storage medium, which can be disposed in a server to store at least one instruction or at least one program related to implementing a vehicle intelligent interaction method in the method embodiment. The at least one instruction or the at least one program is loaded and executed by the processor to implement the vehicle intelligent interaction method provided in the above method embodiment.
[0090] Optionally, in this embodiment, the storage medium may be located at at least one of the multiple network servers in a computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0091] This invention also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform a vehicle intelligent interaction method provided in the various optional embodiments described above.
[0092] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are also possible or may be advantageous.
[0093] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device, equipment, and storage medium embodiments are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0094] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. A program for a vehicle intelligent interaction method can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0095] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for intelligent interaction of a vehicle, characterized in that, The method comprises: in response to an interactive instruction carrying vehicle driving query information of a user, obtaining query intent semantic information corresponding to the vehicle driving query information; based on the query intent semantic information, performing query knowledge category identification to obtain a target category, the target category being used to represent a knowledge data category corresponding to the query intent semantic information; in a case where the target category comprises first category information and second category information, respectively retrieving target driving behavior data matching the query intent semantic information from a first database corresponding to the first category information and retrieving target rule data matching the query intent semantic information from a second database corresponding to the second category information, the target rule data being used to indicate a behavior standard corresponding to the target driving behavior data; performing fusion semantic analysis on the target driving behavior data and the target rule data to obtain a fusion query result; performing voice conversion processing on the fusion query result to obtain target voice information corresponding to the fusion query result; displaying the target voice information.
2. The method of claim 1, wherein, The method further comprises: if the fusion query result comprises data with visual features, generating display data for the data with visual features to obtain target chart information corresponding to the fusion query result; displaying the target chart information.
3. The method of claim 1, wherein, The voice conversion processing on the fusion query result to obtain target voice information corresponding to the fusion query result comprises: performing semantic analysis on the fusion query result to obtain a target semantic category, the target semantic category being used to represent a content category corresponding to the fusion query result; based on a preset mapping relationship, determining a target tone category corresponding to the target semantic category, the preset tone mapping relationship being used to indicate a mapping relationship between a plurality of semantic categories and a plurality of tone categories; inputting the fusion query result into a voice synthesis model for voice conversion to obtain the target voice information; the displaying of the target voice information comprises: playing the target voice information in a voice tone corresponding to the target tone category.
4. The method according to any one of claims 1 to 3, characterized in that, The displaying of the target voice information comprises: performing voice feature analysis on the target voice information to obtain a target phoneme sequence corresponding to the target voice information and target prosody parameters corresponding to each phoneme in the target phoneme sequence, the target prosody parameters being used to represent prosody features of the phonemes; performing visual conversion on the target prosody parameters corresponding to each phoneme to obtain a mouth shape feature sequence and an expression feature sequence corresponding to the target voice information; playing back the target voice information through a digital human in a mouth shape action corresponding to the mouth shape feature sequence and an expression action corresponding to the expression feature sequence.
5. The method according to any one of claims 1-3, characterized in that, The retrieving of the target driving behavior data and the target rule data in a case where the target category comprises first category information and second category information comprises: retrieve target driving behavior data matching the query intent semantic information from a first database corresponding to the first category information; perform data analysis processing on the target driving behavior data to obtain driving category information corresponding to the target driving behavior and behavior characteristic parameters, the driving category information being used to represent a behavior category to which the target driving behavior belongs, and the behavior characteristic parameters being used to represent driving characteristic values of the target driving behavior; retrieve target rule data matching the category label and the behavior characteristic parameters from a second database corresponding to the second category information.
6. The method according to any one of claims 1-3, characterized in that, The fusion semantic analysis on the target driving behavior data and the target rule data includes: performing compliance detection on the target driving behavior data based on the target rule data to obtain a target detection result, the target detection result being used to represent matching between the target driving behavior data and behavior standards indicated by the target rule data; performing semantic fusion processing on the target driving behavior data, the target rule data, and the target detection result to obtain the fusion query result.
7. The method according to any one of claims 1-3, characterized in that, The vehicle driving query information is original user natural language; The response to the interactive instruction carrying the vehicle driving query information of the user includes: performing preprocessing on the vehicle driving query information to obtain a target text sequence, the target text sequence including standardized text representing the vehicle driving query information; inputting the target text sequence into an intent recognition model to perform intent recognition and obtain a target intent category, the target intent category being used to represent a query type to which the vehicle driving query information belongs; extracting target slot information corresponding to the target intent category from the target text sequence, the target slot information being used to represent key query information corresponding to the query type; performing structural packaging on the target intent category and the target slot information to obtain the query intent semantic information.
8. An intelligent interaction device of a vehicle, characterized in that, The apparatus includes: a query intent acquisition module configured to, in response to an interactive instruction carrying vehicle driving query information of a user, acquire query intent semantic information corresponding to the vehicle driving query information; a target category determination module configured to perform query knowledge category identification based on the query intent semantic information to obtain a target category, the target category being used to represent a knowledge data category corresponding to the query intent semantic information; a retrieval module configured to, in a case where the target category includes first category information and second category information, respectively retrieve target driving behavior data matching the query intent semantic information from a first database corresponding to the first category information and retrieve target rule data matching the query intent semantic information from a second database corresponding to the second category information, the target rule data being used to indicate behavior standards corresponding to the target driving behavior data; a fusion module configured to perform fusion semantic analysis on the target driving behavior data and the target rule data to obtain a fusion query result; and a response module configured to perform compliance detection on the target driving behavior data based on the target rule data to obtain a target detection result, the target detection result being used to represent matching between the target driving behavior data and behavior standards indicated by the target rule data. The voice conversion module performs voice conversion processing on the fusion query result to obtain target voice information corresponding to the fusion query result. The display module is configured to display the target voice information.
9. An intelligent interaction device of a vehicle, characterized by, The device includes a processor and a memory, and the memory stores at least one instruction or at least one program, which is loaded and executed by the processor to implement the intelligent interaction method of the vehicle as claimed in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction or at least one program, which is loaded and executed by the processor to implement the intelligent interaction method of the vehicle as claimed in any one of claims 1 to 7.