Vehicle-machine conversation quality detection method and device, vehicle and program product

By employing multi-dimensional detection methods, including sensitive word detection, contextual understanding, emotion recognition, and dialogue content prediction, the problem of low accuracy in vehicle-to-machine dialogue quality detection in existing technologies has been solved, achieving more accurate vehicle-to-machine dialogue content quality detection.

CN121148366APending Publication Date: 2025-12-16FAW CAR CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511305483.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

In existing technologies, the quality detection of vehicle-to-vehicle dialogue relies solely on sensitive word detection, which makes it difficult to fully cover the actual quality of the dialogue content, resulting in low detection accuracy.

Method used

The quality of in-vehicle dialogue content is assessed from four dimensions: sensitive word detection, contextual understanding, sentiment recognition, and dialogue content predictability. Sensitive word probability, contextual information, sentiment type, and dialogue similarity are extracted using in-vehicle dialogue features to conduct multi-dimensional quality assessment.

Benefits of technology

It improves the quality detection accuracy of vehicle-to-vehicle dialogue content, comprehensively covers the actual quality of vehicle-to-vehicle dialogue content, and reduces the risk of misjudgment caused by a single factor.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121148366A_ABST
    Figure CN121148366A_ABST
Patent Text Reader

Abstract

The invention discloses an in-vehicle conversation quality detection method and device, a vehicle and a program product, and belongs to the technical field of in-vehicle conversation. The method comprises the following steps: acquiring a vehicle-mounted terminal dialogue content; carrying out preprocessing and feature extraction on the in-vehicle dialogue content to obtain in-vehicle dialogue features; carrying out recognition processing by utilizing the in-vehicle conversation characteristics and the in-vehicle conversation content to obtain a sensitive word probability, context information, an emotion type and conversation similarity of the in-vehicle conversation content; wherein the dialogue similarity represents the similarity between the in-vehicle dialogue content and the expected dialogue content; according to the sensitive word probability, the context information, the emotion type and the dialogue similarity, performing quality detection on the in-vehicle dialogue content to obtain quality detection information of the in-vehicle dialogue content; wherein the quality detection information indicates whether the conversation content of the vehicle-mounted terminal is compliant or not. The method can effectively improve the quality detection precision of the conversation content of the vehicle-mounted terminal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle-to-everything (V2X) dialogue technology, and in particular to a method, apparatus, vehicle, and program product for detecting the quality of V2X dialogue. Background Technology

[0002] Vehicle-to-everything (V2X) dialogue refers to the process of information transmission and command interaction between a user and the vehicle's infotainment system. Its core lies in controlling vehicle functions through human-machine dialogue, thereby improving driving convenience, safety, and the passenger experience. Some related technologies detect the quality of V2X dialogue content by identifying the presence of specific sensitive words, such as to check the compliance of the dialogue content. However, these technologies rely solely on this single dimension of sensitive words for quality detection, making it difficult to comprehensively cover the actual quality of the V2X dialogue content, resulting in low accuracy in quality detection. Summary of the Invention

[0003] The main objective of this application is to provide a method, apparatus, vehicle, and program product for detecting the quality of vehicle-to-machine dialogue, aiming to improve the accuracy of detecting the quality of vehicle-to-machine dialogue content.

[0004] To achieve the above objectives, one aspect of this application proposes a method for detecting the quality of vehicle-to-everything (V2X) dialogue, the method comprising: Obtain the content of the in-vehicle infotainment system conversation; The vehicle-to-machine dialogue content is preprocessed and features are extracted to obtain vehicle-to-machine dialogue features; The vehicle-mounted system dialogue features and dialogue content are used for identification processing to obtain the sensitive word probability, contextual information, sentiment type, and dialogue similarity of the dialogue content; wherein, the dialogue similarity represents the similarity between the vehicle-mounted system dialogue content and the expected dialogue content; Based on the probability of sensitive words, the contextual information, the sentiment type, and the dialogue similarity, the in-vehicle dialogue content is subjected to quality detection to obtain quality detection information of the in-vehicle dialogue content; wherein, the quality detection information indicates whether the in-vehicle dialogue content is compliant.

[0005] In some embodiments, the preprocessing and feature extraction of the vehicle-to-vehicle dialogue content to obtain vehicle-to-vehicle dialogue features includes: The in-vehicle dialogue content is preprocessed to obtain dialogue preprocessing data; Feature extraction is performed on the preprocessed dialogue data to obtain vehicle-machine vocabulary features, vehicle-machine grammatical features, and vehicle-machine semantic features as the vehicle-machine dialogue features; The vehicle-mounted system vocabulary features are used to indicate the frequency, diversity, and relevance of words in the vehicle-mounted system dialogue content; the vehicle-mounted system grammar features are used to indicate the length, type, and grammar of sentences in the vehicle-mounted system dialogue content; and the vehicle-mounted system semantic features are used to indicate the theme, emotion, intention, and context of the vehicle-mounted system dialogue content.

[0006] In some embodiments, the contextual information includes intent type and contextual contradiction probability; the step of using the vehicle-mounted dialogue features and the vehicle-mounted dialogue content for identification processing to obtain the sensitive word probability, contextual information, sentiment type, and dialogue similarity of the vehicle-mounted dialogue content includes: The vehicle-mounted system dialogue features are used to perform sensitive word identification processing to obtain the probability of the sensitive words; The vehicle-to-machine dialogue features are input into a pre-trained intent recognition model to obtain the intent type; The vehicle-to-machine dialogue features are input into a pre-trained context recognition model to obtain the context contradiction probability; The vehicle-machine dialogue features are input into a pre-trained emotion recognition model to obtain the emotion type; The similarity score is obtained by performing a similarity calculation on the in-vehicle dialogue content and the expected dialogue content.

[0007] In some embodiments, the contextual information includes intent type and contextual contradiction probability; the step of performing quality detection on the vehicle-to-machine dialogue content based on the sensitive word probability, the contextual information, the sentiment type, and the dialogue similarity to obtain quality detection information for the vehicle-to-machine dialogue content includes: The target quality score of the vehicle-mounted dialogue content is obtained by calculating the scores based on the probability of sensitive words, the type of intent, the probability of contextual contradiction, the type of emotion, and the similarity of dialogue. Based on the target quality score, the vehicle-to-machine dialogue content is subjected to quality detection to obtain the quality detection information.

[0008] In some embodiments, the step of calculating the target quality score of the in-vehicle dialogue content based on the sensitive word probability, the intent type, the contextual contradiction probability, the sentiment type, and the dialogue similarity includes: The first quality score of the vehicle-machine dialogue content is obtained by calculating the score based on the probability of the sensitive words. A second quality score for the vehicle-machine dialogue content is obtained by calculating a score based on the intent type and the contextual contradiction probability. A third quality score for the in-vehicle dialogue content is obtained by scoring based on the emotion type. A fourth quality score for the in-vehicle dialogue content is obtained by calculating a score based on the dialogue similarity. The target quality score is obtained by weighted summing of the first quality score, the second quality score, the third quality score, and the fourth quality score.

[0009] In some embodiments, the step of performing quality detection on the vehicle-to-everything (V2X) dialogue content based on the target quality score to obtain the quality detection information includes: If the target quality score is greater than the compliance threshold, the quality detection information is determined to indicate that the vehicle-mounted system dialogue content is compliant. Alternatively, if the target quality score is less than or equal to the compliance threshold, the quality detection information is determined to indicate that the vehicle-mounted system dialogue content is non-compliant.

[0010] In some embodiments, the method further includes: If the quality inspection information indicates that the vehicle-mounted system dialogue content is non-compliant, the vehicle-mounted system dialogue content is subjected to violation identification processing to obtain the violation content, violation type, and violation index of the vehicle-mounted system dialogue content; Based on the violation content, violation type, and violation index, and combined with the feedback information from the vehicle-to-vehicle dialogue content, a suggestion analysis is performed to obtain improvement suggestions for the vehicle-to-vehicle dialogue content.

[0011] To achieve the above objectives, another aspect of this application provides a vehicle-to-everything (V2X) dialogue quality detection device, the device comprising: The acquisition module is used to acquire the content of the in-vehicle system dialogue. The first processing module is used to preprocess and extract features from the vehicle-machine dialogue content to obtain vehicle-machine dialogue features; The second processing module is used to perform recognition processing using the vehicle-machine dialogue features and the vehicle-machine dialogue content to obtain the sensitive word probability, contextual information, sentiment type and dialogue similarity of the vehicle-machine dialogue content; wherein, the dialogue similarity represents the similarity between the vehicle-machine dialogue content and the expected dialogue content; The third processing module is used to perform quality detection on the vehicle-to-machine dialogue content based on the sensitive word probability, the context information, the sentiment type, and the dialogue similarity, to obtain quality detection information of the vehicle-to-machine dialogue content; wherein, the quality detection information indicates whether the vehicle-to-machine dialogue content is compliant.

[0012] To achieve the above objectives, another aspect of this application provides a vehicle, the vehicle comprising: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the above-described vehicle-to-machine dialogue quality detection method.

[0013] To achieve the above objectives, another aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the above-described vehicle-to-machine dialogue quality detection method.

[0014] According to an embodiment of this application, a method, apparatus, vehicle, and program product for vehicle-to-everything (V2X) dialogue quality detection are provided. The method involves acquiring V2X dialogue content; preprocessing and extracting features from the V2X dialogue content to obtain V2X dialogue features; using the V2X dialogue features and the V2X dialogue content for identification processing to obtain the sensitive word probability, contextual information, sentiment type, and dialogue similarity of the V2X dialogue content; wherein, dialogue similarity represents the similarity between the V2X dialogue content and the expected dialogue content; and based on the sensitive word probability, contextual information, sentiment type, and dialogue similarity, the V2X dialogue content is subjected to quality detection to obtain quality detection information of the V2X dialogue content; wherein, the quality detection information indicates whether the V2X dialogue content is compliant. According to the technical solution of this application embodiment, the quality detection of V2X dialogue content is achieved from four dimensions: sensitive word detection, contextual understanding, sentiment recognition, and dialogue content predictability, thereby effectively improving the accuracy of V2X dialogue content quality detection.

[0015] Other features and advantages of this application will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the description, claims and drawings. Attached Figure Description

[0016] Figure 1 This is a flowchart of a vehicle-to-machine dialogue quality testing method provided in this application; Figure 2 This is another flowchart of the vehicle-to-machine dialogue quality testing method provided in this application; Figure 3 This is a schematic diagram of the vehicle-to-machine dialogue quality testing method provided in this application; Figure 4 This is a structural diagram of the vehicle-to-machine dialogue quality detection device provided in this application; Figure 5 This is an example image of a vehicle provided in this application. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.

[0018] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”

[0019] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0021] In-vehicle dialogue refers to the process of information transmission and command interaction between a user and the in-vehicle system. Its core lies in controlling vehicle functions through human-machine dialogue, thereby improving driving convenience, safety, and the passenger experience. With the development of voice technology, voice interaction has become the mainstream interaction method for in-vehicle dialogue. Users can activate the in-vehicle system using a specific wake word. Once activated, the user can ask questions via voice. The system then uses a dialogue generation model to generate corresponding answers and outputs them in voice form, thus achieving in-vehicle dialogue. The dialogue generation model deployed in the in-vehicle system typically requires a large amount of training data. During training, the model may learn content that violates social ethics, laws and regulations, and user privacy protection requirements, leading to compliance issues. For example, the dialogue content may contain insulting language, discriminatory remarks, violent content, or leak sensitive user data.

[0022] In related technologies, the quality of in-vehicle infotainment system (IVS) dialogue content is detected by identifying the presence of specific sensitive words, such as to check the compliance of the dialogue content. However, these technologies rely solely on the single dimension of sensitive words to achieve quality detection, which makes it difficult to comprehensively cover the actual quality of the IVS dialogue content, resulting in low accuracy in quality detection.

[0023] In view of this, embodiments of this application provide a method, apparatus, vehicle, and program product for detecting the quality of vehicle-to-machine dialogue, aiming to achieve quality detection of vehicle-to-machine dialogue content from four dimensions: sensitive word detection, contextual understanding, emotion recognition, and dialogue content predictability, thereby effectively improving the accuracy of vehicle-to-machine dialogue content quality detection.

[0024] It should be noted that in all specific embodiments of this application, when processing is required based on data such as vehicle-to-vehicle dialogue content, the permission or consent of the target user is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require the acquisition of data such as vehicle-to-vehicle dialogue content, separate permission or consent from the target user is obtained through pop-ups or redirection to a confirmation page. Only after obtaining the target user's separate permission or consent is the necessary data, such as vehicle-to-vehicle dialogue content, required for the proper functioning of embodiments of this application is acquired.

[0025] First, the implementation steps of a vehicle-to-machine dialogue quality detection method provided in this application will be described in detail below with reference to the accompanying drawings.

[0026] This application provides a method for detecting the quality of vehicle-to-everything (V2X) dialogue. This method can be applied to a terminal, a server, or software running on either a terminal or server. The terminal can be a tablet, laptop, desktop computer, etc., but is not limited to these. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Furthermore, the server can be a node server in a blockchain network, but is not limited to these. Blockchain is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms.

[0027] Reference Figure 1 , Figure 1 This is a flowchart of a vehicle-to-machine dialogue quality detection method provided in this application, which may include the following steps S101-S104.

[0028] S101, retrieve the content of the in-vehicle infotainment system dialogue.

[0029] It should be noted that the in-vehicle dialogue content refers to the dialogue content between the user and the in-vehicle system, recorded in voice format. Here, the user refers to the entity engaging in the in-vehicle dialogue; in some embodiments, the user may be a real user; in other embodiments, the user may be an account.

[0030] In this step, the content of the in-vehicle dialogue is obtained, which includes questions or answers and instructions raised by the user via voice, as well as questions or answers raised by the in-vehicle system in response to the user's voice information via voice.

[0031] S102, preprocess and extract features from the vehicle-machine dialogue content to obtain vehicle-machine dialogue features.

[0032] In this step, after obtaining the vehicle-machine dialogue content, the vehicle-machine dialogue content is first preprocessed to ensure the quality of the vehicle-machine dialogue content. Then, feature extraction is performed on the preprocessed vehicle-machine dialogue content to extract feature information that can highlight the characteristics of the vehicle-machine dialogue content (such as vocabulary, grammar, semantics, etc.), thereby obtaining vehicle-machine dialogue features.

[0033] S103, using the vehicle-machine dialogue features and vehicle-machine dialogue content for identification processing, to obtain the sensitive word probability, contextual information, sentiment type and dialogue similarity of the vehicle-machine dialogue content; where, dialogue similarity represents the similarity between the vehicle-machine dialogue content and the expected dialogue content.

[0034] It should be noted that the "sensitive word probability" represents the likelihood of sensitive words appearing in the in-vehicle dialogue content. A higher sensitive word probability indicates that the dialogue content is more likely to contain sensitive words, and vice versa. Contextual information represents the context of the dialogue content, such as the intended intent and the possibility of contradictions within the context. Sentiment type represents the emotional polarity of the dialogue content, such as positive, negative, and neutral. Dialogue similarity is the semantic similarity between the in-vehicle dialogue content and the expected dialogue content. Expected dialogue content refers to the dialogue content that should occur in the in-vehicle dialogue in response to a user's question or response and instruction.

[0035] In this step, recognition processing is performed based on the features and content of the vehicle-machine dialogue, thereby obtaining four recognition results: the probability of sensitive words, contextual information, sentiment type, and dialogue similarity of the vehicle-machine dialogue content. This allows for subsequent steps to detect whether the vehicle-machine dialogue content is compliant based on multi-dimensional factors.

[0036] It is worth noting that the probability of sensitive words reflects whether there are any non-compliant or high-risk words in the in-vehicle dialogue content, thus directly indicating the compliance of the in-vehicle dialogue content. The closer the probability of sensitive words is to one, the higher the probability that the in-vehicle dialogue content contains sensitive words, and the higher the probability that the in-vehicle dialogue content is non-compliant. Conversely, the lower the probability, the lower the probability that the in-vehicle dialogue content contains sensitive words, and the lower the probability that the in-vehicle dialogue content is non-compliant.

[0037] Contextual information can reflect the reasonableness of the context of in-vehicle dialogue content. For example, it can indicate whether there is any non-compliant intent or contradictory context within the dialogue, thus indirectly indicating the compliance of the dialogue content. For instance, if a user asks, "I want to scam someone," the contextual information of the dialogue content indicates a sensitive intent, suggesting an overall tendency towards non-compliance. Besides indirectly indicating compliance, contextual information can also cover compliance situations that cannot be indicated by the probability of sensitive words. For example, if a user asks, "How do I report a scam call?", which contains the word "scam," and has a high probability of being a sensitive word, the contextual information of the dialogue content indicates a non-sensitive intent, suggesting an overall tendency towards compliance. Thus, contextual information can correct for compliance situations indicated by the probability of sensitive words, thereby improving the overall coverage of the compliance of in-vehicle dialogue content.

[0038] Emotion type reflects the subjective emotional tendency displayed in the in-vehicle dialogue content, such as negative, positive, and neutral emotions. It has an indirect correlation with the compliance of the dialogue content. For example, negative emotions such as aggression, threats, and insults are usually associated with illegal content, while positive emotions are usually associated with compliant content. Neutral emotions require consideration of other factors to comprehensively determine whether the content is compliant. Combining emotion type with other factors such as the probability of sensitive words and contextual information can achieve accurate detection of in-vehicle dialogue content.

[0039] Dialogue similarity focuses on whether the in-vehicle infotainment system (IVS) dialogue content matches the expected dialogue content, indirectly indicating the compliance of the IVS dialogue content. Under the preset scenario, the higher the dialogue similarity, the more the IVS dialogue content conforms to the standards defined by the preset scenario's business, and the higher the likelihood of the IVS dialogue content being compliant. Conversely, the lower the dialogue similarity, the less the IVS dialogue content conforms to the standards defined by the preset scenario's business, and the lower the likelihood of the IVS dialogue content being compliant.

[0040] S104. Based on the probability of sensitive words, contextual information, sentiment type, and dialogue similarity, perform quality detection on the in-vehicle dialogue content to obtain quality detection information of the in-vehicle dialogue content; wherein, the quality detection information indicates whether the in-vehicle dialogue content is compliant.

[0041] It should be noted that compliance refers to conformity with social ethics, laws and regulations, and user privacy protection requirements.

[0042] In this step, the quality of the in-vehicle dialogue content is tested based on the probability of sensitive words, contextual information, sentiment type, and dialogue similarity. The aim is to detect whether the in-vehicle dialogue content is compliant, that is, whether it complies with social ethics, laws and regulations, and user privacy protection requirements. This results in quality test information for the in-vehicle dialogue content, which indicates whether the in-vehicle dialogue content is compliant. This is how the quality test of the in-vehicle dialogue content is achieved.

[0043] Therefore, this application embodiment achieves quality detection of vehicle-to-vehicle dialogue content from four dimensions: sensitive word detection, contextual understanding, sentiment detection, and dialogue content predictability. Sensitive word detection can directly reflect the compliance of vehicle-to-vehicle dialogue content, while contextual understanding, sentiment detection, and dialogue content predictability can indirectly reflect the compliance of vehicle-to-vehicle dialogue content. Through the dual detection mechanism of direct and indirect feedback, the explicit and implicit quality of vehicle-to-vehicle dialogue content can be fully explored, thereby comprehensively covering the actual quality of vehicle-to-vehicle dialogue content and reducing the risk of misjudgment caused by a single factor. This can effectively improve the accuracy of vehicle-to-vehicle dialogue content quality detection.

[0044] The steps described above will be explained in further detail below.

[0045] In some implementations, step S102 above, which involves preprocessing and feature extraction of the vehicle-to-machine dialogue content to obtain vehicle-to-machine dialogue features, may include: The dialogue content of the vehicle-to-vehicle system is preprocessed to obtain dialogue preprocessing data; Feature extraction is performed on the preprocessed dialogue data to obtain vehicle-machine lexical features, vehicle-machine grammatical features, and vehicle-machine semantic features as vehicle-machine dialogue features.

[0046] It should be noted that the preprocessed dialogue data refers to the preprocessed vehicle-to-vehicle dialogue content; vehicle-to-vehicle lexical features are used to indicate the frequency, diversity, and relevance of words in the vehicle-to-vehicle dialogue content; vehicle-to-vehicle grammatical features are used to indicate the length, type, and grammar of sentences in the vehicle-to-vehicle dialogue content; and vehicle-to-vehicle semantic features are used to indicate the topic, sentiment, intent, and context of the vehicle-to-vehicle dialogue content.

[0047] In this embodiment, after obtaining the vehicle-to-vehicle (V2V) dialogue content, the V2V dialogue content is first preprocessed to obtain preprocessed V2V dialogue content, i.e., dialogue preprocessing data. Then, feature extraction is performed on the preprocessed V2V dialogue content to extract feature information highlighting the vocabulary, grammar, and semantic characteristics of the V2V dialogue content. This yields V2V vocabulary features, V2V grammar features, and V2V semantic features, which are then used as V2V dialogue features. Thus, this embodiment effectively improves the quality of the V2V dialogue content through preprocessing. Subsequent feature extraction fully captures feature information such as vocabulary, grammar, and semantic characteristics from the V2V dialogue content, thereby helping to improve the accuracy of subsequent recognition and processing.

[0048] Optionally, the preprocessing method can be set according to the actual situation, and this embodiment does not impose specific limitations on it.

[0049] For example, during preprocessing, firstly, the in-vehicle dialogue content undergoes speech preprocessing to obtain preprocessed in-vehicle dialogue content. This example does not specifically limit the speech preprocessing method; for example, it could be speech denoising, endpoint detection, etc. Speech denoising refers to removing background noise such as ambient sound and current noise through methods such as spectral subtraction and wavelet transform. Endpoint detection refers to identifying the start and end points of speech and removing silent segments such as pauses and silences. Then, the preprocessed in-vehicle dialogue content is converted into text format to obtain text-based in-vehicle dialogue content, i.e., in-vehicle dialogue text data. This example does not specifically limit the text conversion method; for example, it could use an existing Automatic Speech Recognition (ASR) model to convert the speech-based in-vehicle dialogue content into text format, but is not limited to this. Finally, the in-vehicle dialogue text data undergoes text preprocessing to obtain preprocessed dialogue data. This example does not specify a particular text preprocessing method. For example, text preprocessing methods could include text cleaning, text segmentation, and text vectorization. Text cleaning aims to remove noise from the in-vehicle dialogue text data, thereby improving the efficiency and accuracy of subsequent processing. Text segmentation refers to dividing continuous in-vehicle dialogue text data into multiple independent paragraphs or sentences for subsequent processing and analysis. Text vectorization involves using algorithms such as Word2Vec, GloVe, and BERT to convert the in-vehicle dialogue text data into a series of machine-readable vectors that can express the semantics of the text. These vectors can accurately capture information such as semantic characteristics in the in-vehicle dialogue text data. Thus, this example can effectively improve the quality of the in-vehicle dialogue content, thereby helping to improve the accuracy of subsequent feature extraction.

[0050] Optionally, the feature extraction method can be set according to the actual situation, and this implementation method does not impose specific limitations on it.

[0051] For example, the vehicle-mounted system's vocabulary features focus on word-level attributes of the vehicle-mounted system's dialogue content, which may include, but are not limited to, vocabulary frequency, vocabulary diversity features, and vocabulary relevance features. Vocabulary frequency indicates the frequency of word occurrences in the vehicle-mounted system's dialogue content. In this example, the frequency is obtained as the ratio of the occurrence frequency of a specific sensitive word to the total occurrence frequency of all words in the dialogue preprocessing data. Vocabulary diversity indicates the richness of the vocabulary in the vehicle-mounted system's dialogue content. In this example, the average word length of the dialogue preprocessing data is used as the vocabulary diversity feature. The average word length is the ratio of the total number of characters in the dialogue preprocessing data to the number of words in the dialogue preprocessing data, where the total number of characters is the sum of the character counts of all words in the dialogue preprocessing data. Lexical association features indicate the co-occurrence patterns among words in the vehicle-to-everything (V2X) dialogue. This example sets the window size and divides all words in the preprocessed dialogue data into several text windows based on the window size. For any two words, one is defined as the first word and the other as the second word. The number of times the first and second words appear simultaneously in the same text window is taken as the co-occurrence count. The number of times the first word appears in all text windows is taken as the individual occurrence count of the first word. The number of times the second word appears in all text windows is taken as the individual occurrence count of the second word. Then, the product of the individual occurrence counts of the first and second words is calculated. The ratio of the co-occurrence count to this product is calculated and its logarithm is taken to obtain the lexical association features between the first and second words. By traversing all words, the lexical association features among all words can be obtained.

[0052] For example, the vehicle-to-everything (V2X) grammatical features focus on sentence-level attributes of the V2X dialogue content, which may include, but are not limited to, sentence length, sentence type, and sentence grammatical features. Sentence length indicates the sentence density of the V2X dialogue content. In this example, the preprocessed dialogue data is split into multiple sentences based on periods, exclamation marks, and question marks, and the number of sentences is determined as the sentence length. Sentence type indicates the functional type of sentences in the V2X dialogue content, such as declarative sentences, interrogative sentences, imperative sentences, etc. In this example, after splitting the sentences, the type of each sentence is identified based on the punctuation marks at the end of each sentence. Sentence grammatical features indicate the syntactic structure of each sentence in the V2X dialogue content. In this example, after splitting the sentences, subject-verb-object triples are extracted from each sentence as sentence grammatical features.

[0053] For example, the vehicle-mounted system semantic features focus on the contextual semantics of the vehicle-mounted system dialogue content, indirectly reflecting the core intent, emotional tendency, and subject orientation of the dialogue content. In this example, semantic features are extracted from the preprocessed dialogue data using a pre-defined feature extraction model to obtain the vehicle-mounted system semantic features. The feature extraction model is not specifically limited; for example, it can be a Transformer model, but it is not limited to this.

[0054] In some implementations, the aforementioned contextual information may include, but is not limited to, intent type and contextual contradiction probability. Intent type represents the intent displayed in the vehicle-to-machine dialogue content, and contextual contradiction probability represents the probability that the vehicle-to-machine dialogue content contains contextual contradictions. In step S103, the vehicle-to-machine dialogue features and content are used for identification processing to obtain the sensitive word probability, contextual information, sentiment type, and dialogue similarity of the vehicle-to-machine dialogue content. This may include: Sensitive word identification is performed using the features of the vehicle-mounted system dialogue to obtain the probability of sensitive words; The in-vehicle dialogue features are input into a pre-trained intent recognition model to obtain the intent type; The in-vehicle dialogue features are input into a pre-trained context recognition model to obtain the context contradiction probability; The in-vehicle dialogue features are input into a pre-trained emotion recognition model to obtain the emotion type; The similarity score is obtained by performing a similarity calculation between the in-vehicle dialogue content and the expected dialogue content.

[0055] It should be noted that intent type includes either sensitive intent type or non-sensitive intent type. Sensitive intent type indicates that the intent displayed in the vehicle-to-everything (V2X) dialogue is illegal, while non-sensitive intent type indicates that the intent displayed in the V2X dialogue is compliant. Contextual contradiction probability represents the probability of a contextual contradiction in the V2X dialogue content. A higher contextual contradiction probability indicates that the V2X dialogue content is more likely to have a contextual contradiction, and vice versa. Sentiment type includes any one of positive sentiment type, negative sentiment type, or neutral sentiment type. Positive sentiment type indicates that the V2X dialogue content has a positive sentiment towards compliant behavior, negative sentiment type indicates that the V2X dialogue content has a negative sentiment towards compliant behavior, and neutral sentiment type indicates that the V2X dialogue content maintains a neutral sentiment towards compliant behavior.

[0056] In this embodiment, for sensitive word detection, sensitive word recognition processing is performed based on the vehicle-to-vehicle (V2V) dialogue features to determine the probability of sensitive words appearing in the V2V dialogue content. For contextual understanding, it is divided into two aspects: intent recognition and contextual contradiction detection. In intent recognition, the V2V dialogue features are input into a pre-trained intent recognition model, which identifies whether the intent displayed in the V2V dialogue content is illegal, thus obtaining the intent type. In contextual contradiction detection, the V2V dialogue features are input into a pre-trained contextual recognition model, which determines the probability of contextual contradictions in the V2V dialogue content, thus obtaining the contextual contradiction probability. For sentiment recognition, the V2V dialogue features are input into a pre-trained sentiment recognition model, which identifies the sentiment polarity displayed in the V2V dialogue content, thus obtaining the sentiment type. For dialogue content anticipation, similarity calculation is performed between the V2V dialogue content and the expected dialogue content to obtain the dialogue similarity. Thus, this embodiment can quickly and accurately determine the sensitive word probability, contextual information, sentiment type, and dialogue similarity of the V2V dialogue content, improving the efficiency and accuracy of acquiring each factor.

[0057] For example, during sensitive word identification, the in-vehicle dialogue features can be input into a pre-trained sensitive word identification model, which then outputs the probability of sensitive words. The pre-trained sensitive word identification model is trained using several in-vehicle dialogue feature samples and the corresponding sensitive word labels for each sample. The sensitive word label indicates whether a sensitive word exists; a label of one indicates the presence of a sensitive word, otherwise, it indicates the absence of a sensitive word. In this way, the probability of sensitive words appearing in the in-vehicle dialogue content can be quickly determined through model identification, improving the accuracy of sensitive word recognition.

[0058] This example does not specifically limit the type of sensitive word recognition model. For example, the sensitive word recognition model can be a traditional machine learning model such as Support Vector Machine (SVM) or Logistic Regression (LR), or a deep learning model such as Transformer or Convolutional Neural Network (CNN), but it is not limited to these.

[0059] Support vector machines (SVMs) typically contain linear and nonlinear kernel functions. Linear SVMs are usually used to describe a hyperplane that separates sample points of different classes. The equation of the hyperplane can be expressed as: , The weight vector determines the vector of the hyperplane; The input feature vector represents the coordinates of the sample points; The bias term represents the distance between the hyperplane and the origin. The goal of a support vector machine is to find a hyperplane that maximizes the distance to the nearest point between the two classes of samples, which can be represented as... , Let be the norm of the weight vector. For non-linearly separable data, support vector machines (SVMs) use kernel functions to map the data to a high-dimensional space, making the data linearly separable in that space. Commonly used kernel functions include linear kernels, polynomial kernels, radial basis function kernels, and sigmoid kernels. When using a kernel function, the inner product in the original space is replaced with the kernel function value, thus extending a linear SVM to a non-linear SVM.

[0060] Logistic regression typically satisfies: , This represents the sigmoid function. Represents the weight matrix. This represents the input feature vector. For output, The weight matrix and the bias term are solved by maximizing the likelihood function.

[0061] The convolution operation in a convolutional neural network typically satisfies: , Represents the convolution kernel function. The input feature vector (in matrix form). For bias terms, For activation functions such as the sigmoid function, the weights of the convolution kernel and bias terms are updated through the backpropagation algorithm.

[0062] For example, during sensitive word identification, the in-vehicle dialogue features can be input into a pre-trained sensitive word identification model. The sensitive word identification model outputs sensitive word probabilities, and dictionary matching is performed based on the in-vehicle dialogue features. If any sensitive word stored in the dictionary is found in the in-vehicle dialogue content, the sensitive word probability is set to 100%; otherwise, it is set to 0%. Subsequently, the sensitive word probabilities output by the sensitive word identification model and those obtained from dictionary matching are weighted and summed to obtain the final sensitive word probability. In this way, explicit sensitive words in the in-vehicle dialogue content can be directly identified through dictionary matching, while implicit sensitive words can be accurately mined through model identification. The final probability of the presence of sensitive words in the in-vehicle dialogue content is determined from both explicit and implicit sensitive word dimensions, thus effectively improving the accuracy of sensitive word identification. The values ​​of all sensitive word probabilities range from zero to one. Furthermore, the weights of the sensitive word probabilities output by the sensitive word recognition model and the sensitive word probabilities obtained by dictionary matching can be flexibly set according to the actual situation. The sum of the weights of the two is one. For example, the weight of the sensitive word probability output by the sensitive word recognition model can be 0.6, and the weight of the sensitive word probability obtained by dictionary matching can be 0.4, but it is not limited to this.

[0063] It should be understood that the intent recognition model is trained using several in-vehicle dialogue feature samples and the intent type labels corresponding to each sample. The context recognition model is trained using several in-vehicle dialogue feature samples and the context contradiction labels corresponding to each sample. The context contradiction label indicates whether a contextual contradiction exists; a label of one indicates a contextual contradiction exists, otherwise, no contradiction exists. The emotion recognition model is trained using several in-vehicle dialogue feature samples and the emotion types corresponding to each sample.

[0064] Optionally, the types of intent recognition model, context recognition model, and emotion recognition model can be set according to the actual situation, and this implementation does not impose specific limitations on them. For example, the model type can be a traditional machine learning model such as support vector machine or logistic regression, or a deep learning model such as Transformer or convolutional neural network, but is not limited to these.

[0065] For example, during similarity calculation, firstly, the vehicle-to-everything (V2X) dialogue content is divided into several sub-dialogue contents, each containing a single question and its corresponding answer. Then, for each sub-dialogue content, the expected answer corresponding to the question in the current sub-dialogue content is searched from a preset database. The database pre-stores several preset questions and their corresponding expected answers. The answer and expected answer of the current sub-dialogue content are converted into vector form to obtain the answer vector features and expected answer vector features of the current sub-dialogue content. Subsequently, a similarity calculation is performed on the answer vector features and expected answer vector features of the current sub-dialogue content to obtain the dialogue similarity of the current sub-dialogue content. The dialogue similarity of each sub-dialogue content can be obtained by traversing the similarity of each sub-dialogue content. Finally, the dialogue similarities of each sub-dialogue content are averaged to obtain the final dialogue similarity.

[0066] Here, this example first breaks down the in-vehicle dialogue into several question-and-answer pairs. Then, for each question-and-answer pair, the expected answer is obtained through database retrieval as a reference benchmark, and similarity is calculated based on this benchmark to obtain the dialogue similarity of the question-and-answer pair. This process can eliminate interference from irrelevant context, allowing each similarity calculation to focus on a single question-and-answer pair, thereby accurately capturing the similarity between the actual answer and the expected answer of the question-and-answer pair. Finally, the average dialogue similarity of each question-and-answer pair is determined as the final dialogue similarity. This approach can comprehensively cover the dialogue similarity status of different question-and-answer pairs, thereby effectively improving the accuracy of similarity calculation.

[0067] This example does not specifically limit the type of similarity.

[0068] For example, similarity can be cosine similarity, which calculates the cosine similarity between the response vector features of the current sub-dialogue content and the expected response vector features as the dialogue similarity of the current sub-dialogue content. The value of cosine similarity ranges from zero to one. Here, cosine similarity can accurately capture the deep semantic similarity between the actual response and the expected response of the question-response pair, thereby improving the accuracy of similarity calculation.

[0069] For example, similarity can be calculated using cosine similarity and BLEU (Bilingual Evaluation Understudy) scores. This involves calculating the cosine similarity between the current sub-dialogue's response vector features and the expected response vector features, and calculating the BLEU score between these two features. The cosine similarity and BLEU score are then averaged or weighted to obtain the dialogue similarity of the current sub-dialogue. The values ​​of cosine similarity and BLEU score range from zero to one. Here, cosine similarity accurately captures the deep semantic similarity between the actual and expected responses of a question-and-answer pair, while BLEU score accurately captures the surface structural consistency. Dialogue similarity is generated based on cosine similarity and BLEU score. Cosine similarity compensates for the deep semantic similarity ignored by BLEU score, and BLEU score compensates for the surface structural consistency ignored by cosine similarity. This multi-dimensional similarity calculation provides a more comprehensive coverage of the dialogue similarity of question-and-answer pairs, thereby further improving the accuracy of similarity calculation. When using a weighted average, there are no specific restrictions on the weights of cosine similarity and BLEU score; the sum of their weights is one. For example, the weight of cosine similarity can be 0.4 and the weight of BLEU score can be 0.6, but it is not limited to these.

[0070] It should be understood that the BLEU score is a commonly used automatic evaluation metric in Natural Language Processing (NLP). Its core idea is to achieve a rapid and objective evaluation of the quality of actual text by quantifying the degree of n-gram (i.e., a sequence of n consecutive words) matching between the actual text and the expected text, combined with a penalty mechanism for excessively short texts. The BLEU score can be expressed as the following formula (1): (1); In equation (1), This represents the weight of each n-gram, with values ​​ranging from zero to one, and is usually treated as a uniform distribution. ; Represents the maximum order of an n-gram, for example, when Both the actual text and the desired text are cut into several consecutive segments of four words each. Indicates the first The cropping precision of an n-gram is typically the ratio of the number of times each n-gram in the actual text appears in the desired text to the total number of n-grams in the actual text. The full name of is Brevity Penalty, which satisfies the following formula (2). Indicates the actual text length. Indicates the shortest expected text length: (2).

[0071] In some implementations, contextual information may include, but is not limited to, intent type and contextual contradiction probability; in step S104 above, the quality detection of the vehicle-to-machine dialogue content is performed based on sensitive word probability, contextual information, sentiment type, and dialogue similarity to obtain quality detection information of the vehicle-to-machine dialogue content, which may include: The target quality score of the vehicle-machine dialogue content is calculated based on the probability of sensitive words, intent type, probability of contextual contradiction, sentiment type, and dialogue similarity. Based on the target quality score, the quality of the in-vehicle dialogue content is checked to obtain quality check information.

[0072] It should be noted that the target quality score is positively correlated with the compliance of the vehicle-to-vehicle dialogue content. That is, the higher the score, the higher the likelihood that the vehicle-to-vehicle dialogue content is compliant, and vice versa.

[0073] In this embodiment, firstly, a score is calculated based on the probability of sensitive words, intent type, probability of contextual contradiction, sentiment type, and dialogue similarity of the vehicle-to-vehicle (V2V) dialogue content. This results in a target quality score for the V2V dialogue content, indicating its compliance level. A higher score suggests a higher probability of compliance, while a lower score suggests a higher probability of non-compliance. Then, based on the target quality score, a quality inspection is performed on the V2V dialogue content to determine its compliance, thus obtaining quality inspection information that indicates whether the content is compliant. This achieves the quality inspection of the V2V dialogue content. Therefore, this embodiment effectively improves the accuracy of V2V dialogue content quality inspection by quantifying four dimensions—sensitive word detection, contextual understanding, sentiment detection, and dialogue content predictability—into a specific quality score.

[0074] In some implementations, the target quality score of the vehicle-to-everything (V2X) dialogue content is calculated based on the probability of sensitive words, intent type, probability of contextual contradiction, sentiment type, and dialogue similarity. This may include: The first quality score of the in-vehicle dialogue content is obtained by scoring based on the probability of sensitive words. The second quality score of the vehicle-machine dialogue content is obtained by calculating the score based on the intent type and the probability of contextual contradiction. The third quality score of the in-vehicle dialogue content is obtained by scoring based on the emotion type. The fourth quality score of the in-vehicle dialogue content is obtained by calculating the score based on the similarity of the dialogue. The target quality score is obtained by weighted summing of the first, second, third, and fourth quality scores.

[0075] It should be noted that each quality score is positively correlated with the compliance level of the in-vehicle dialogue content; that is, the higher the score, the higher the likelihood that the in-vehicle dialogue content is compliant, and vice versa. The quality scores range from zero to ten.

[0076] In this implementation, firstly, independent scores are calculated for four dimensions: sensitive word detection, contextual understanding, sentiment detection, and dialogue content predictability. Specifically, in the sensitive word detection dimension, the score calculation is performed based on the sensitive word probability, resulting in a first quality score for the vehicle-to-vehicle dialogue content. In the contextual detection dimension, the score calculation is performed based on both intent type and contextual contradiction probability, resulting in a second quality score. In the sentiment detection dimension, the score calculation is performed based on sentiment type, resulting in a third quality score. In the dialogue content predictability dimension, the score calculation is performed based on dialogue similarity, resulting in a fourth quality score. Then, weight values ​​are assigned to each of the four quality scores. The weight values ​​for each quality score can be flexibly set according to the actual situation; for example, the weights for the first and second quality scores are both 0.2, and the weights for the third and fourth quality scores are both 0.4, but this is not a limitation. Finally, the four weighted quality scores are summed to obtain the target quality score. Therefore, this implementation method calculates scores independently for four dimensions: sensitive word detection, contextual understanding, sentiment detection, and dialogue content predictability. This accurately determines the compliance level of the vehicle-to-machine dialogue content under multiple factors. Subsequently, the quality scores of the four dimensions are weighted and summed to obtain the final target quality score. This approach can fully explore the explicit and implicit quality of the vehicle-to-machine dialogue content, thereby comprehensively covering the actual quality of the vehicle-to-machine dialogue content, reducing the risk of misjudgment caused by a single factor, and effectively improving the accuracy of quality detection of vehicle-to-machine dialogue content.

[0077] For example, in the sensitive word detection dimension, the score value corresponding to the sensitive word probability is retrieved from the first mapping data as the first quality score. The first mapping data pre-stores several preset sensitive word probabilities and the score values ​​corresponding to each preset sensitive word probability. Also for example, in the sensitive word detection dimension, the first quality score is obtained based on the sensitive word probabilities using machine learning methods.

[0078] For example, in the context detection dimension, if the intent type is a sensitive intent type, the first sub-quality score of the vehicle-to-everything (V2X) dialogue content is set to zero; otherwise, it is set to ten. Simultaneously, the score value corresponding to the contextual contradiction probability is retrieved from the second mapping data as the second sub-quality score of the V2X dialogue content. The second mapping data pre-stores several preset contextual contradiction probabilities and their corresponding score values. Finally, the first and second sub-quality scores are weighted and summed to obtain the second quality score. Alternatively, the second quality score can be obtained by combining the intent type and contextual contradiction probability with machine learning methods.

[0079] For example, in the sentiment detection dimension, if the sentiment type is positive, the third quality score is set to ten; if the sentiment type is neutral, the third quality score is set to five; and if the sentiment type is negative, the third quality score is set to zero. As another example, the third quality score is obtained based on the sentiment type using machine learning methods.

[0080] For example, in the dialogue content predictability dimension, the score value corresponding to the dialogue similarity is retrieved from the third mapping data as the fourth quality score. The third mapping data pre-stores several preset dialogue similarities and the score values ​​corresponding to each preset dialogue similarity. Also for example, the fourth quality score is obtained based on the dialogue similarity using machine learning methods.

[0081] The format of the mapping data is not specifically limited; for example, the mapping data can be chart data or tabular data, but is not limited to these. Furthermore, the machine learning methods are not specifically limited; for example, the machine learning methods can be support vector machines, logistic regression, etc., but are not limited to these.

[0082] In some implementations, the above-mentioned quality inspection of the vehicle-to-everything (V2X) dialogue content based on the target quality score to obtain quality inspection information may include: If the target quality score is greater than the compliance threshold, the quality inspection information will be determined as indicating that the in-vehicle dialogue content is compliant. Alternatively, if the target quality score is less than or equal to the compliance threshold, the quality inspection information will be identified as indicating that the vehicle-mounted system dialogue content is non-compliant.

[0083] In this embodiment, the target quality score is compared with a compliance threshold. If the target quality score is greater than the compliance threshold, the vehicle-to-everything (V2X) dialogue content is considered compliant, and the quality detection information is determined as indicating compliance. If the target quality score is less than or equal to the compliance threshold, the V2X dialogue content is considered non-compliant, and the quality detection information is determined as indicating non-compliance. Therefore, this embodiment achieves quality detection of V2X dialogue content by comparing the target quality score with a compliance threshold, thus improving the accuracy and efficiency of V2X dialogue content quality detection.

[0084] Optionally, the compliance threshold can be set according to actual circumstances, and this implementation does not impose a specific limitation on it. For example, the compliance threshold can be eight, but it is not limited to this.

[0085] In some implementations, refer to Figure 2 The above method may further include the following steps S105-S106: S105, if the quality inspection information indicates that the content of the vehicle-to-vehicle dialogue is non-compliant, the vehicle-to-vehicle dialogue content is subjected to violation identification processing to obtain the violation content, violation type and violation index of the vehicle-to-vehicle dialogue content; S106. Based on the content, type, and index of the violation, and combined with the feedback information from the vehicle-mounted system dialogue, a suggestion analysis is conducted to obtain suggestions for improving the vehicle-mounted system dialogue content.

[0086] In this embodiment, when the quality inspection information indicates that the in-vehicle dialogue content is non-compliant, the in-vehicle content is interpreted and analyzed, and feedback and optimization suggestions are provided accordingly. This process mainly includes two parts: analysis and optimization. First, the evaluation results are interpreted to understand the compliance status of the in-vehicle dialogue content. Then, based on the interpretation results, potential violations in the in-vehicle dialogue content are identified, such as sensitive words. These issues can be further refined into three parts: the content of the violation, the type of violation, and the violation index. The violation index indicates the degree of violation; the higher the index, the higher the degree of violation, and vice versa. Then, corresponding improvement suggestions are provided for the identified issues.

[0087] Specifically, firstly, dictionary matching is performed based on the features of the vehicle-to-vehicle (V2V) dialogue to identify sensitive words as violations. Then, machine learning methods are used to determine the violation type and violation index, effectively improving the accuracy and efficiency of violation identification and processing. Next, corresponding prompts are set, such as "The current V2V dialogue content violates the rules of…, the violation type is…, the violation index is high, and the user feedback is…; please design improvement suggestions for the current state of this V2V dialogue content." Feedback information from users regarding the V2V dialogue content is also collected, reflecting the performance and existing problems of the V2V dialogue in practical applications. Finally, the prompts, feedback information, violation content, violation type, and violation index are input into a large language model for suggestion analysis, and corresponding improvement suggestions are output.

[0088] Optionally, the type of large language model can be set according to the actual situation, and this application embodiment does not limit it. For example, the large language model can be DeepSeek, ChatGPT, Claude, Gemini, etc., but is not limited to this.

[0089] For example, improvement suggestions may include adjusting the parameters and training data of the dialogue generation model in the vehicle infotainment system, or adding compliance requirements. For instance, adjusting the parameters of the dialogue generation model could involve modifying the model structure, changing the learning rate, or adding regularization terms, but is not limited to these. Similarly, adjusting the training data of the dialogue generation model in the vehicle infotainment system could involve adding new samples or deleting outdated or useless samples, but is not limited to these.

[0090] Therefore, when the quality inspection information indicates that the vehicle-to-everything (V2X) dialogue content is non-compliant, this implementation method performs violation identification processing on the V2X dialogue content to accurately locate the specific violation status of the V2X dialogue content, such as the violation content, violation type, and violation index. Subsequently, based on the specific violation status of the V2X dialogue content, combined with feedback information reflecting the performance of the V2X dialogue in actual applications and existing problems, feedback optimization suggestions are provided. This can effectively optimize the V2X dialogue capabilities of the vehicle system, improve the quality of the V2X dialogue content, and further enhance the driving convenience, safety, and passenger experience of the vehicle.

[0091] Optionally, the quality inspection information and improvement suggestions can be output as an evaluation report. The evaluation report may include, but is not limited to, the target quality score of the vehicle-to-everything (V2X) dialogue content, existing problems, and improvement suggestions, in order to facilitate subsequent analysis and optimization.

[0092] To facilitate understanding of the vehicle-to-machine dialogue quality detection method described in this application, an example of its actual application scenario is provided below. (Refer to...) Figure 3 In this example, the dialogue generation model configured in the vehicle system is any large language model. The principle of the vehicle dialogue quality detection method provided in this example is as follows: S201, Data Acquisition: The system acquires the content of dialogue between the vehicle and the infotainment system, which includes questions or answers and instructions raised by the user via voice, as well as questions or answers raised by the vehicle system in response to the user's voice information.

[0093] S202, Data Preprocessing: First, the in-vehicle dialogue content undergoes speech preprocessing such as noise reduction and endpoint detection to obtain preprocessed dialogue content. Then, an automatic speech recognition model converts the preprocessed dialogue content into text, resulting in text-based dialogue data. Finally, the text data undergoes further preprocessing such as text cleaning, segmentation, and vectorization to obtain preprocessed dialogue data.

[0094] S203, Feature Extraction: Feature extraction is performed on the preprocessed dialogue data to obtain vehicle-machine dialogue features, including lexical features, grammatical features, and semantic features. Lexical features may include, but are not limited to, word frequency, word diversity features, and word association features; grammatical features may include, but are not limited to, sentence length, sentence type, and sentence grammar features.

[0095] S204, Multi-dimensional Recognition Processing: Sensitive word identification is performed using the vehicle-machine dialogue features to obtain the sensitive word probability; the vehicle-machine dialogue features are input into a pre-trained intent recognition model to obtain the intent type; the vehicle-machine dialogue features are input into a pre-trained context recognition model to obtain the context contradiction probability; the vehicle-machine dialogue features are input into a pre-trained sentiment recognition model to obtain the sentiment type; and the similarity between the vehicle-machine dialogue content and the expected dialogue content is calculated to obtain the dialogue similarity.

[0096] S205, Score Calculation: First, a first quality score for the vehicle-to-vehicle dialogue content is calculated based on the probability of sensitive words. A second quality score is calculated based on the intent type and the probability of contextual contradiction. A third quality score is calculated based on the sentiment type. A fourth quality score is calculated based on the dialogue similarity. Then, the first, second, third, and fourth quality scores are weighted and summed to obtain the target quality score.

[0097] S206, Quality Inspection: The target quality score is compared with the compliance threshold. If the target quality score is greater than the compliance threshold, it means that the vehicle-to-vehicle dialogue content is compliant, and the quality detection information is determined to indicate that the vehicle-to-vehicle dialogue content is compliant. If the target quality score is less than or equal to the compliance threshold, it means that the vehicle-to-vehicle dialogue content is non-compliant, and the quality detection information is determined to indicate that the vehicle-to-vehicle dialogue content is non-compliant.

[0098] S207, Optimization Suggestions: When quality inspection information indicates that the in-vehicle dialogue content is non-compliant, the process begins with dictionary matching based on the dialogue features to identify sensitive words as violations. Then, a violation type and violation index are determined using machine learning methods. Next, corresponding prompts are set, and feedback on the dialogue content is collected. The prompts, feedback, violations, violation type, and violation index are then input into a large language model for suggestion analysis, which outputs corresponding improvement suggestions. Finally, the quality inspection information and improvement suggestions are output as an evaluation report. This report may include, but is not limited to, the target quality score of the in-vehicle dialogue content, existing problems, and improvement suggestions, facilitating subsequent analysis and optimization.

[0099] Furthermore, refer to Figure 4 This application also provides a vehicle-to-everything (V2X) dialogue quality detection device, which can implement the above-described V2X dialogue quality detection method. The device includes: Module 301 is used to acquire the content of the in-vehicle system dialogue; The first processing module 302 is used to preprocess and extract features from the vehicle-machine dialogue content to obtain vehicle-machine dialogue features; The second processing module 303 is used to perform recognition processing using the vehicle-machine dialogue features and vehicle-machine dialogue content to obtain the sensitive word probability, contextual information, sentiment type and dialogue similarity of the vehicle-machine dialogue content; wherein, the dialogue similarity represents the similarity between the vehicle-machine dialogue content and the expected dialogue content. The third processing module 304 is used to perform quality detection on the vehicle-to-vehicle dialogue content based on the probability of sensitive words, contextual information, sentiment type and dialogue similarity, and obtain quality detection information of the vehicle-to-vehicle dialogue content; wherein, the quality detection information indicates whether the vehicle-to-vehicle dialogue content is compliant.

[0100] The content of the above method embodiments is applicable to the device embodiments. The specific functions implemented by the device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0101] In addition, refer to Figure 5 This application also provides a vehicle, which includes: At least one processor 401; At least one memory 402 is used to store at least one program; When at least one program is executed by at least one processor 401, the at least one processor 401 implements the above-described vehicle-to-machine dialogue quality detection method.

[0102] The aforementioned vehicles can be private cars, such as sedans, sport utility vehicles (SUVs), multi-purpose vehicles (MPVs), or pickup trucks, or commercial vehicles, such as vans, buses, small trucks, or large trailers, or gasoline vehicles or new energy vehicles such as hybrid or pure electric vehicles.

[0103] The aforementioned memory 402, as a non-transitory network system, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory 402 may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory 402 may optionally include memory 402 remotely located relative to processor 401, and these remote memories 402 can be connected to processor 401 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0104] The aforementioned memory 402 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). Memory 402 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in memory 402 and called by processor 401 to execute the methods of the embodiments of this application.

[0105] The processor 401 described above can be implemented using a general-purpose central processing unit (CPU), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0106] In some embodiments, the vehicle may further include: Input / output interfaces are used to implement information input and output; The communication interface is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). The bus transmits information between various components of the device (such as processor 401, memory 402, input / output interface and communication interface); The processor 401, memory 402, input / output interface, and communication interface can communicate with each other within the device via a bus.

[0107] The content of the above method embodiments is applicable to this vehicle embodiment. The specific functions implemented in this vehicle embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0108] Finally, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described vehicle-to-machine dialogue quality detection method.

[0109] The content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0110] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0111] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0112] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0113] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0114] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0115] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0116] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0117] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0118] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0119] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0120] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0121] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A method for detecting the quality of vehicle-to-everything (V2X) dialogue, characterized in that, The method includes: Obtain the content of the in-vehicle infotainment system conversation; The vehicle-to-machine dialogue content is preprocessed and features are extracted to obtain vehicle-to-machine dialogue features; The vehicle-mounted system dialogue features and dialogue content are used for identification processing to obtain the sensitive word probability, contextual information, sentiment type, and dialogue similarity of the dialogue content; wherein, the dialogue similarity represents the similarity between the vehicle-mounted system dialogue content and the expected dialogue content; Based on the probability of sensitive words, the contextual information, the sentiment type, and the dialogue similarity, the in-vehicle dialogue content is subjected to quality detection to obtain quality detection information of the in-vehicle dialogue content; wherein, the quality detection information indicates whether the in-vehicle dialogue content is compliant.

2. The method according to claim 1, characterized in that, The preprocessing and feature extraction of the vehicle-to-vehicle dialogue content to obtain vehicle-to-vehicle dialogue features includes: The in-vehicle dialogue content is preprocessed to obtain dialogue preprocessing data; Feature extraction is performed on the preprocessed dialogue data to obtain vehicle-machine vocabulary features, vehicle-machine grammatical features, and vehicle-machine semantic features as the vehicle-machine dialogue features; The vehicle-mounted system vocabulary features are used to indicate the frequency, diversity, and relevance of words in the vehicle-mounted system dialogue content; the vehicle-mounted system grammar features are used to indicate the length, type, and grammar of sentences in the vehicle-mounted system dialogue content; and the vehicle-mounted system semantic features are used to indicate the theme, emotion, intention, and context of the vehicle-mounted system dialogue content.

3. The method according to claim 1, characterized in that, The contextual information includes intent type and contextual contradiction probability; the process of using the vehicle-to-vehicle dialogue features and the vehicle-to-vehicle dialogue content for identification processing to obtain the sensitive word probability, contextual information, sentiment type, and dialogue similarity of the vehicle-to-vehicle dialogue content includes: The vehicle-mounted system dialogue features are used to perform sensitive word identification processing to obtain the probability of the sensitive words; The vehicle-to-machine dialogue features are input into a pre-trained intent recognition model to obtain the intent type; The vehicle-to-machine dialogue features are input into a pre-trained context recognition model to obtain the context contradiction probability; The vehicle-machine dialogue features are input into a pre-trained emotion recognition model to obtain the emotion type; The similarity score is obtained by performing a similarity calculation on the in-vehicle dialogue content and the expected dialogue content.

4. The method according to claim 1, characterized in that, The contextual information includes intent type and contextual contradiction probability; the quality detection of the vehicle-to-machine dialogue content based on the sensitive word probability, the contextual information, the sentiment type, and the dialogue similarity to obtain quality detection information of the vehicle-to-machine dialogue content includes: The target quality score of the vehicle-mounted dialogue content is obtained by calculating the scores based on the probability of sensitive words, the type of intent, the probability of contextual contradiction, the type of emotion, and the similarity of dialogue. Based on the target quality score, the vehicle-machine dialogue content is subjected to quality detection to obtain the quality detection information.

5. The method according to claim 4, characterized in that, The process of calculating a target quality score for the vehicle-mounted system dialogue content based on the probability of sensitive words, the intent type, the probability of contextual contradiction, the sentiment type, and the dialogue similarity includes: The first quality score of the vehicle-machine dialogue content is obtained by calculating the score based on the probability of the sensitive words. A second quality score for the vehicle-machine dialogue content is obtained by calculating a score based on the intent type and the contextual contradiction probability. A third quality score for the in-vehicle dialogue content is obtained by scoring based on the emotion type. A fourth quality score for the in-vehicle dialogue content is obtained by calculating a score based on the dialogue similarity. The target quality score is obtained by weighted summing of the first quality score, the second quality score, the third quality score, and the fourth quality score.

6. The method according to claim 4, characterized in that, The step of performing quality detection on the vehicle-to-everything (V2X) dialogue content based on the target quality score to obtain the quality detection information includes: If the target quality score is greater than the compliance threshold, the quality detection information is determined to indicate that the vehicle-mounted system dialogue content is compliant. Alternatively, if the target quality score is less than or equal to the compliance threshold, the quality detection information is determined to indicate that the vehicle-mounted system dialogue content is non-compliant.

7. The method according to claim 1, characterized in that, The method further includes: If the quality inspection information indicates that the vehicle-mounted system dialogue content is non-compliant, the vehicle-mounted system dialogue content is subjected to violation identification processing to obtain the violation content, violation type, and violation index of the vehicle-mounted system dialogue content; Based on the violation content, violation type, and violation index, and combined with the feedback information from the vehicle-to-vehicle dialogue content, a suggestion analysis is performed to obtain improvement suggestions for the vehicle-to-vehicle dialogue content.

8. A vehicle-mounted communication quality detection device, characterized in that, include: The acquisition module is used to acquire the content of the in-vehicle system dialogue. The first processing module is used to preprocess and extract features from the vehicle-machine dialogue content to obtain vehicle-machine dialogue features; The second processing module is used to perform recognition processing using the vehicle-machine dialogue features and the vehicle-machine dialogue content to obtain the sensitive word probability, contextual information, sentiment type and dialogue similarity of the vehicle-machine dialogue content; wherein, the dialogue similarity represents the similarity between the vehicle-machine dialogue content and the expected dialogue content; The third processing module is used to perform quality detection on the vehicle-to-machine dialogue content based on the sensitive word probability, the context information, the sentiment type, and the dialogue similarity, to obtain quality detection information of the vehicle-to-machine dialogue content; wherein, the quality detection information indicates whether the vehicle-to-machine dialogue content is compliant.

9. A vehicle, characterized in that, include: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the vehicle-to-machine dialogue quality detection method as described in any one of claims 1-7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the vehicle-machine dialogue quality detection method according to any one of claims 1 to 7.