On-board intelligent interaction method, apparatus and system, and computer readable storage medium
Through on-board image acquisition and large language model judgment, the problem that drivers find it difficult to obtain and process roadside text information is solved, efficient screening and output of text information along the way is achieved, and driving safety is improved.
Patent Information
- Application Number
- PCT/CN2023/134099
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-24
- Publication Date
- 2025-05-30
AI Technical Summary
It is difficult for drivers to effectively obtain and process roadside text information during driving, resulting in safety hazards and difficulty in collecting information.
Images are collected through the vehicle-mounted image acquisition device, detect whether they contain text content, extract multiple frames of images, select one or more frames as the subject to be identified, and use a large language model to determine whether the text content belongs to the topic of interest to the user, and output relevant text information.
It realizes customized screening and tracking of text information along the way, reduces the cost of information acquisition and improves driving safety.
Smart Images

Figure CN2023134099_30052025_PF_FP_ABST
Abstract
Description
Vehicle-mounted intelligent interaction method, device, system and computer-readable storage medium Technical Field
[0001] The embodiments of the present disclosure relate to, but are not limited to, the field of artificial intelligence technology, and in particular to an in-vehicle intelligent interaction method, device, system, and computer-readable storage medium. Background Art
[0002] While driving, drivers' hands, feet, and even eyes are restricted, forcing them to maintain a single position for extended periods. If they're not careful, they might miss roadside textual information, such as warning signs, store names, and roadside advertisements. Furthermore, when searching for a product, drivers must constantly glance left and right to locate the desired product information among numerous license plate information with text, posing a serious safety hazard. If drivers miss important information, they may need to check their dashcam, but when dashcam footage isn't particularly clear, collecting information becomes quite cumbersome.
[0003] Summary of the Invention
[0004] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.
[0005] The present disclosure also provides an in-vehicle intelligent interaction method, including:
[0006] Acquire an image captured by a vehicle-mounted image acquisition device, and detect whether the captured image includes text content;
[0007] When the captured image includes text content, extracting multiple frames of images including the text content;
[0008] Select one or more frames of images from the multiple frames of images as the subject to be identified;
[0009] Obtaining topics of interest to the user, and inputting the topics of interest to the user and the text content of the subject to be identified into the large language model, so that the large language model determines whether the text content of the subject to be identified belongs to the topics of interest to the user;
[0010] When the text content of the subject to be identified belongs to a topic that the user is interested in, the text content of the subject to be identified is output to the user.
[0011] An embodiment of the present disclosure also provides an in-vehicle intelligent interaction device, comprising a memory; and a processor connected to the memory, wherein the memory is used to store instructions, and the processor is configured to execute the steps of the in-vehicle intelligent interaction method described in any embodiment of the present disclosure based on the instructions stored in the memory.
[0012] An embodiment of the present disclosure also provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the in-vehicle intelligent interaction method described in any embodiment of the present disclosure is implemented.
[0013] An embodiment of the present disclosure further provides an in-vehicle intelligent interaction system, comprising an in-vehicle image acquisition device and an in-vehicle intelligent interaction device as described in any embodiment of the present disclosure.
[0014] Other aspects will become apparent upon reading and understanding the drawings and detailed description.
[0015] Summary of the Figures
[0016] The accompanying drawings are intended to provide a further understanding of the technical solutions of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solutions of the present disclosure and do not constitute a limitation of the technical solutions of the present disclosure. The shapes and sizes of the components in the drawings do not reflect the actual scale and are intended only to illustrate the contents of the present disclosure.
[0017] FIG1A is a schematic diagram of a flow chart of an in-vehicle intelligent interaction method provided by an exemplary embodiment of the present disclosure;
[0018] FIG1B is a schematic diagram of a process of an image acquisition device acquiring a text image to be recognized provided by an exemplary embodiment of the present disclosure;
[0019] FIG2 is a schematic diagram of a subject text recognition process to be recognized provided by an exemplary embodiment of the present disclosure;
[0020] FIG3 is a schematic diagram of a process of filtering information according to user preference settings using a large language model provided by an exemplary embodiment of the present disclosure;
[0021] FIG4 is a schematic diagram of a process in which a large language model determines whether an online search is required based on its own knowledge base, provided by an exemplary embodiment of the present disclosure;
[0022] FIG5 is a schematic diagram of the structure of a large language model provided by an exemplary embodiment of the present disclosure;
[0023] FIG6 is a schematic diagram of a method for fine-tuning a large language model provided by an exemplary embodiment of the present disclosure;
[0024] FIG7 is a schematic flow chart of another in-vehicle intelligent interaction method provided by an exemplary embodiment of the present disclosure;
[0025] FIG8 is a schematic diagram of an application scenario provided by an exemplary embodiment of the present disclosure;
[0026] FIG9 is a schematic diagram of an interface for setting topics of interest to a user provided by an exemplary embodiment of the present disclosure;
[0027] FIG10A is a schematic diagram of the automatic wake-up process during the in-vehicle intelligent interaction process of the application scenario shown in FIG8 ;
[0028] FIG10B is a schematic diagram of a road sign text provided by an exemplary embodiment of the present disclosure;
[0029] FIG11 is a schematic diagram of the voice interaction process during the in-vehicle intelligent interaction process of the application scenario shown in FIG8 ;
[0030] FIG12 is a schematic diagram of the information sorting and packaging process during the in-vehicle intelligent interaction process of the application scenario shown in FIG8 ;
[0031] FIG13 is a schematic diagram of information packaged and sent during the information packaging process shown in FIG12;
[0032] FIG14 is a schematic diagram of another interface for setting topics of interest to a user provided by an exemplary embodiment of the present disclosure;
[0033] FIG15 is a schematic diagram of another application scenario provided by an exemplary embodiment of the present disclosure;
[0034] FIG16 is a schematic structural diagram of an in-vehicle intelligent interaction device provided by an exemplary embodiment of the present disclosure.
[0035] Details
[0036] To make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. It should be noted that, unless there is a conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other in any manner.
[0037] Unless otherwise defined, the technical or scientific terms used in the embodiments of the present disclosure should have the ordinary meaning understood by people with ordinary skills in the field to which the present disclosure belongs. The words "first", "second" and similar words used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. The words "include" or "comprising" and similar words mean that the elements or objects preceding the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects.
[0038] As shown in FIG1A , an embodiment of the present disclosure provides an in-vehicle intelligent interaction method, including:
[0039] Step 101: Acquire an image captured by a vehicle-mounted image acquisition device and detect whether the captured image contains text content;
[0040] Step 102: When the captured image includes text content, extract multiple frames of images including the text content;
[0041] Step 103: selecting one or more frames of images from the multiple frames of images as a subject to be identified;
[0042] Step 104: Obtain the topic of interest to the user, input the topic of interest to the user and the text content of the subject to be identified into a large language model (LLM), so that the large language model determines whether the text content of the subject to be identified belongs to the topic of interest to the user;
[0043] Step 105: When the text content of the subject to be identified belongs to a topic that the user is interested in, the text content of the subject to be identified is output to the user.
[0044] While driving, drivers often miss roadside text messages, such as warning signs, store names, and roadside advertisements. This disclosure provides an in-vehicle intelligent interaction method that primarily addresses the issue of collecting information along a user's route. By presetting topics of interest, users can filter relevant text messages, enabling customized filtering and tracking of route information. This reduces the cost of obtaining information for users and improves driving safety for drivers.
[0045] In the embodiments of the present disclosure, the subject to be identified may be a warning road sign, a store name, a roadside advertisement, a scenic spot name, a community name, a hotel name, etc. For example, when a vehicle passes through a scenic spot, the on-board image acquisition device may capture an image containing the name of the scenic spot, and the captured image containing the name of the scenic spot may be used as the subject to be identified; when a vehicle passes through a hotel, the on-board image acquisition device may capture an image containing the name of the hotel, and the captured image containing the name of the hotel may also be used as the subject to be identified, and so on.
[0046] In some exemplary embodiments, the vehicle-mounted image acquisition device may be a driving recorder or any other image acquisition device.
[0047] A driving recorder is typically used to record images, sounds, and other related information while a vehicle is driving. The present invention detects images captured by the driving recorder without the need for additional accessories.
[0048] In some exemplary embodiments, the in-vehicle intelligent interaction method of the embodiments of the present disclosure can be implemented through a driving computer, which is also called an on-board computer. It is equivalent to the brain of the car. It can read various data of the car (including images collected by the driving recorder) and display key data through an electronic display screen.
[0049] In some exemplary embodiments, in step 101, some text detection technology may be used to detect whether the captured image contains text content. In this case, the text area may be located and marked in the captured image (in this case, text recognition is not necessary, but text recognition is also possible, and this disclosure is not limited to this). Exemplary text detection technologies may include edge detection-based methods, connected region-based methods, sliding window-based methods, deep learning-based methods, and the like.
[0050] In some exemplary embodiments, the collected images may be preprocessed to improve the detection effect. Exemplarily, the preprocessing may include: image scaling, grayscale conversion, binarization, denoising, etc.
[0051] In some exemplary embodiments, when the captured image includes text content, extracting multiple frames of images including the text content includes:
[0052] Determine the time when the captured image is first detected to include text content as the entry time t0, and the time when the captured image is first not detected to include text content after t0 as the exit time t1;
[0053] Determine the number of frames of the extracted image and the frame extraction time according to the current vehicle speed;
[0054] Multiple frames of images are extracted between t0 and t1 according to the number of frames of the extracted image and the frame extraction time.
[0055] On-board images (i.e., images taken by a driving recorder) will be affected by vehicle speed. When the vehicle is traveling at a faster speed, the on-board images may appear blurry, jittery, or distorted. Therefore, in the design and use of the on-board imaging system, it is necessary to consider the impact of vehicle speed on image quality and take corresponding measures to ensure the clarity and stability of the image. The text content in the on-board image is dynamically amplified with the shooting time. When the distance between the subject to be identified (including warning signs, store names, roadside advertisements, etc.) and the car is 0, the subject to be identified in the on-board image basically disappears, and the text content also disappears. The present disclosure determines the number of frames of the extracted image based on the current vehicle speed, and takes the time when the subject text to be identified is detected as the entry time t0, and the time when the subject text to be identified is not detected as the exit time t1. That is, the time that the subject text to be identified is on the screen is t=t1-t0.
[0056] As shown in Figure 1B, when the subject text to be recognized has just entered the picture (i.e., from the entry time to the first dotted line), recognition is inaccurate due to the distance. When the subject text to be recognized is closer (i.e., from the second dotted line to the exit time), part of the subject text to be recognized will be out of the picture, resulting in incomplete recognition. Both situations are unsuitable for text recognition. Therefore, text recognition must be performed within the valid time period, which is the period between the two dotted lines.
[0057] In some exemplary embodiments, the number of frames n of the extracted image may be v is the current vehicle speed, v max is the maximum vehicle speed, and n0 is the default number of video frames extracted.
[0058] For example, n0 can be 20, that is, 20 frames are extracted by default. Assume that the vehicle is traveling on a city road with a maximum speed v max The preset speed is 60 km / h. If the current vehicle speed is 30 km / h, the number of frames of the extracted image is Since the faster the vehicle speed, the less clear the picture, in the embodiment of the present disclosure, the number of extracted frames is proportional to the vehicle speed, that is, the faster the vehicle speed, the higher the frequency of extraction.
[0059] In some exemplary embodiments, the time of the extracted first frame image is I is between 1 and n, and t is the duration of the text content in the captured image, t=t1-t0.
[0060] For example, the time of the first frame image extracted is The time of the second frame image extracted is The time of the extracted third frame image is
[0061] In the embodiment of the present disclosure, the actual extraction time of the image frame is The frames after the time point are clearer. Frame extraction starts at the time point where the distance is relatively close. If multiple image frames are extracted, the extraction time increases slowly. The time for extracting the nth frame is:
[0062] In some exemplary embodiments, as shown in FIG2 , selecting one or more frames of images from a plurality of frames of images as a subject to be identified includes:
[0063] Determine the resolution of the extracted multiple frames of images, and filter out images with resolutions lower than the resolution threshold according to a preset resolution threshold;
[0064] Identify the text in each frame of the image and the corresponding coordinates and background color of the text;
[0065] Group texts according to their corresponding background colors;
[0066] Eliminate the images corresponding to each group of texts with low similarity and incomplete texts;
[0067] The text coordinates of each group of unremoved texts are unified, and the text in the largest text box of each group is used as the subject to be identified.
[0068] In some exemplary embodiments, similarity detection is performed according to the following formula:
[0069] The Levenshtein distance is the minimum number of edits between string S1 and string S2, S1.length is the length of string S1, and S2.length is the length of string S2.
[0070] In the disclosed embodiment, incomplete text is usually a short text in which some characters cannot be recognized due to problems such as the clarity or angle of the picture. The incomplete text can be determined by the text recognition results of multiple frames of images. It is impossible to determine whether it is incomplete text by only one frame of image. It is necessary to compare the text recognition results of multiple frames of images to determine whether it is incomplete text. Text with low similarity is usually text containing some typos or missing characters (missing characters can be filtered out by incomplete text) due to recognition errors and other reasons. When eliminating each group of texts with low similarity and images corresponding to incomplete texts, the incomplete text can be determined based on the text recognition results of multiple frames of images, and the images corresponding to the incomplete text can be eliminated. Then, the texts that are not eliminated are tested for similarity in pairs, and the texts with low similarity to other texts are eliminated. In this way, the texts that are not eliminated are long texts that are high-frequency (i.e., exist in multiple frames of images) and complete or relatively complete in the multiple frames of images.
[0071] In some exemplary embodiments, optical character recognition (OCR) technology can be used to identify the text and the corresponding coordinates in each image frame. OCR technology analyzes and recognizes image files containing textual materials to obtain text and layout information. The text and the corresponding coordinates in each image frame can also be identified based on other text recognition technologies, such as feature point-based methods or deep learning-based methods, which are not limited in the present embodiment.
[0072] Exemplarily, each frame of image uses OCR technology to identify the text and text coordinates, and uses image detection technology to identify the background color corresponding to the text (using the main color extraction algorithm, the most representative or most frequent color can be selected from the image). If there are multiple text areas, multiple text areas under the same background color can be divided into a group (as the overall information of the main text to be identified). In other exemplary embodiments, multiple text areas can also be grouped according to the identified text content. For example, when the same advertisement uses different background colors to mark the advertising content, contact number, contact address and other information, the multiple text areas can be divided into a group according to the text content usually contained in the advertisement.
[0073] Since multiple frames of images are taken at different distances, there will definitely be differences in the text recognition of the multiple frames. These differences include differences in the position coordinates of the text in the image and differences in the clarity of the text. Therefore, after acquiring multiple frames of images, the resolution of the multiple frames is first determined, and a portion of the images with lower resolution are filtered out based on the resolution threshold. Then, all the text in each frame of the image is identified through OCR technology, and the text, coordinates, and background color corresponding to each frame of the image are output; based on the background color and / or text content corresponding to the text, the text is classified and integrated to ensure that all text content of a subject to be identified is grouped together, avoiding information confusion of multiple subjects to be identified and facilitating the next step of similarity detection.
[0074] Similarity detection is performed on text with the same color threshold in different image frames. For example, in the disclosed embodiment, similarity detection can be performed using the Levenshtein distance, which is the minimum number of edits (including insertions, deletions, or substitutions) required to convert a string S1 into another string S2. Text with low similarity and incomplete text are eliminated based on the similarity. The text coordinates of the text that is not eliminated are then unified, and the text in the largest text box is used as the subject to be identified. The text coordinates are then sorted from left to right and from top to bottom to obtain the final text content of the subject to be identified.
[0075] In some exemplary embodiments, the topics of interest to the user are set in any one or more of the following ways:
[0076] Pre-stored in a storage unit;
[0077] By user input via voice and / or text;
[0078] Obtained through big data analysis.
[0079] In order to prevent the vehicle system from constantly collecting roadside text information and sending information to users without selection, the present disclosure sets an information filtering mode, and users can filter the roadside information according to their own preferences and filter out other uninteresting information.
[0080] In some examples, users can pre-store topics of interest to them in a storage unit (for example, a car memo). In this way, when the system detects that the text content of the subject to be identified belongs to the topic of interest to the user, the text content of the subject to be identified can be output to the user, thereby prompting the user.
[0081] In other examples, users can input topics of interest in real time through voice and / or text. In this way, users can set topics of interest for each trip according to their travel plans. After the system detects that the text content of the subject to be identified belongs to the topic of interest to the user, it outputs the text content of the subject to be identified to the user, thereby facilitating the user to find products and improving driving safety.
[0082] In other examples, topics of interest to users can also be based on their historical habits, such as big data analysis of their search history, browser history, click history, and product browsing history on their personal mobile phones. Through big data analysis, this information can be deeply mined to help the system continuously optimize personalized services and provide content that better suits user interests.
[0083] As shown in Figure 3, when the system detects a subject to be identified (using a road sign as an example), it selects the topic of interest to the user as the target and feeds the text of the subject to be identified into the large language model. The large language model, with its text classification capabilities, classifies the road sign based on its existing knowledge and determines whether the text category is included in the user's topic of interest. If not, it proceeds to the next road sign. If so, it outputs the text of the subject to be identified to the user (in the form of voice, image, document, video, etc.).
[0084] In some exemplary embodiments, topics of interest to users can be set through an information publishing system (referred to as an information publishing system). Since information publishing systems generally have functions such as subscription, identification, classification, push, storage, query, and sending of information, topics of interest to users can be set through the information publishing system. When the system detects that the target text belongs to the topic of interest to the user, the target text can be automatically pushed to the user, thereby reducing the cost of the user obtaining information.
[0085] The in-vehicle intelligent interaction method provided by the embodiments of the present disclosure is mainly used for in-vehicle information collection and intelligent question-answering. The system can include the following two voice wake-up functions, and users can disable the automatic wake-up function as needed:
[0086] (1) When the user actively asks, the user can wake up the system through the wake-up word;
[0087] (2) Without the user's active inquiry, the system automatically wakes up when it retrieves the target text.
[0088] In some exemplary embodiments, outputting the text content of the subject to be recognized to the user includes at least one of the following:
[0089] Input the text content of the subject to be identified into the large language model, so that the large language model generates a summary based on the text content of the subject to be identified and broadcasts the generated summary to the user;
[0090] Output an image containing the subject to be identified and a label to the user, where the label is the category of the subject to be identified;
[0091] Organize the text content of the subject to be identified into a document and send it to the user;
[0092] Output the video containing the subject to be recognized to the user.
[0093] In the disclosed embodiments, a variety of output methods can be used to output the text content of the subject to be identified to the user. For example, a large language model can be used to integrate scattered text into a smooth summary, and the information can be broadcast to the user by voice. Alternatively, an image or video output method can be used to output an image or video containing the subject to be identified to the user through the vehicle display screen. Each frame of the output image or video can be assigned a label corresponding to the category of the subject to be identified. The user can use the output image or video to generate a travel diary. In addition, the text content of the subject to be identified collected along the way can be organized into a Word, PDF, Excel, or other document and sent to the user.
[0094] In one exemplary application scenario, when using image output, key target areas can be highlighted in the output vehicle display using detection frames, allowing users to more easily and clearly focus on the text content they are interested in. For example, suppose a user is driving and looking for a hotel, and the user's focus is on the topic of "nearby hotels." The system can collect hotel information along the way and scroll through the detected images at different times on the vehicle display at a certain playback speed, highlighting the hotels in the images with detection frames.
[0095] When a user is interested in a hotel, they can voice-ask questions about it, such as its ranking and its unique features. The system can then determine whether an online search is necessary based on its own knowledge and respond based on its own knowledge or the results of the online search.
[0096] In some exemplary embodiments, a large language model can be trained using pre-training technology. When the number of parameters of a language model expands to over 100 billion, pre-training a large language model from scratch becomes a very difficult and challenging task. Pre-training refers to the initial training of the model using large-scale data sets and unsupervised learning methods before the target task. During the pre-training stage, the model acquires knowledge and features by learning the internal representation of the input data so as to perform fine-tuning or transfer learning on subsequent specific tasks.
[0097] In some exemplary embodiments, the large language model can interact with the user in a zero-start manner. During this interaction, the large language model acquires or collects an in-vehicle user dataset, and trains and calibrates the large language model based on the acquired or collected in-vehicle user dataset to refine model parameters. The in-vehicle user dataset can be images captured by an image acquisition device during driving, and these images can be labeled by the user.
[0098] In some exemplary embodiments, the large language model may be located in a cloud server or in an onboard storage unit.
[0099] When the large language model resides on a cloud server, it can access vehicle user datasets from multiple different terminals. This data can be used to train and calibrate the large language model, improving model parameters. When uploading vehicle user datasets to the cloud server, different terminals can upload them in real time after the data is collected to generate the dataset, or they can first store the dataset locally and upload it when pre-set conditions are met. For example, the pre-set condition may be that the current network is WiFi, etc.
[0100] In some exemplary embodiments, as shown in FIG4 , the method further includes:
[0101] Receive questions from users;
[0102] Input the user's question into the large language model;
[0103] The large language model determines whether an online search is necessary. If so, it retrieves multiple text segments, determines the similarity between these segments and the user's question, generates summaries for the top N segments with the highest similarity, and outputs the generated summaries to the user. If an online search is not necessary, the large language model provides an answer based on its own knowledge, where N is a natural number greater than or equal to 1. For example, N can be 3.
[0104] In the disclosed embodiment, when a user's question is about the interior information of the currently driven vehicle, the large language model determines that answers to such questions do not require online search, thereby improving the accuracy and credibility of the answer. For example, the interior information of the currently driven vehicle may include: What is the current air conditioning temperature? What is the volume? How many kilometers can the fuel level remain? Such non-open-ended questions related to the interior information of the currently driven vehicle can be answered without online search.
[0105] When users are interested in detected text information, they often ask questions, such as "What is this company's local ranking?", "Who is the chairman of this hotel?", and "When was this scenic spot built?" This leads to two problems: drivers cannot manually search for questions, and the answers retrieved are often lengthy. This disclosure addresses these two issues by fine-tuning a large language model.
[0106] Fine-tune the large language model so that it can easily call the web search plug-in. The challenge here is how to make the large language model determine that the current question requires the web search plug-in. This paper borrows the ideas and corpus of the MOSS plug-in to fine-tune the large language model to do the following: 1. What type of question is this? 2. Can its existing knowledge answer it? If not, a web search is required; 3. Provide a brief summary of the retrieved text.
[0107] First, the large language model is fine-tuned to determine the nature of the question. For example, for a question like "The history of this scenic spot," the large language model generates a thought and judgment based on the question: "This is a question that requires retrieval. Search("scenic spot," "history")." Based on this, the large language model searches the network for information. The retrieved text information is all text fragments containing keywords, which inevitably contain some irrelevant information. The similarity model is used to determine the similarity between the answer and the question. The text fragments of the top three most similar answers are input into the large language model, which then briefly summarizes the content of the text fragments (i.e., generates a summary) to avoid cluttering the answers and affecting the user experience.
[0108] When generating summaries, the large language model can adapt the user's character settings to different styles. For example, if the user sets the character style to "child," the large language model will use simpler, easier-to-understand language during summary generation and add more interjections, such as "ah, oh, yeah...", to make the user sound more cheerful and endearing. If the user sets the character style to "adult," the model will use a formal style for summary generation.
[0109] During long journeys, users may feel bored and long for companionship. However, the topics raised by commercial companion robots are not likely to arouse users' interest, and they cannot have natural and smooth conversations with users. Although the in-vehicle systems currently on the market have simple voice Q&A operations and provide answers by entering them into a knowledge base, they cannot answer real-time questions raised by users, such as: "What is the average market price of houses in this area?" "When were the houses in this community built?" etc. Some in-vehicle systems can also perform information retrieval, but lack the ability to integrate information. They can only play some long-winded news articles step by step without focusing on the key points and cannot provide targeted answers based on the user's information. The disclosed embodiments can bring a lot of fun to users' journeys by collecting text information along the way, filtering topics of interest to users through a large language model, and interacting with users through voice using the large language model.
[0110] In some exemplary embodiments, the large language model of the disclosed embodiments can be constructed based on the decoder portion of a Transformer. The structure of an exemplary large language model is shown in FIG5 . Each Transformer block mainly includes three parts: an attention layer, a root mean square layer normalization (RMSNorm), and a multilayer perceptron (MLP). The attention layer adds a rotational position encoding (RoPE) to the query tensor and the key-value tensor (QK tensor). This position encoding formally relies on the absolute position encoding and can be converted into a relative position encoding during calculation, which is very helpful for learning the contextual semantics of long texts.
[0111] A multilayer perceptron (MLP), also known as a feedforward neural network (FFN), consists of two fully connected layers and a nonlinear activation function. The FFN layer needs to learn two linear transformations, separated by a nonlinear activation function. The mathematical expression for the FFN layer is as follows, where x represents the hidden representation of a specific position in the sequence, W1 and W2 represent coefficient matrices, and b1 and b2 represent bias vectors: FFN(x, W1, W2, b1, b2) = max(0, xW1 + b1)W2 + b2.
[0112] The large language model of this disclosure replaces the ReLU nonlinearity with the SwiGLU nonlinear activation function to improve the performance of the FFN layer. SwiGLU is a combination of the Swish and Gate Liner Unit (GLU) activation functions. In SwiGLU, the Swish function is a self-gated activation function, which enables SwiGLU to capture the advantages of Swish and GLU while overcoming their shortcomings. The mathematical expression of the SwiGLU activation function is as follows:
[0113] Among them, Swish β (x) = xσ(βx), β is a specified constant, σ() is the Sigmoid function, Represents a point-wise matrix multiplication operation.
[0114] As shown in Figure 6, this paper uses the Low-Rank Adaptation of Quantized LLMs (QloRA) method to fine-tune large language models. This technique enables low-precision quantization of neural networks and high-fidelity 4-bit fine-tuning. QloRA includes two data types: a low-precision storage data type, typically 4 bits; and a computational data type, typically BFloat16. This dequantizes weight tensors to BFloat16, enabling matrix multiplication operations with 16-bit computational precision. The model itself is loaded using 4 bits, and during training, the values are dequantized to BFloat16 before training. The quantization technique uses 4-bit NormalFloat (NF4) quantization and double quantization. NF4 estimates the quantiles of the input tensor to ensure equal values are assigned to each quantization bin, while double quantization requantizes the quantization constant to reduce average memory usage. The QLoRA method allocates paged memory for the optimizer state, automatically unloading it to the CPU memory when the CPU memory is insufficient, and loading it back into the CPU memory when the optimizer update step requires it. The QLoRA method inserts adapters at all fully connected layers, increasing the training parameters to compensate for the performance loss caused by accuracy. In the adapter layer, the input is first projected down to a smaller dimension (Feedforward Down-project), passes through a layer of nonlinear activation function (Nonlinearity), and then projected up to the original dimension (Feedforward Up-project). In addition, there is a residual connection between the input and output of the entire adapter layer.
[0115] The large language model of the disclosed embodiment is a generative model. When making predictions, the generative model uses the information of K-1 words as prior information, without considering the information of the Kth and subsequent words, and directly predicts the Kth word. This is also a self-supervised learning method. The Kth token (token, usually used to represent the smallest unit in text or sequence data) is used as the K-1th target, and the value of the loss function is calculated, where:
[0116] Among them, loss represents the sum of the losses of m words.
[0117] The large language model in the disclosed embodiments needs to handle a variety of text tasks, including text classification, summary generation, plugin judgment, and similarity detection. Therefore, it is necessary to fine-tune the instructions to better adapt the large language model to these tasks. When performing different tasks, the system organizes the text data into JSON format and inputs different "instructions" to the large language model at different task stages.
[0118] (1) Text classification task:
[0119] For example, when performing a text classification task, the system inputs text data (text) and instructions (instruction) similar to the following into the large language model, and the large language model outputs (target) similar to the following:
[0120] {“text”: “Cinema promotion, I’ll pay for your movie, special prices, happy shopping, everyone can enjoy the discount.”,
[0121] "instruction": "Which of the following categories does this road sign belong to: movies, restaurants, housing prices, other",
[0122] "target": "movie"}
[0123] The number and types of categories included in the instruction are set by the user (i.e., the user's preset topics of interest). The instruction guides the large language model to classify the road sign within the set categories. If the road sign does not belong to the user's preset categories, it is directly classified as "other."
[0124] (2) Summary generation task:
[0125] For example, when performing a summary generation task, the system inputs text data (text) and instructions (instruction) similar to the following into the large language model, and the large language model outputs (target) similar to the following:
[0126] {"text":"(1),(2),(3)",
[0127] "instruction": "Based on the above content, summarize the natural and smooth judgment",
[0128] "target": "..."}
[0129] (1), (2), (3) represent three pieces of text data, ... represents a natural and fluent summary text. The large language model expands and organizes the "fragmented" input content into a natural and fluent text.
[0130] (3) Plug-in judgment task:
[0131] For example, when executing the plug-in judgment task, the system inputs text data (text) and instructions (instruction) similar to the following into the large language model, and the large language model outputs (target) similar to the following:
[0132] {"text":"What is the main story of Fast and Furious 10?",
[0133] "instruction": "Please determine whether you need to use external tools to answer the question."
[0134] "target": "This is a real-time question that requires a search engine tool. Search (Fast & Furious 10; Story)"}
[0135] The large language model determines whether external plug-ins are needed to expand the system's functions based on user needs. If so, it determines the required plug-in tools and transmission parameters.
[0136] (4) Similarity detection task:
[0137] For example, when performing a similarity detection task, the system inputs text data (text) and instructions (instruction) similar to the following into the large language model, and the large language model outputs (target) similar to the following:
[0138] {"text":"(1)(2)(3)......",
[0139] "instruction": "The three most similar fragments to the first text fragment are",
[0140] "target": "(2), (4), (5)"}
[0141] The text includes the question text and multiple (for example, 5) retrieved texts, and three text segments that are most similar to the original text are selected from the multiple retrieved texts.
[0142] In some exemplary embodiments, the method further comprises:
[0143] The user's questions and the answers generated by the large language model are screened and integrated to generate a document and sent to the user.
[0144] As shown in Figure 7, the in-vehicle intelligent interaction method provided by the embodiment of the present disclosure is based on multimodal (text, image, voice) input, helps users filter relevant text information through set topics of interest, and then integrates the interactive information along the way and outputs it to the user in the form of documents, images or voice.
[0145] The in-vehicle intelligent interaction method provided by the disclosed embodiments can integrate text information (including but not limited to road sign warnings, plaques, building signs, community names, etc.) and / or interactive information along the way and send it to the user's mobile phone in the form of an Excel, Word, or PDF document with one click. The fields included in the document can be specified by the user through voice or text input, or can be automatically specified by a large language model based on the user's voice interaction process (for example, determined by the user's high-frequency questions).
[0146] The following uses two actual application scenarios as examples to further explain the in-vehicle intelligent interaction method provided by the embodiments of the present disclosure.
[0147] Scenario 1: Information Collection Function
[0148] As shown in Figure 8, the user drives to a nearby residential area and wants to collect information about surrounding real estate projects. First, on the system interface shown in Figure 9, the user sets the topic of interest (i.e., the information to be collected along the way), such as housing prices. When arriving at a certain residential area, as shown in Figure 10A, the text information of the subject to be identified along the way is obtained through the on-board camera and OCR technology. For example, if the image taken along the way is shown in Figure 10B, the text information in the image is recognized as follows: {text1: "Zhongnan Garden, fixed price: 12800 yuan / M 2 , duplex, free a brand new Honda Accord", text2: "Champagne Plaza, sky villa, fixed price: 14800 yuan / M 2 , get a brand new BMW", text3: "Contact number: Mr. Liu 13802267", text4: "Ad space for rent, phone: 1371533609"}. The system inputs all identified text information and topics of interest to the user into the large language model, which then initiates text classification. The large language model classifies the text based on the user-set categories "housing prices, other", resulting in the ad classification result of {text1: "housing prices", text2: "housing prices", text3: "other", text4: "other"}. At this point, the system detects the housing price information and directly alerts the user through voice without requiring a wake-up operation. For example, the voice prompt can be as follows:
[0149] Bot (i.e., large language model): "Hello, we have detected housing price information for this community. Do you need detailed information?"
[0150] User: “Yes.”
[0151] As shown in Figure 11, after receiving the user's permission, the large language model can initiate the summary generation task. It generates text1 as "Zhongnan Garden duplex, fixed price 12,800 yuan / m², and even more surprise! Buy a house and receive a brand new Honda Accord!" and then inputs this text into speech for synthesized broadcasting.
[0152] The large language model can receive the user's speech and perform voice interaction with the user. For example, the voice interaction process can be as follows:
[0153] User: "Is this a school district housing complex?"
[0154] Bot: "This community is a school district housing, there is a xxx primary school and yyy middle school nearby."
[0155] User: "How does yyy Middle School rank in the city?"
[0156] bot: "This school's admission rate ranks fifth in the city."
[0157] User: “How is the teaching staff?”
[0158] bot: "Strong faculty, bilingual teaching, 32 national outstanding teachers"
[0159] User: “Sounds good.”
[0160] As shown in Figure 12, the large language model can organize the conversation into text and send it to the user's phone. Based on the positive and negative feedback from the user during the conversation, the positive feedback segments are filtered and integrated to form a document. If the user wants to view housing prices along the route, they can directly ask the in-car voice assistant. The assistant will organize the key information based on the user's needs and send it to the phone. The user can click to download it. For example, a user can ask, "Please give me an Excel file of the housing prices and amenities of the community I just saw." The large language model understands the user's request, generates an Excel file, and sends it to the user's phone, as shown in Figure 13.
[0161] Scenario 2: Memo function
[0162] Because the journey is long, if the user does not want to make a special trip for something they need but not in a rush, they can first record the item in a note. When they pass by a store that sells the item, a voice prompt will be sent to the user to buy the item.
[0163] For example, if a light bulb at home is broken, to avoid forgetting or letting other things clutter up your memory, you can tell the in-car voice assistant, "I need to buy a light bulb. I need the type, style, and size of the light bulb." The system will then label "light bulb" as a category, and this label will remain until the user removes the information from the memo, as shown in Figure 14.
[0164] As shown in Figure 15, some time later, when the user goes to a nearby park and passes by a lighting store, they may forget about buying the lights and miss the opportunity to buy them. When the system detects the name of the lighting store, it reminds the user: "There is a lighting store here that may have the lights you need." This reminds the user of the purchase.
[0165] The present invention aggregates multimodal information of voice, text, and images to solve multiple continuity tasks in vehicle-mounted scenarios. The number of video frames extracted is determined by vehicle speed. The dynamic adjustment process not only reduces the number of extracted frames, but also improves the clarity of video frames. A large language model is used to perform multiple text tasks in vehicle-mounted scenarios, including text classification, summary generation, plug-in judgment, similarity detection, etc. Internet plug-ins are used to expand the flexibility, richness, and consistency of question-and-answer scenarios, thereby improving the high availability of the vehicle-mounted assistant.
[0166] An embodiment of the present disclosure also provides an in-vehicle intelligent interaction device, comprising a memory; and a processor connected to the memory, wherein the memory is used to store instructions, and the processor is configured to execute the steps of the in-vehicle intelligent interaction method as described in any embodiment of the present disclosure based on the instructions stored in the memory.
[0167] As shown in Figure 16, in one example, the in-vehicle intelligent interaction device may include: a processor 1610, a memory 1620, a bus system 1630 and a transceiver 1640, wherein the processor 1610, the memory 1620 and the transceiver 1640 are connected through the bus system 1630, the memory 1620 is used to store instructions, and the processor 1610 is used to execute the instructions stored in the memory 1620 to control the transceiver 1640 to send and receive signals. Specifically, the transceiver 1640 can obtain an image captured by the image acquisition device under the control of the processor 1610, and the processor 1610 detects whether the captured image includes text content; when the captured image includes text content, extract multiple frames of images including text content; select one or more frames of images from the multiple frames as the subject to be identified; obtain the topic of interest to the user, and input the topic of interest to the user and the text content of the subject to be identified into the large language model, so that the large language model determines whether the text content of the subject to be identified belongs to the topic of interest to the user; when the text content of the subject to be identified belongs to the topic of interest to the user, output the text content of the subject to be identified to the user.
[0168] It should be understood that the processor 1610 may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0169] The memory 1620 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1610. A portion of the memory 1620 may also include a non-volatile random access memory. For example, the memory 1620 may also store information about the device type.
[0170] In addition to the data bus, the bus system 1630 may also include a power bus, a control bus, a status signal bus, etc. However, for the sake of clarity, various buses are labeled as the bus system 1630 in FIG.
[0171] During implementation, the processing performed by the processing device can be completed by the hardware integrated logic circuit in the processor 1610 or by instructions in the form of software. That is, the method steps of the embodiment of the present disclosure can be embodied as being executed by a hardware processor, or being executed by a combination of hardware and software modules in the processor. The software module can be located in a storage medium such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 1620, and the processor 1610 reads the information in the memory 1620 and completes the steps of the above method in combination with its hardware. To avoid repetition, it will not be described in detail here.
[0172] The present disclosure also provides an in-vehicle intelligent interaction system, including: an image acquisition device and an in-vehicle intelligent interaction device, wherein the in-vehicle intelligent interaction device is configured to:
[0173] Acquire an image captured by a vehicle-mounted image acquisition device, and detect whether the captured image includes text content;
[0174] When the captured image includes text content, extracting multiple frames of images including the text content;
[0175] Select one or more frames of images from the multiple frames of images as the subject to be identified;
[0176] Obtaining a topic of interest to the user, and inputting the topic of interest to the user and the text content of the subject to be identified into a large language model, so that the large language model determines whether the text content of the subject to be identified belongs to the topic of interest to the user;
[0177] When the text content of the subject to be identified belongs to a topic that the user is interested in, the text content of the subject to be identified is output to the user.
[0178] In some exemplary embodiments, the in-vehicle intelligent interactive device may include: a messaging system, which stores topics of interest to the user, and when the text content of the subject to be identified belongs to the topic of interest to the user, the messaging system pushes the text content of the subject to be identified to the user.
[0179] In the embodiment of the present disclosure, how the in-vehicle intelligent interaction device specifically performs in-vehicle intelligent interaction can refer to the embodiment of the in-vehicle intelligent interaction method described above, and will not be repeated here.
[0180] The present disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the in-vehicle intelligent interaction method described in any of the embodiments of the present disclosure. The method for driving in-vehicle intelligent interaction by executing executable instructions is substantially the same as the in-vehicle intelligent interaction method provided in the aforementioned embodiments of the present disclosure and is not further described here.
[0181] In some possible implementations, various aspects of the in-vehicle intelligent interaction method provided by the present disclosure may also be implemented in the form of a program product, which includes program code. When the program product is run on a computer device, the program code is used to enable the computer device to execute the steps of the in-vehicle intelligent interaction method according to various exemplary embodiments of the present disclosure described above in this specification. For example, the computer device may execute the in-vehicle intelligent interaction method recorded in the embodiments of the present disclosure.
[0182] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0183] It will be appreciated by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In hardware implementations, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As is well known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable, and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media generally embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0184] It should be noted that the above-described embodiments or implementations are merely illustrative and not restrictive. Therefore, the present disclosure is not limited to what is specifically shown and described herein. Various modifications, substitutions, or omissions may be made to the forms and details of the implementations without departing from the scope of the present disclosure.
Claims
1. An in-vehicle intelligent interaction method, including: Obtain the image collected by the in-vehicle image acquisition device, and detect whether the collected image includes text content; When the collected image includes text content, extract multiple frames of images including the text content; Select one or more frames of images from the multiple frames of images as the subject to be recognized; Obtain the topic that the user is interested in, and input the topic that the user is interested in and the text content of the subject to be recognized into the large language model, so that the large language model determines whether the text content of the subject to be recognized belongs to the topic that the user is interested in; When the text content of the subject to be recognized belongs to the topic that the user is interested in, output the text content of the subject to be recognized to the user.
2. The method according to claim 1, wherein, The extraction of multiple frames of images including text content includes: Determine that the time when it is first detected that the collected image includes text content is the entry time t0, and the time when it is first not detected that the collected image includes text content after t0 is the exit time t1; calculate the duration of the text content in the collected image as t = t1 - t0; Determine the number of frames of the extracted images and the frame extraction time according to the current vehicle speed and the duration of the text content in the collected image; Extract multiple frames of images according to the number of frames of the extracted images and the frame extraction time.
3. The method according to claim 2, wherein, The number of frames of the extracted image is v is the current vehicle speed, v max is the maximum vehicle speed, n 0 is the default number of video frames to be extracted.
4. The method according to claim 3, wherein, The time of the extracted I-th frame image is I is between 1 and n.
5. The method according to claim 1, wherein, The selection of one or more frames of images from the multiple frames of images as the subject to be recognized includes: Determine the resolution of the multiple frames of extracted images, and filter out the images with a resolution lower than the resolution threshold according to the preset resolution threshold; Identify the text in each frame of image, as well as the coordinates and background color corresponding to the text; Group the text according to the text and the corresponding background color; Perform similarity detection on the text in the same group, and delete the text with lower similarity and the images corresponding to the incomplete text in each group; Unify the text coordinates of the text in each group that is not deleted, and use the text in the largest text box in each group as the subject to be recognized.
6. The method according to claim 5, wherein, Perform similarity detection according to the following formula: The Levenshtein distance is the minimum number of edits between string S 1 and string S 2 . S 1 .length is the length of string S 1 , and S 2 .length is the length of string S 2 .
7. The method according to claim 1, wherein, The topic that the user is interested in is set by any one or more of the following methods: Pre-stored in the storage unit; Input by the user through voice and / or text; Obtained through big data analysis.
8. The method according to claim 1, wherein, The output of the text content of the subject to be recognized to the user includes at least one of the following: Input the text content of the subject to be recognized into the large language model, so that the large language model generates an abstract according to the text content of the subject to be recognized, and broadcasts the generated abstract to the user; Output an image containing the subject to be recognized and a label to the user, where the label is the category of the subject to be recognized; Organize the text content of the subject to be recognized into a document and send it to the user; Output a video containing the subject to be recognized to the user.
9. The method according to claim 1, wherein, The method further includes: Receive the user's question; Input the user's question into the large language model; The large language model determines whether to perform an online search. When an online search is required, the large language model retrieves multiple text fragments through an online search, determines the similarity between the multiple text fragments and the user's question, generates a summary of the content of the top N fragments with higher similarity, and outputs the generated summary to the user. When an online search is not required, the large language model answers based on its own knowledge.
10. The method according to claim 9, wherein, the method further includes: Screen and integrate the user's question and the answer generated by the large language model, and generate a document to be sent to the user.
11. The method according to claim 1, wherein, the large language model includes multiple decoder blocks, each decoder block includes an attention layer, a root mean square layer normalization, and a feed-forward neural network layer. The attention layer adds rotational position encoding to the query tensor and the key-value tensor. The feed-forward neural network layer includes two fully connected layers and a SwiGLU non-linear activation function. The large language model is fine-tuned through a low-rank adaptation method for quantized large models.
12. A vehicle-mounted intelligent interaction device, including a memory; and a processor connected to the memory. The memory is used to store instructions, and the processor is configured to execute the steps of the vehicle-mounted intelligent interaction method according to any one of claims 1 to 11 based on the instructions stored in the memory.
13. A vehicle-mounted intelligent interaction system, including a vehicle-mounted image acquisition device and the vehicle-mounted intelligent interaction device according to claim 12.
14. A computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the vehicle-mounted intelligent interaction method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Road sign recognition
CN108021862A
Street-along shop identification and recommendation system and method based on vehicle-mounted camera shooting
CN113298001A
Traffic text detection method and system under expert knowledge guidance mechanism
CN113505625A
Voice question answering method and device in driving scene and vehicle-mounted terminal
CN115312061A
Car window intelligent display method, system and device, electronic equipment and storage medium
CN116834663A
Cited By
Prediction interaction system and method for user behaviors in intelligent cabin
CN121278270A