Voice question and answer method and device in driving scene and vehicle-mounted terminal

By installing an environmental information collection component in the vehicle to acquire external environmental information and generate question-and-answer results, the problem of insufficient intelligence in question-and-answer systems in driving scenarios is solved, realizing intelligent question-and-answer and efficient interaction with the real-time environment during driving.

CN115312061BActive Publication Date: 2026-02-10GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210952625.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-09
Publication Date
2026-02-10
Estimated Expiration
2042-08-09

AI Technical Summary

Technical Problem

In existing driving scenarios, question-and-answer systems can only handle user questions about the status of intelligent vehicle devices or questions that can be answered directly by searching the Internet. The level of intelligence is low, and they cannot interact based on the current driving environment during driving.

Method used

By installing environmental information collection components in vehicles, external environmental information is acquired, and question-and-answer results are generated based on this information and voice question-and-answer commands. Data processing is performed using an in-vehicle terminal or server, and finally, interaction is achieved through voice broadcast.

Benefits of technology

It enables intelligent question-and-answer responses to the real-time external environment during driving, improving the success rate and intelligence of human-computer interaction. Users can ask and answer questions in real time based on the vehicle's external environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115312061B_ABST
    Figure CN115312061B_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a voice question and answer method and device in a driving scene and a vehicle-mounted terminal, and belongs to the technical field of human-computer interaction. The method comprises the following steps: in the case that a voice question and answer instruction is received, external environment information is acquired, the external environment information is collected by an environment information collection component during the driving of a carrier, and the external environment information is used to represent an external environment in which the carrier is located; based on the external environment information and the voice question and answer instruction, a question and answer result corresponding to the voice question and answer instruction is acquired; and voice broadcasting is performed based on the question and answer result. According to the scheme provided in the embodiment, a user can ask questions about the external environment of a cab, and the vehicle-mounted terminal can answer questions according to the environment, thereby improving the intelligent degree of a human-vehicle interaction question and answer system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of human-computer interaction technology, and in particular to a voice question-and-answer method, device and vehicle terminal for driving scenarios. Background Technology

[0002] With the rapid development of vehicle networking systems, human-vehicle interaction functions in driving scenarios are becoming increasingly popular. Among them, voice, as a convenient interaction method, has been widely used in human-vehicle interaction, and as a result, question-and-answer systems in driving scenarios are constantly being improved.

[0003] In related technologies, intelligent vehicle systems are equipped with voice interaction systems that can acquire user voice commands in driving scenarios to control devices or answer questions via voice. The voice question-and-answer function first converts speech into text using speech recognition technology, then finds the answer that matches the text through data querying, and finally returns the result to the user through screen display or voice broadcast to achieve the purpose of intelligent question-and-answer.

[0004] However, the question-and-answer system described above is only applicable to questions about the status of intelligent vehicle devices or questions that can be answered directly by searching the Internet. It can only achieve simple vehicle control and navigation functions, and its level of intelligence is relatively low. Summary of the Invention

[0005] This application provides a voice question-and-answer method, device, and vehicle-mounted terminal for driving scenarios. The technical solution is as follows:

[0006] On one hand, embodiments of this application provide a voice question-and-answer method in a driving scenario, the method comprising:

[0007] Upon receiving a voice question-and-answer command, external environment information is acquired. This external environment information is collected by an environmental information acquisition component during the vehicle's operation and is used to characterize the external environment in which the vehicle is located.

[0008] Based on the external environment information and the voice question-and-answer command, obtain the question-and-answer result corresponding to the voice question-and-answer command;

[0009] The audio is broadcast based on the question and answer results.

[0010] On the other hand, embodiments of this application provide a voice question-and-answer device for driving scenarios, the device comprising:

[0011] The information acquisition module is used to acquire external environment information upon receiving a voice question-and-answer command. The external environment information is acquired by the environmental information acquisition component during the vehicle's operation, and the external environment information is used to characterize the external environment in which the vehicle is located.

[0012] The result acquisition module is used to acquire the question and answer result corresponding to the voice question and answer command based on the external environment information and the voice question and answer command;

[0013] The voice broadcast module is used to broadcast voice information based on the question and answer results.

[0014] On the other hand, embodiments of this application provide a terminal, the terminal including a processor and a memory; the memory stores at least one instruction, the at least one instruction being executed by the processor to implement the voice question-and-answer method in a driving scenario as described above.

[0015] On the other hand, embodiments of this application provide a computer-readable storage medium storing at least one piece of program code, which is loaded and executed by a processor to implement the voice question-and-answer method in a driving scenario as described above.

[0016] On the other hand, embodiments of this application provide a computer program product including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the voice question-and-answer method in a driving scenario provided in various optional implementations of the above aspects.

[0017] In this embodiment of the application, upon receiving a voice question-and-answer command, the vehicle terminal can determine the question-and-answer result of the voice question-and-answer command based on external environment information that characterizes the external environment, and broadcast the question-and-answer result by voice, thereby realizing intelligent question-and-answer in response to the real-time external environment during driving, improving the success rate and intelligence level of human-computer interaction during driving. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of an implementation environment provided by an exemplary embodiment of this application;

[0019] Figure 2 This is a block diagram of the main components of a voice question-answering system provided in an exemplary embodiment of this application;

[0020] Figure 3 This is a flowchart of a voice question-and-answer method in a driving scenario provided by an exemplary embodiment of this application;

[0021] Figure 4 This is a flowchart of a voice question-and-answer method in a driving scenario provided by another exemplary embodiment of this application;

[0022] Figure 5This is a flowchart of a voice question-and-answer method in a driving scenario provided by another exemplary embodiment of this application;

[0023] Figure 6 This is a schematic diagram illustrating the process of acquiring external environment images provided in an exemplary embodiment of this application;

[0024] Figure 7 This is a schematic diagram illustrating the process of obtaining the question and answer text corresponding to a voice question and answer command provided in an exemplary embodiment of this application;

[0025] Figure 8 This is a flowchart of a voice question-and-answer method in a driving scenario provided in another exemplary embodiment of this application;

[0026] Figure 9 This is a schematic diagram illustrating the process of determining the first data collection period provided in an exemplary embodiment of this application;

[0027] Figure 10 This is a flowchart illustrating the process of obtaining the question-and-answer result corresponding to a voice question-and-answer command, provided in an exemplary embodiment of this application.

[0028] Figure 11 This is a schematic diagram of the question-answering analysis algorithm provided in the embodiments of this application;

[0029] Figure 12 This is a flowchart illustrating a method for extracting features from external environment information to obtain external environment features, provided by an exemplary embodiment of this application.

[0030] Figure 13 This is a schematic diagram illustrating the difference between the observer's perspective and the shooting perspective in an exemplary embodiment of this application;

[0031] Figure 14 This is a schematic diagram of a voice response application scenario provided by an exemplary embodiment of this application;

[0032] Figure 15 This is a schematic diagram of the state transition of an in-vehicle terminal according to an exemplary embodiment of this application;

[0033] Figure 16 This is a structural block diagram of a voice question-and-answer device in a driving scenario provided in an exemplary embodiment of this application;

[0034] Figure 17 This is a structural block diagram of an in-vehicle terminal provided in an exemplary embodiment of this application. Detailed Implementation

[0035] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0036] Figure 1 This is a schematic diagram of an implementation environment provided by an exemplary embodiment of this application. The implementation environment may include: a vehicle 110, an in-vehicle terminal 120, and a server 130.

[0037] The vehicle 110 may be a vehicle, a ship, an aircraft, etc. The following embodiments are all illustrated using a vehicle as an example, but this does not constitute a limitation.

[0038] The vehicle 110 is externally equipped with an environmental information acquisition component, which may include an image acquisition component 140 and an audio acquisition component 150. The image acquisition component is used to acquire images of the external environment, specifically scenes that can be seen by the human eye, such as buildings and vehicles in the external environment. The audio acquisition component is used to acquire audio of the external environment, specifically sounds that can be heard by the human ear, such as horns and birdsong in the external environment.

[0039] In this embodiment, the vehicle-mounted terminal 120 is disposed in the vehicle 110. The vehicle-mounted terminal 120 can be a vehicle-mounted system 121 or a mobile terminal 122 that establishes a communication connection with the vehicle-mounted system, such as an electronic device like a smartphone, laptop, or wearable device. Figure 1 The example described uses a smartphone as an example of a mobile terminal 122. The communication connection between the vehicle system 121 and the mobile terminal 122 can be established through wired or wireless means, such as Bluetooth connection, Universal Serial Bus (USB), Wireless Fidelity (WiFi) connection, or mobile data network connection, etc. This embodiment does not limit this.

[0040] The vehicle-mounted terminal 120 is used to process external environment information and user voice commands. Specifically, it extracts target external environment information from the external environment information, converts voice question-and-answer commands into corresponding voice question-and-answer text, and generates question-and-answer results based on the target external environment information and the voice question-and-answer text corresponding to the voice question-and-answer commands.

[0041] In this embodiment, the vehicle-mounted terminal 120 has the function of communicating with the server 130 via wireless communication, thereby establishing a connection and conducting data communication through the connection. This communication connection can be a Wi-Fi connection or a mobile data network connection, etc., and this embodiment does not limit it to any particular type.

[0042] In this embodiment of the application, when the vehicle terminal generates question-and-answer results based on voice question-and-answer commands and external environment information, it can be processed through the vehicle's infotainment system or mobile terminal, or it can generate question-and-answer results with the help of the server 130.

[0043] It should be noted that subsequent steps can only be executed after the voice question-and-answer program of the vehicle terminal is woken up. The wake-up command is preset. The steps in this application embodiment are executed after the voice question-and-answer program of the vehicle terminal is woken up. This application embodiment does not limit the method of waking up the voice question-and-answer program.

[0044] Indicative, such as Figure 1 As shown, the vehicle 110 is equipped with environmental information acquisition components in the front, rear, left, and right directions. Each environmental information acquisition component includes an image acquisition component 140 and an audio acquisition component 150. The image acquisition component 140 can be an external vehicle-mounted sensing camera, a dashcam, etc., and the audio acquisition component 150 can be a microphone, etc.

[0045] The image acquisition component 140 and the audio acquisition component 150 work together, and each of the image acquisition component 140 and the audio acquisition component 150 is connected to the vehicle terminal 120 to acquire external environmental information.

[0046] Optionally, an auxiliary imaging component can be set in the image acquisition component for auxiliary imaging. The auxiliary imaging component can be a millimeter radar or an infrared imaging instrument, etc.

[0047] Optionally, the image acquisition component and the audio acquisition component can establish a connection with the vehicle terminal 120 through methods such as Low-Voltage Differential Signaling (LVDS), Composite Video Broadcast Signal (CVBS), Controller Area Network (CAN), and Local Interconnect Network (LIN). The corresponding vehicle terminal 120 can read the external environment information acquired by the environmental acquisition component through this connection.

[0048] In addition, the vehicle terminal 120 is also equipped with a voice broadcast component and an image display component, which are used to broadcast or display the question and answer results corresponding to the voice question and answer commands.

[0049] Indicative, such as Figure 1 As shown, during driving, the image acquisition component 140 and the audio acquisition component 150 acquire external environmental information and cache the acquired information in the vehicle terminal's built-in memory. When a passenger in the vehicle sees scenery outside the window, they issue a voice question-and-answer command to a specific scene. After receiving the voice question-and-answer command, the vehicle terminal retrieves the cached external environmental information from its built-in memory, processes the data through the vehicle terminal 120 or the server 130, obtains the voice question-and-answer result, and then broadcasts the result.

[0050] In an illustrative example, the main components of a voice question-and-answer system are as follows: Figure 2 As shown, it mainly consists of an in-vehicle terminal 210, an environmental information acquisition component 220, a voice acquisition component 250, a voice broadcasting component 260, an image display component 270, a near-field analysis component 240, and a remote information analysis component 230. The remote information analysis component 230 is primarily a remote server that establishes a communication connection with the in-vehicle terminal, and the near-field information analysis component 240 is primarily a near-field mobile terminal that establishes a connection with the in-vehicle terminal.

[0051] Figure 2 The arrows in the diagram indicate the direction of information flow. The environmental information acquisition component 220 collects the vehicle's external environment in real time and sends the external environment information data to the vehicle terminal 210 for caching. After the voice acquisition component 250 collects the user's voice question-and-answer command, it sends the command to the vehicle terminal 210, which then extracts a certain duration of external environment information from the cached information. Subsequently, the vehicle terminal 210 performs calculations and analyses on the external environment information and the voice question-and-answer command. Alternatively, the vehicle terminal 210 sends the extracted external environment information and the command to the near-field information analysis component 240 or the far-field information analysis component 230 for calculations and analyses, and then returns the generated question-and-answer results to the vehicle terminal 210. The near-field information analysis component 240 can also perform distributed calculations in conjunction with the far-field information analysis component 230. After receiving the question and answer results, the vehicle terminal 210 processes the results and sends the voice question and answer results to the image display component 270 and the voice broadcast component 260 respectively. The image display component 270 and the voice broadcast component 260 then provide feedback on the question and answer results to the user.

[0052] Figure 3 This is a flowchart of a voice question-and-answer method in a driving scenario provided in an exemplary embodiment of this application. This embodiment uses this method for... Figure 1 The following explanation uses the vehicle-mounted terminal shown as an example. The method includes the following steps:

[0053] Step 301: Upon receiving a voice question-and-answer instruction, obtain external environment information.

[0054] External environmental information is collected by the environmental information acquisition component during vehicle operation, and the external environmental information is used to characterize the external environment in which the vehicle is located.

[0055] In one possible implementation, external environment information can characterize the external environment in which the vehicle is located from at least two dimensions, which may include an image dimension and an audio dimension.

[0056] In step 302, based on external environment information and voice question-and-answer instructions, the question-and-answer results corresponding to the voice question-and-answer instructions are obtained.

[0057] Based on external environment information and voice Q&A commands, the vehicle terminal selects the corresponding external environment information from the cached external environment information data according to the user's voice Q&A command, and then processes the external environment information and the voice Q&A command to obtain the Q&A result corresponding to the voice Q&A command.

[0058] Optionally, the question-and-answer results corresponding to the voice question-and-answer commands can be obtained through the local in-vehicle terminal, i.e., the vehicle's infotainment system or a mobile terminal, or the question-and-answer results can be obtained after data processing by the server.

[0059] In step 303, voice broadcast is performed based on the question-and-answer results.

[0060] After obtaining the question-and-answer results, the vehicle terminal will automatically fill the analysis results into the preset language broadcast template. The language broadcast template will contain polite language that is more in line with human interaction language, as well as some safe driving tips.

[0061] In some embodiments, while obtaining the question-and-answer results corresponding to the voice question-and-answer command, the in-vehicle terminal also extracts the current location information of the vehicle and queries the pre-built navigation map knowledge graph for any additional information that may exist at the current location, such as nearby landmarks, nearby restaurants, gas stations, etc. This additional information can be populated into the broadcast template along with the question-and-answer results for voice broadcast.

[0062] In summary, the voice question-and-answer method for driving scenarios provided in this application embodiment allows the in-vehicle terminal to obtain external environment information through received voice question-and-answer commands, generate corresponding question-and-answer results based on the obtained external environment information and the voice question-and-answer commands, and then broadcast them. This solves the problem that question-and-answer systems cannot interact based on the current driving environment during driving, achieving the effect of allowing users to ask and answer questions as soon as they see them in a driving scenario.

[0063] After the voice question-and-answer program of the vehicle terminal is activated, it needs to classify the voice commands issued by the user and implement different functions according to different user needs. When it is determined that the voice command issued by the user is a voice question-and-answer command, the vehicle terminal executes the subsequent steps of the embodiments of this application.

[0064] Figure 4 This is a flowchart illustrating a voice question-answering method in a driving scenario provided by another exemplary embodiment of this application. The method includes the following steps:

[0065] Step 401: Upon receiving a voice command, perform command type recognition on the voice command.

[0066] Voice commands are functionally divided into two categories: voice question-and-answer commands and non-voice question-and-answer commands. Non-voice question-and-answer commands include device control commands and navigation commands. Device control commands are used to adjust the working status of in-vehicle devices such as air conditioning, in-vehicle TV, and in-vehicle audio through voice interaction. Navigation commands are used to start the in-vehicle navigation system through human-computer interaction.

[0067] Voice question-and-answer commands are question-based commands, such as "Will it rain today?", "What's the temperature outside?", and "What kind of car is that white car?". The subsequent steps in this embodiment are all executed when the voice question-and-answer command relates to the vehicle's external environment. For conventional voice question-and-answer commands, the in-vehicle terminal can obtain the answer through methods such as searching a network database and then broadcast it via voice; this will not be elaborated upon here.

[0068] In one possible approach, voice command type recognition can be achieved through a command classification model. The command classification model is pre-trained to calculate the probability of the command type for each input voice command. The in-vehicle terminal inputs the voice text corresponding to the voice command into the command classification model, obtaining an output that represents the probability of the command type. Command types with probabilities higher than a threshold are identified as the final command types. Command types include device control commands, navigation commands, voice question-and-answer commands, etc. This command classification model can be trained based on a large number of sample commands and their corresponding command labels (indicating command types).

[0069] In another possible approach, the in-vehicle terminal can determine the type of voice command based on keyword recognition. For example, if a voice command contains keywords such as "air conditioning temperature" and "speaker volume," it is highly likely to be a device control command. If a voice command contains interrogative keywords such as "what is it," it is highly likely to be a question-and-answer command.

[0070] This application does not limit the specific way of classifying voice commands.

[0071] Step 402: If the voice command type is a question-and-answer command, determine that the voice question-and-answer command has been received and obtain external environment information.

[0072] After receiving a voice question-and-answer command, the vehicle terminal extracts external environment information and executes the subsequent steps of analyzing the voice question-and-answer command in this embodiment. If the voice command is not a voice question-and-answer command, the vehicle terminal does not execute the subsequent steps in this embodiment, but executes the program corresponding to the voice command. A non-voice question-and-answer command may be a parameter adjustment command or a navigation command for the vehicle's equipment; in this case, the vehicle terminal also executes the corresponding equipment adjustment program or navigation program. For example, if a user issues the command "Raise the air conditioning temperature," this is a non-question-and-answer command, so the subsequent steps are not executed, and the vehicle terminal controls the vehicle's air conditioning to raise the temperature.

[0073] Step 403: Based on external environment information and voice question-and-answer commands, obtain the question-and-answer results corresponding to the voice question-and-answer commands.

[0074] The implementation method of this step can refer to step 302 above, and will not be repeated here.

[0075] Step 404: Perform voice broadcast based on the question and answer results.

[0076] The implementation method of this step can refer to step 303 above, and will not be repeated here.

[0077] In summary, in real-world scenarios, the in-vehicle terminal determines the type of command issued by the user and executes subsequent steps only after confirming that the command is a voice question-and-answer command. This avoids the in-vehicle terminal processing non-voice question-and-answer commands based on external environmental information, thus avoiding the waste of processing resources.

[0078] In this embodiment, the vehicle terminal obtains the question-and-answer results corresponding to the voice question-and-answer command based on external environment information and voice commands. The external environment information includes all information in multiple dimensions or within a certain period of time, but the voice question-and-answer command may only be made for a certain dimension or a certain period of time. If the question-and-answer results are obtained based on all information content, it will not only cause the vehicle terminal to perform unnecessary data processing, but also affect the accuracy of the question-and-answer results. Therefore, it is necessary to filter the external environment information first.

[0079] In one possible implementation, the vehicle terminal extracts target external environment information from external environment information based on voice question-and-answer commands. The correlation between the target external environment information and the voice question-and-answer commands is higher than the correlation between other external environment information and the voice question-and-answer commands.

[0080] The correlation can include at least one of dimensional correlation or temporal correlation. Accordingly, the target external environment information can be information of a specific dimension or information collected over a specific period of time.

[0081] Therefore, target external environment information can be extracted from external environment information from two aspects: identifying the problem dimension and determining a specific time period. The following will describe these two methods of extracting target external environment information through two exemplary embodiments.

[0082] Figure 5 This is a flowchart illustrating a voice question-answering method in a driving scenario provided by another exemplary embodiment of this application. The method includes the following steps:

[0083] Step 501: Upon receiving a voice question-and-answer instruction, obtain external environment information.

[0084] External environment information contains information in multiple dimensions, including at least image and sound dimensions. The image dimension corresponds to the external environment images in the external environment information, and the sound dimension corresponds to the external environment audio in the external environment information.

[0085] The process of acquiring external environmental information is explained below. For example... Figure 6 As shown, in one possible implementation, firstly, the vehicle's external image acquisition component takes pictures of the external environment, and then the captured content is processed to obtain video images. The image acquisition component includes vehicle-mounted external sensing cameras or dashcams and other photography devices.

[0086] Optionally, auxiliary imaging devices, such as millimeter-wave radar or infrared imagers, can be used for auxiliary imaging processing when acquiring external environmental information. These devices can perform orthogonal processing on image frames based on non-visible light band images acquired by the auxiliary imaging devices, making the positional information in the image more accurate, and ultimately resulting in more accurate question-and-answer results. For example, when there are multiple vehicles ahead, and a user asks a question about a specific vehicle, it is difficult to accurately locate the target vehicle based solely on footage from a camera or dashcam. Adding auxiliary imaging devices allows for further confirmation of the target vehicle based on factors such as distance and orientation. This enables users to ask targeted questions, such as "What is the second car ahead?"

[0087] After the image acquisition component captures images of the external environment, the vehicle terminal caches them for later retrieval. Since image information occupies a large amount of storage space, the caching time cannot be too long. Furthermore, users will ask questions based on the vehicle's real-time environment while driving, so the image caching time should be set to within two minutes. Once the preset caching time is reached, the oldest image frame is deleted, and the latest image frame is written.

[0088] The vehicle-mounted terminal reads image frames from the camera at a fixed frame rate. After reading the image frames, it uses an image filtering algorithm to quickly eliminate noise in each frame. If an auxiliary imaging device is used, it will also perform orthogonal processing on the image frames of the non-visible light band images acquired by the auxiliary imaging device.

[0089] Optionally, the frame rate is generally set to 20fps, that is, 20 frames of images are read per second. It can also be adjusted according to different shooting devices and application scenarios. This application embodiment does not limit this.

[0090] The storage method for external environmental audio is similar to that for image caching. It also requires noise reduction processing of the acquired audio or processing by other algorithms before caching, which will not be elaborated here.

[0091] Step 502: Perform question dimension recognition on the voice question and answer text corresponding to the voice question and answer command to obtain the question dimension corresponding to the voice question and answer text. The question dimension includes at least one of image dimension and sound dimension.

[0092] The question-and-answer text corresponding to the voice question-and-answer command is obtained by the in-vehicle terminal through sequentially executing beamforming algorithm, front-end signal processing, and ASR (Automatic Speech Recognition) algorithm on the voice question-and-answer command, such as... Figure 7 As shown.

[0093] Optionally, the front-end signal processing uses the ANC (Active Noise Cancellation) algorithm to eliminate ambient noise; the AEC (Acoutic Echo Cancellation) algorithm to eliminate the voice echo broadcast by the vehicle terminal; and the AGC (Automatic Gain Control) algorithm to adjust the amplitude range of the voice signal so that the amplitude of the processed output signal is stable.

[0094] In some use cases, the ASR recognition result may be empty. In this case, the vehicle terminal will not perform subsequent steps, but will wait for a period of time and return to standby mode.

[0095] After the vehicle terminal receives the voice question and answer text corresponding to the voice question and answer command, it identifies the question dimension. After identifying the question dimension, it extracts the corresponding target external environment information from the external environment information based on the question dimension and the type of external environment information.

[0096] In one possible approach, a question classification model is pre-trained to calculate the probability representing a question dimension. The in-vehicle terminal inputs the voice-answered text corresponding to the voice command into the question classification model, obtaining an output that represents the probability of a question dimension. Dimensions with probabilities higher than a threshold are determined as the final question dimensions. This question classification model can be trained based on a large number of sample questions and their corresponding question labels (indicating question dimensions).

[0097] In another possible approach, the in-vehicle terminal can determine the question dimension (image-related keywords, sound-related keywords) corresponding to the voice Q&A command based on keyword matching. Words that characterize the appearance of objects, such as color and shape, can be used as image-related keywords, for example, red, green, spherical, largest, etc. Words that characterize sound, such as sound and onomatopoeia, can be used as sound-related keywords, for example, birdsong, beeping sounds, etc.

[0098] This application does not limit the specific problem dimension identification method.

[0099] Step 503: When the problem dimension is the image dimension, extract the external environment image from the external environment information as the target external environment information.

[0100] In this context, the image dimension refers to the user's question from an image perspective. Descriptions of shape, color, and size—information observable by the human eye—all fall under the image dimension. For example, "What is that H-shaped building?" Clearly, the user's voice command is a description of the target's shape; therefore, this question belongs to the image dimension. The in-vehicle terminal extracts external environmental images from the external environment information as the target's external environment information.

[0101] Step 504: When the problem dimension is the sound dimension, extract the external environment audio from the external environment information as the target external environment information.

[0102] The sound dimension refers to the user's question from the perspective of sound. Descriptions of sound volume, characteristics, and presence or absence that can be captured by the human ear all fall under the sound dimension. For example, "What kind of bird is calling?" is clearly a voice question-and-answer command targeting sounds in the external environment. Therefore, the in-vehicle terminal extracts audio from the external environment information as the target external environment information.

[0103] In one possible implementation, the question dimension includes both image and sound dimensions. In this case, both external environment images and external environment audio are extracted simultaneously as the target external environment information. For example, a voice question-and-answer command might be "Which car is honking its horn now?" Clearly, this command targets both sound and image elements within the external environment. Therefore, the in-vehicle terminal needs to extract both external environment images and external environment audio as the target external environment information.

[0104] Step 505: Based on the target's external environment information and the voice question-and-answer command, obtain the question-and-answer results corresponding to the voice question-and-answer command.

[0105] After extracting the target's external environment information, the vehicle-mounted terminal analyzes and processes the target's external environment information and voice question-and-answer commands to obtain the question-and-answer results corresponding to the voice question-and-answer commands.

[0106] When the target external environment information is an external environment image, the vehicle terminal processes the external environment image and voice question-and-answer commands to obtain the question-and-answer result; when the target external environment information is an external environment audio, the vehicle terminal processes the external environment audio and voice question-and-answer commands to obtain the question-and-answer result; when the target external environment information includes both external environment images and external environment audio, the vehicle terminal processes the external environment images, external environment audio, and voice question-and-answer commands to obtain the question-and-answer result corresponding to the voice question-and-answer command.

[0107] In step 504, the target external environment information has been extracted. In step 505, the range of external environment information that the vehicle terminal needs to analyze is reduced, thereby reducing the amount of computation when analyzing external environment information and voice commands.

[0108] Step 506: If the external environment information includes external environment images, determine the associated image frames corresponding to the question-and-answer results in the external environment images, and display the associated image frames.

[0109] In some application scenarios, simply broadcasting the voice answer to a user's voice question-and-answer command makes it difficult for the user to intuitively understand the answer. For example, if the user's voice question-and-answer command is "Where did the road sign just now pass?", simply broadcasting the answer to such a voice question-and-answer command is unlikely to provide the user with enough information. Therefore, using associated image frames to display the answer allows the user to more intuitively obtain the corresponding question-and-answer result and acquire more information.

[0110] Optionally, based on the question-and-answer results corresponding to the voice question-and-answer command, the image frame in which the target object indicated by the question-and-answer result is located in the image is determined, and then the frame with the best image quality is selected from several frames before and after the image frame as the associated image frame.

[0111] For example, if a user issues a voice command asking "Where does the road sign I just passed point to?", the vehicle terminal will execute the steps in this embodiment and obtain the answer "The road sign points to the department store, the food street, and Central Park". The image of the road sign captured by the image acquisition component will be displayed on the vehicle display screen, allowing the user to see the various locations and directions indicated by the road sign more intuitively, and obtaining more information than voice broadcast.

[0112] Step 507: Perform voice broadcast based on the question and answer results.

[0113] The implementation method of this step can refer to step 303 above, and will not be repeated here.

[0114] In summary, the question-answering method for driving scenarios provided in this embodiment, by identifying the question dimension and then extracting target external environment information based on the question dimension, enables the vehicle terminal to selectively extract some data from the external environment information as target external environment information, reducing the pressure on the vehicle terminal to process external environment information and improving the efficiency of the vehicle terminal in handling questions.

[0115] Furthermore, the method of displaying associated image frames provided in this embodiment allows users to not only learn the question-and-answer results by listening to voice broadcasts, but also by visualizing the results, further ensuring the reliability and accuracy of the question-and-answer results.

[0116] Figure 8 This is a flowchart of a voice question-answering method in a driving scenario provided by another exemplary embodiment of this application. The method includes the following steps:

[0117] Step 801: Upon receiving a voice question-and-answer instruction, perform time keyword recognition on the voice instruction.

[0118] Step 802: If the voice question and answer text corresponding to the voice question and answer command contains a time keyword, determine the first collection period based on the time keyword and the receiving time.

[0119] After receiving a voice question-and-answer command, the vehicle-mounted terminal needs to analyze and process the external environment information and the voice question-and-answer command. Since the data volume of images and audio is large, processing the external environment information within the entire preset buffer time is also very costly. Therefore, the vehicle-mounted terminal can first perform time keyword recognition on the voice question-and-answer text corresponding to the voice question-and-answer command, then determine a specific time period based on the recognized time keywords, and then perform data analysis and processing on the external environment information within that specific time period. This significantly reduces the computational load and lowers the overhead.

[0120] When the vehicle terminal receives a voice question-and-answer command, it performs time keyword recognition on the corresponding question-and-answer text. For example, "just now," "five seconds ago," and "within one minute" are all time keywords. After recognizing the time keywords contained in the voice question-and-answer command, the time obtained by subtracting the duration expressed by the time keyword from the time when the vehicle terminal receives the voice question-and-answer command is taken as the start time of the first collection period. The time from this start time to the time when the voice question-and-answer command is received is taken as the first collection period.

[0121] like Figure 9 As shown, assuming t2 is the moment the voice Q&A command is received, the time from t1 to t2 is the duration described by the time keyword. Therefore, the time from t1 to t2 is set as the first data collection period. For example, if the vehicle terminal receives the user's voice Q&A command "How many convenience stores did we pass in the past minute?" at 17:33, the vehicle terminal will use 17:32-17:33 as the first data collection period.

[0122] Step 803: The external environment information collected during the first collection period is determined as the target external environment information.

[0123] After determining the first collection period, the vehicle-mounted terminal extracts the external environment information within the first collection period from the external environment information and uses it as the target external environment information.

[0124] For example, after determining the first collection period as 17:32-17:33, the vehicle terminal extracts the data from the cached external information for the period of 17:32-17:33 as the target external information.

[0125] Step 804: If the voice question and answer text corresponding to the voice question and answer command does not contain time keywords, determine the second collection period based on the receiving time.

[0126] This application applies to a driving scenario, where voice question-and-answer commands are issued by the user based on the vehicle's environment and are typically for short-term content. Therefore, to reduce data processing pressure, the in-vehicle terminal can determine a relatively short time period as the second collection period based on the moment the voice question-and-answer command is received. The second collection period is a short time preceding the moment the voice question-and-answer command is received. For example, if the in-vehicle terminal receives the user's voice question-and-answer command "What is that blue building on the left?" at 17:50:30, and the vehicle is in motion, and the user's question is based on the vehicle's real-time environment, then the 10 seconds preceding the moment the in-vehicle terminal receives the voice question-and-answer command, i.e., 17:50:20-17:50:30, can be designated as the second collection period.

[0127] Step 805: The external environment information collected during the second collection period is determined as the target external environment information.

[0128] After determining the second collection period, the vehicle-mounted terminal extracts the external environment information within the second collection period from the external environment information and uses it as the target external environment information.

[0129] For example, after determining the second collection period as 17:50:20-17:50:30, the vehicle terminal extracts the data from the cached external information for the period of 17:50:20-17:50:30 as the target external information.

[0130] Step 806: Based on the target external environment information and the voice question-and-answer command, obtain the question-and-answer result corresponding to the voice question-and-answer command.

[0131] In one possible implementation scenario, if the vehicle terminal does not obtain a question-and-answer result corresponding to the voice question-and-answer command based on the target external environment information within the first or second collection period, it still needs to process all external environment information within the preset cache period to obtain the question-and-answer result.

[0132] Step 807: Perform voice broadcast based on the question and answer results.

[0133] The implementation method of this step can refer to step 303 above, and will not be repeated here.

[0134] In summary, in this embodiment, when the vehicle terminal recognizes that the voice question-and-answer command contains time keywords, it determines the first collection period based on the reception time of the voice question-and-answer command and the time keywords. If the command does not contain time keywords, it determines the second collection period. Furthermore, it extracts and analyzes the corresponding external environmental information from the external environment information, making the extraction of target environmental information by the vehicle terminal more targeted, reducing the data processing burden on the vehicle terminal, and improving the efficiency of obtaining question-and-answer results.

[0135] This application embodiment obtains the question-and-answer results corresponding to the voice question-and-answer commands based on external environment information and voice question-and-answer commands. Therefore, the vehicle terminal needs to analyze the target external environment information and the voice question-and-answer commands. Since a large amount of external environment information is collected during driving, although extracting the target external environment information based on the receiving time reduces data processing time to some extent, in most cases, processing image data solely through the local processor still presents a certain burden. Therefore, this application provides the following three methods, all of which can generate question-and-answer results, depending on the driving scenario and other factors.

[0136] 1. The vehicle terminal generates the question and answer results corresponding to the voice question and answer commands based on external environment information and voice question and answer commands.

[0137] 2. When network conditions meet transmission requirements, the vehicle-mounted terminal reports external environment information and voice Q&A commands to the server, so that the server can generate the corresponding Q&A results based on the external environment information and voice Q&A commands. The terminal then receives the Q&A results from the server.

[0138] Third, when network conditions do not meet transmission requirements, the vehicle-mounted terminal determines the target near-field device from among the near-field devices based on its computing power; it sends external environment information and voice question-and-answer commands to the target near-field device, so that the target near-field device can generate the corresponding question-and-answer results based on the external environment information and the voice question-and-answer commands; and it receives the question-and-answer results sent by the target near-field device. The near-field device can be a smartphone or tablet computer located inside the vehicle.

[0139] Near-field devices are mobile terminals that establish communication connections with in-vehicle terminals within the vehicle. These connections can be established via Bluetooth, WiFi, or other methods, and the in-vehicle terminal can identify near-field devices through Bluetooth scanning or other means.

[0140] After identifying the near-field device, the vehicle-mounted terminal determines the target near-field device from among the available near-field devices based on the device's computing power. Device computing power refers to the computing capability of a device to achieve specific output results by processing data. Computing power can be measured using objective data, and the computing power performance of different devices is pre-tested using dedicated testing programs.

[0141] Optionally, the computing power performance of different pre-tested devices can be sorted and assigned different priorities, with a highest priority of 1, representing the device with the strongest computing power. This prioritization is then stored in the vehicle terminal's built-in memory. When it's necessary to determine a target near-field device, the vehicle terminal selects the device with the highest relative computing power priority from among the near-field devices and identifies it as the target near-field device. For example, typically, a laptop's computing power is greater than a smartphone's, which is greater than a smartwatch's. The computing power priority of a laptop is set to 1, a smartphone's to 2, and a smartwatch's to 3. This priority is then stored in the vehicle terminal's built-in memory. When the vehicle terminal needs to determine a target near-field device, if a laptop is among the near-field devices, it is identified as the target near-field device; otherwise, the device with the highest priority among the remaining near-field devices is identified as the target near-field device.

[0142] Optionally, network latency and transmission speed thresholds can be set. If the current network latency is greater than the set threshold, or if the current network transmission speed is less than the set threshold, the vehicle terminal determines that the current network conditions do not meet the transmission requirements. If the current network latency is less than the set threshold and the current network transmission speed is greater than the set threshold, the vehicle terminal determines that the current network conditions meet the transmission requirements.

[0143] Of course, other parameters can also be used to determine whether the current network status meets the transmission conditions, but this application embodiment does not limit this.

[0144] In one possible implementation, when the vehicle-mounted terminal needs to process data via a near-field device or a remote server, it can directly transmit the received voice question-and-answer commands to the near-field device or the remote server, or it can convert the received voice question-and-answer commands into corresponding voice question-and-answer text before transmitting them to the near-field device or the remote server. This embodiment does not limit this approach.

[0145] Figure 10 This is a flowchart illustrating the process of obtaining the question-and-answer result corresponding to a voice question-and-answer command, provided in an exemplary embodiment of this application. Based on the aforementioned method one, this embodiment describes the method steps for obtaining the question-and-answer result corresponding to a voice question-and-answer command, as follows: Figure 10 As shown, the method includes the following steps:

[0146] Step 1001: Extract features from the external environment information to obtain external environment features.

[0147] The vehicle-mounted terminal extracts features from external environmental information. External image features include color, texture, and shape features, while external audio features include loudness, pitch, and timbre. When extracting external image features, the vehicle-mounted terminal can employ algorithms such as non-local or slow-fast models.

[0148] Step 1002: Perform feature concatenation on the external environment features and the text features of the voice question and answer text corresponding to the voice question and answer command to obtain fused features.

[0149] After the vehicle terminal extracts features from the external environment information, if the external environment information includes image information, it reduces the three-dimensional tensor representing the environmental image information to a one-dimensional vector. Then, it concatenates the one-dimensional vector representing the image information with the text vector corresponding to the voice question and answer text to obtain the fused features.

[0150] Step 1003: Input the fused features into the question answering model to obtain the question answering results output by the question answering model.

[0151] The input of the question-answering model is a fusion feature vector of external environment information and voice question-answering instructions, and the output is the question-answering result corresponding to the voice question-answering instructions. The question-answering model can adopt algorithms such as convolutional neural networks, recurrent neural networks or Transformer models.

[0152] Here, taking external environment information at the image dimension as an example, the above steps are explained. In this embodiment, the vehicle terminal uses a question-and-answer analysis algorithm to implement the above steps, such as... Figure 11 As shown. Figure 11 This is a schematic diagram of a question-answering analysis algorithm provided in an embodiment of this application.

[0153] The target external environment and the voice question-and-answer text corresponding to the voice question-and-answer command serve as the input to the question-and-answer analysis algorithm, and the question-and-answer results serve as the output of the question-and-answer analysis algorithm.

[0154] First, the Slow-fast model 1102 is used to extract feature information from the target's external environment information 1101, thus obtaining the external environment features. Among them, the fast branch network 11021 has low computational overhead and is used to analyze dynamic change information in the video sequence, while the slow branch network 11022 has high computational overhead and a slightly larger number of parameters and is used to analyze information such as color, texture, and lighting changes in the video sequence.

[0155] After the fast and slow branch networks extract feature information respectively, they are fused through the feature fusion network 11023 to obtain a three-dimensional tensor representing image information. Then, through the dimensionality reduction network 1104, a one-dimensional vector 1106 representing image information is generated.

[0156] Meanwhile, the voice question and answer text 1103 corresponding to the voice question and answer command is generated by the text vector generation 1105, then word segmentation is performed, and then the word vector of each word is queried. The word vectors of these words are weighted and averaged to obtain the one-dimensional text vector 1107 of the voice question and answer text corresponding to the voice question and answer command.

[0157] Finally, the one-dimensional vector representing the image information and the one-dimensional text vector corresponding to the voice question-and-answer command are concatenated to obtain the fused feature vector 1108. The fused feature vector is then input into the Transformer model 1109 to generate the question-and-answer result.

[0158] When using sound as external environment information to obtain question-and-answer results, it is also necessary to extract features from the external environment audio, concatenate the text features of the voice question-and-answer text corresponding to the voice question-and-answer command to obtain a fusion vector, and then input it into the question-and-answer system model. This embodiment will not be elaborated here.

[0159] In this embodiment, the vehicle terminal extracts features from external environmental information and fuses these features to obtain question-and-answer results. This operation matches the obtained question-and-answer results with the features of the user's voice question-and-answer commands, enabling the user to ask questions based on vehicle environmental perception in driving scenarios and obtain accurate question-and-answer results, thus achieving a higher level of intelligence.

[0160] Furthermore, in practical applications, the user who triggers the voice question-and-answer command is not limited to the driver, but may be a user sitting in another seat. In this case, the observer's perspective is different from the shooting perspective of the external environment image. Therefore, the embodiments of this application provide another way to obtain the external environment features.

[0161] Figure 12 This is a flowchart illustrating the process of extracting features from external environment information to obtain external environment features, provided by an exemplary embodiment of this application. The method may include the following steps:

[0162] Step 1201: Determine the observation perspective, which is the perspective of the observer who triggered the voice question-and-answer command.

[0163] The observer's perspective refers to the perspective of the observer who issues the voice question-and-answer command when observing the external environment.

[0164] Optionally, the viewing angle is related to the observer's height, age, and seating position. Therefore, when determining the observer's viewing angle, the vehicle terminal uses sound source localization technology to roughly judge the spatial position of the user who issued the voice question and answer command. The spatial position includes, but is not limited to, the seating position and the height of the voice, and then reasonably infers the observer's viewing angle.

[0165] In one possible implementation, the vehicle terminal is equipped with a sound source localization device. The vehicle terminal uses the sound source localization device to locate the position of the user who triggered the voice question-and-answer command inside the vehicle, and then processes the external environment image according to the observer's perspective to generate a more accurate question-and-answer result.

[0166] Step 1202: Based on the observation perspective and the shooting perspective of the external environment image, perform an image affine transformation on the external environment image to obtain the transformed external environment image.

[0167] The shooting perspective and the observation perspective cannot be consistent to a large extent. Figure 13 This is a diagram illustrating the difference between the observer's perspective and the shooting perspective at a certain moment in an application scenario.

[0168] exist Figure 13 In the diagram, 1301 represents a vehicle in motion, and 1302 represents a building the vehicle passes by. This building can be simultaneously captured by both the vehicle's external camera 1303 and the user 1304 riding in the vehicle. As can be seen from the diagram, at this moment, the shooting angle of the vehicle's external camera differs from the observer's perspective. For the same object, because the image observed by the observer differs from the image captured by the camera, the user's question about the building may not correspond to the voice command.

[0169] Affine transformation refers to the process of performing one linear transformation and one translation in one vector space to transform into another vector space. Affine transformations include scaling, translation, rotation, reflection, and shearing. A straight line in the original image remains a straight line after an affine transformation, and parallel lines in the original image remain parallel lines after an affine transformation. This is the nature of affine transformation.

[0170] The purpose of the vehicle-mounted terminal to perform affine transformation on the external environment image is to transform the image captured by the external shooting device into an image that is more in line with the observer's perspective through image affine transformation, so that the image features can correspond to the features described in the voice question and answer command, thereby obtaining a more accurate question and answer result.

[0171] Step 1203: Extract features from the transformed external environment image to obtain external environment features.

[0172] It should be noted that this embodiment provides a method for extracting features from external environment information to obtain external environment features. This method can also be used for... Figure 10 The illustrated embodiment serves as... Figure 10 Step 1001 in the illustrated embodiment has a better implementation effect when the observer's perspective and the shooting perspective of the external environment image are different.

[0173] In summary, in real-world driving environments, the scenery outside the car window may appear differently depending on the observer's perspective, causing discrepancies between the description in the user's question-and-answer command and the features of the image captured by the camera. In this embodiment, an affine transformation is performed on the external environment image based on the perspective of observers at different positions within the vehicle. This makes the transformed image features more closely match the features in the user's question-and-answer command, thus leading to a more accurate answer.

[0174] Figure 14 This is a schematic diagram of a voice question and answer application scenario provided by an exemplary embodiment of this application.

[0175] exist Figure 14 In this embodiment, both the in-vehicle terminal and the external image acquisition device are turned on. The user initiates a voice question-and-answer program, asking questions about the scenery in the current vehicle environment. The in-vehicle terminal analyzes the user's voice commands based on the perceived external environmental information, generates corresponding question-and-answer results, and then broadcasts the answers and displays images. This embodiment uses a scenery-related question-and-answer scenario, but this does not constitute a limitation on the embodiment.

[0176] In an illustrative example, during voice question-and-answer sessions in a vehicle driving scenario, the state transition process of the in-vehicle terminal is as follows: Figure 15 As shown in the diagram. The arrows indicate the direction of state transitions for the vehicle-mounted terminal.

[0177] Standby state 1501 refers to the state in which the vehicle terminal is in before the entire voice question-and-answer program begins to run. In standby state 1501, the external information acquisition component continues to operate, collecting real-time information about the vehicle's external environment, but the user does not issue any voice question-and-answer commands at this time.

[0178] When the vehicle terminal is in voice receiving state 1502, the user issues a voice command, the voice acquisition component starts working, and sends the acquired voice command to the vehicle terminal for voice command type determination.

[0179] Information extraction status 1503 refers to the status when the vehicle terminal extracts the target external environment information from the cached external environment information.

[0180] When the vehicle terminal is in the analysis and calculation state 1504, it selects the best computing device and issues the corresponding data calculation command. The corresponding device uses a question-and-answer analysis model to perform calculations and returns the question-and-answer results to the vehicle terminal after generating the question-and-answer results corresponding to the voice question-and-answer command.

[0181] The broadcast result status 1505 means that the vehicle terminal, after further processing based on the question and answer results, enables the image display component and voice broadcast component to output accordingly.

[0182] In standby mode 1501, if no voice question and answer command is received, the vehicle terminal remains in standby mode 1501. If a voice question and answer command is received, the vehicle terminal switches to voice receiving mode 1502.

[0183] In voice receiving state 1502, the vehicle terminal continues to maintain voice receiving state 1502 within the truncation waiting time when the user's voice input is interrupted; when the voice information ASR recognition is empty, or the voice command is determined to be a non-voice question and answer command, the vehicle terminal returns to standby state 1501; when the voice command is determined to be a voice question and answer command, the vehicle terminal transfers to information extraction state 1503.

[0184] In information extraction state 1503, when the intelligent vehicle system fails to extract target environment information from the external environment, it returns to standby state 1501; when the vehicle terminal completes the extraction of target external environment information, the vehicle terminal transfers to analysis and calculation state 1504.

[0185] In the analysis and calculation state 1504, when the text of the question and answer result corresponding to the generated question and answer result is empty, the vehicle terminal returns to the standby state 1501; when the generated voice question and answer result is not empty, the vehicle terminal switches to the broadcast result state 1505.

[0186] When the broadcast result is in state 1505, if the user issues a new voice command, the vehicle terminal will directly switch to voice reception state 1502 and continue to execute the next voice question and answer program.

[0187] In any of the above states, if the user manually interrupts the voice Q&A program, the vehicle terminal will directly return to standby state 1501.

[0188] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.

[0189] Please refer to Figure 16 This illustration shows a structural block diagram of a voice question-and-answer device in a driving scenario provided by an exemplary embodiment of this application. The device may include:

[0190] The information acquisition module 1601 is used to acquire external environment information upon receiving a voice question and answer command. The external environment information is acquired by the environment information acquisition component during the vehicle's operation and is used to characterize the external environment in which the vehicle is located.

[0191] The result acquisition module 1602 is used to acquire the question and answer result corresponding to the voice question and answer command based on the external environment information and the voice question and answer command;

[0192] The voice broadcast module 1603 is used to perform voice broadcast based on the question and answer results.

[0193] Optionally, the result acquisition module 1602 is used for:

[0194] Based on the voice question-and-answer command, target external environment information is extracted from the external environment information, wherein the correlation between the target external environment information and the voice question-and-answer command is higher than the correlation between other external environment information and the voice question-and-answer command; this is used to obtain the question-and-answer result corresponding to the voice question-and-answer command based on the target external environment information and the voice question-and-answer command.

[0195] Optionally, the result acquisition module 1602 is used for:

[0196] Based on the voice question-and-answer command, target external environment information is extracted from the external environment information. The correlation between the target external environment information and the voice question-and-answer command is higher than the correlation between other external environment information and the voice question-and-answer command. Based on the target external environment information and the voice question-and-answer command, the question-and-answer result corresponding to the voice question-and-answer command is obtained.

[0197] Optionally, the result acquisition module 1602 is used for:

[0198] When the problem dimension is the image dimension, external environment images are extracted from the external environment information as the target external environment information; when the problem dimension is the sound dimension, external environment audio is extracted from the external environment information as the target external environment information.

[0199] Optionally, the result acquisition module 1602 is used for:

[0200] Based on the time of receiving the voice question-and-answer command and the time of collecting the external environment information, the target external environment information is extracted from the external environment information.

[0201] Optionally, the result acquisition module 1602 is used for:

[0202] If the voice question-and-answer text corresponding to the voice question-and-answer instruction contains a time keyword, a first collection period is determined based on the time keyword and the receiving time; the external environment information whose collection time is located in the first collection period is determined as the target external environment information; if the voice question-and-answer text corresponding to the voice question-and-answer instruction does not contain a time keyword, a second collection period is determined based on the receiving time; the external environment information whose collection time is located in the second collection period is determined as the target external environment information.

[0203] Optionally, the result acquisition module 1602 is used for:

[0204] Based on the external environment information and the voice question-and-answer command, the question-and-answer result corresponding to the voice question-and-answer command is generated;

[0205] or,

[0206] When the network conditions meet the transmission requirements, the external environment information and the voice question-and-answer command are reported to the server, so that the server can generate the question-and-answer result corresponding to the voice question-and-answer command based on the external environment information and the voice question-and-answer command; and receive the question-and-answer result sent by the server;

[0207] or,

[0208] If the network conditions do not meet the transmission requirements, the target near-field device is determined from the near-field devices based on the device's computing power; the external environment information and the voice question-and-answer command are sent to the target near-field device so that the target near-field device can generate the question-and-answer result corresponding to the voice question-and-answer command based on the external environment information and the voice question-and-answer command; and the question-and-answer result sent by the target near-field device is received.

[0209] Optionally, the result acquisition module 1602 is used for:

[0210] The external environment information is subjected to feature extraction to obtain external environment features; the external environment features and the text features of the voice question and answer text corresponding to the voice question and answer command are concatenated to obtain fused features; the fused features are input into the question and answer model to obtain the question and answer result output by the question and answer model.

[0211] Optionally, the result acquisition module 1602 is used for:

[0212] A viewing angle is determined, which is the perspective of the observer who triggers the voice question-and-answer command; based on the viewing angle and the shooting angle of the external environment image, an image affine transformation is performed on the external environment image to obtain the transformed external environment image; features are extracted from the transformed external environment image to obtain the external environment features.

[0213] Optionally, the information acquisition module 1601 is used for:

[0214] Upon receiving a voice command, the command type is identified; if the command type of the voice command is a question-and-answer command, the voice question-and-answer command is confirmed to have been received, and the external environment information is obtained.

[0215] Optionally, the device further includes:

[0216] The image display module is used to identify the command type of a voice command when it is received; if the command type of the voice command is a question-and-answer command, it determines that the voice question-and-answer command has been received and obtains the external environment information.

[0217] In summary, the voice question-and-answer device for driving scenarios provided in this embodiment can be used to obtain external environment information by receiving voice commands, and to obtain the corresponding question-and-answer results based on the external environment information and the voice commands, and then broadcast them. This solves the problem that question-and-answer systems cannot interact based on the current driving environment during driving, achieving a user-visible, question-and-answer effect in driving scenarios, resulting in a higher level of intelligence.

[0218] Please refer to Figure 17 This diagram illustrates a structural block diagram of an in-vehicle terminal provided in an exemplary embodiment of this application. The terminal 1700 can be implemented as the in-vehicle terminal in the various embodiments described above. The terminal 1700 may include one or more components such as a processor 1710 and a memory 1720.

[0219] Processor 1710 may include one or more processing cores. Processor 1710 connects to various parts within terminal 1700 using various interfaces and lines, and performs various functions and processes data of terminal 1700 by running or executing instructions, programs, code sets, or instruction sets stored in memory 1720, and by calling data stored in memory 1720. Optionally, processor 1710 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). Processor 1710 may integrate one or more of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), Neural-network Processing Unit (NPU), and modem. Specifically, the CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required to be displayed on the touch screen; the NPU is used to implement Artificial Intelligence (AI) functions; and the modem is used to handle wireless communication. Understandably, the aforementioned modem may also be implemented as a separate chip rather than being integrated into the processor 1710.

[0220] The memory 1720 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 1720 may include a non-transitory computer-readable storage medium. The memory 1720 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 1720 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the various method embodiments described above, etc.; the data storage area may store data created based on the use of the terminal 1700, etc.

[0221] In addition, those skilled in the art will understand that the structure of the terminal 1700 shown in the above figures does not constitute a limitation on the terminal. The terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements. For example, the terminal 1700 also includes a display screen, camera assembly, microphone, speaker, radio frequency circuit, sensor, audio circuit, WiFi module, power supply, Bluetooth module, etc., which will not be described in detail here.

[0222] This application also provides a computer-readable storage medium storing at least one piece of program code, which is loaded and executed by a processor to implement the question-and-answer method in the driving scenario described in the above embodiments.

[0223] This application provides a computer program product including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the question-and-answer method in a driving scenario provided in various optional implementations of the above aspects.

[0224] It should be understood that "multiple" as used in this article refers to two or more.

[0225] Furthermore, the step numbers described herein are merely illustrative of one possible execution order between steps. In some other embodiments, the steps may not be executed in the order of their numbers, such as two steps with different numbers being executed simultaneously, or two steps with different numbers being executed in the reverse order of the illustration. This application does not limit this.

[0226] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A voice question-and-answer method for driving scenarios, characterized in that, The method includes: Upon receiving a voice question-and-answer command, external environment information is acquired. This external environment information is collected by the environmental information acquisition component during the vehicle's operation and cached in the vehicle terminal's built-in memory. The external environment information is used to characterize the external environment in which the vehicle is located. If the voice question and answer text corresponding to the voice question and answer command contains a time keyword, a first collection period is determined based on the time keyword and the receiving time; the cached external environment information is extracted from the built-in memory of the vehicle terminal, and the external environment information whose collection time is in the first collection period is determined as the target external environment information. The correlation between the target external environment information and the voice question and answer command is higher than the correlation between other external environment information and the voice question and answer command. If the voice question and answer text corresponding to the voice question and answer instruction does not contain the time keyword, a second collection period is determined based on the receiving time; the cached external environment information is extracted from the built-in memory of the vehicle terminal, and the external environment information whose collection time is in the second collection period is determined as the target external environment information; Based on the target external environment information and the voice question-and-answer command, the question-and-answer result corresponding to the voice question-and-answer command is obtained. Specifically, if the target external environment information is an external environment image, data processing is performed on the external environment image and the voice question-and-answer command to obtain the question-and-answer result; if the target external environment information is external environment audio, data processing is performed on the external environment audio and the voice question-and-answer command to obtain the question-and-answer result; if the target external environment information includes both the external environment image and the external environment audio, data processing is performed on the external environment image, the external environment audio, and the voice question-and-answer command to obtain the question-and-answer result; or... If the network conditions meet the transmission requirements, the external environment information and the voice Q&A command are reported to the server, so that the server can generate the Q&A result corresponding to the voice Q&A command based on the external environment information and the voice Q&A command; or, the server sends the Q&A result; When the network conditions do not meet the transmission requirements, the system determines the target near-field device from among the near-field devices based on the device's computing power. The near-field device is a mobile terminal that establishes a communication connection with the vehicle-mounted terminal inside the vehicle. The system sends the external environment information and the voice question-and-answer command to the target near-field device, so that the target near-field device generates the question-and-answer result corresponding to the voice question-and-answer command based on the external environment information and the voice question-and-answer command. The system then receives the question-and-answer result sent by the target near-field device. Voice broadcast is performed based on the question and answer results; If the external environment information includes the external environment image, determine the associated image frame corresponding to the question-and-answer result in the external environment image; The associated image frames are then displayed.

2. The method according to claim 1, characterized in that, The method further includes: The question dimension is identified by performing question dimension recognition on the voice question-and-answer text corresponding to the voice question-and-answer command, and the question dimension includes at least one of image dimension and sound dimension. Based on the problem dimension and the type of external environment information, the target external environment information is extracted from the external environment information.

3. The method according to claim 2, characterized in that, The step of extracting the target external environment information from the external environment information based on the problem dimension and the type corresponding to the external environment information includes: When the problem dimension is the image dimension, external environment images are extracted from the external environment information as the target external environment information; When the problem dimension is the sound dimension, external environment audio is extracted from the external environment information as the target external environment information.

4. The method according to claim 1, characterized in that, The method further includes: Based on the external environment information and the voice question-and-answer command, the question-and-answer result corresponding to the voice question-and-answer command is generated.

5. The method according to claim 4, characterized in that, The step of generating the question-and-answer result corresponding to the voice question-and-answer command based on the external environment information and the voice question-and-answer command includes: Feature extraction is performed on the external environment information to obtain external environment features; The external environment features and the text features of the voice question and answer text corresponding to the voice question and answer command are concatenated to obtain fused features; The fused features are input into the question-answering model to obtain the question-answering result output by the question-answering model.

6. The method according to claim 5, characterized in that, The external environment information includes external environment images; Before extracting features from the external environment information to obtain the external environment features, the method further includes: Determine the observation perspective, which is the perspective of the observer who triggered the voice question-and-answer command; Based on the observation perspective and the shooting perspective of the external environment image, an image affine transformation is performed on the external environment image to obtain the transformed external environment image. The step of extracting features from the external environment information to obtain external environment features includes: Feature extraction is performed on the transformed external environment image to obtain the external environment features.

7. The method according to claim 1, characterized in that, Upon receiving a voice question-and-answer command, the process of obtaining external environment information includes: Upon receiving a voice command, the command type is identified. If the voice command is a question-and-answer command, it is determined that the voice question-and-answer command has been received, and the external environment information is obtained.

8. A voice question-and-answer device for driving scenarios, characterized in that, The device includes: The information acquisition module is used to acquire external environment information upon receiving a voice question and answer command. The external environment information is acquired by the environmental information acquisition component during the vehicle's operation and cached in the vehicle terminal's built-in memory. The external environment information is used to characterize the external environment in which the vehicle is located. The result acquisition module is used to determine a first collection period based on the time keyword and the receiving time when it is found that the voice question and answer text corresponding to the voice question and answer command contains a time keyword; extract the cached external environment information from the built-in memory of the vehicle terminal, and determine the external environment information whose collection time is in the first collection period as the target external environment information, wherein the correlation between the target external environment information and the voice question and answer command is higher than the correlation between other external environment information and the voice question and answer command; If the voice question and answer text corresponding to the voice question and answer instruction does not contain the time keyword, a second collection period is determined based on the receiving time; the cached external environment information is extracted from the built-in memory of the vehicle terminal, and the external environment information whose collection time is in the second collection period is determined as the target external environment information; Based on the target external environment information and the voice question-and-answer command, the question-and-answer result corresponding to the voice question-and-answer command is obtained. Specifically, if the target external environment information is an external environment image, data processing is performed on the external environment image and the voice question-and-answer command to obtain the question-and-answer result; if the target external environment information is external environment audio, data processing is performed on the external environment audio and the voice question-and-answer command to obtain the question-and-answer result; if the target external environment information includes both the external environment image and the external environment audio, data processing is performed on the external environment image, the external environment audio, and the voice question-and-answer command to obtain the question-and-answer result; or... If the network conditions meet the transmission requirements, the external environment information and the voice Q&A command are reported to the server, so that the server can generate the Q&A result corresponding to the voice Q&A command based on the external environment information and the voice Q&A command; or, the server sends the Q&A result; When the network conditions do not meet the transmission requirements, the system determines the target near-field device from among the near-field devices based on the device's computing power. The near-field device is a mobile terminal that establishes a communication connection with the vehicle-mounted terminal inside the vehicle. The system sends the external environment information and the voice question-and-answer command to the target near-field device, so that the target near-field device generates the question-and-answer result corresponding to the voice question-and-answer command based on the external environment information and the voice question-and-answer command. The system then receives the question-and-answer result sent by the target near-field device. The voice broadcast module is used to broadcast voice information based on the question and answer results; The image display module is used to determine the associated image frame corresponding to the question-and-answer result in the external environment image when the external environment information includes the external environment image; The associated image frames are then displayed.

9. The apparatus according to claim 8, characterized in that, The result acquisition module includes an information extraction unit, which is used for: The question dimension is identified by performing question dimension recognition on the voice question and answer text corresponding to the voice question and answer command, and the question dimension corresponding to the voice question and answer text includes at least one of image dimension and sound dimension; Based on the problem dimension and the type of external environment information, the target external environment information is extracted from the external environment information.

10. The apparatus according to claim 9, characterized in that, The information extraction unit is used for: When the problem dimension is the image dimension, external environment images are extracted from the external environment information as the target external environment information; When the problem dimension is the sound dimension, external environment audio is extracted from the external environment information as the target external environment information.

11. The apparatus according to claim 8, characterized in that, The result acquisition module includes: The first processing unit is configured to generate the question-and-answer result corresponding to the voice question-and-answer instruction based on the external environment information and the voice question-and-answer instruction.

12. The apparatus according to claim 11, characterized in that, The first processing unit is configured to: Feature extraction is performed on the external environment information to obtain external environment features; The external environment features and the text features of the voice question and answer text corresponding to the voice question and answer command are concatenated to obtain fused features; The fused features are input into the question-answering model to obtain the question-answering result output by the question-answering model.

13. The apparatus according to claim 12, characterized in that, The external environment information includes external environment images; The device further includes: A perspective determination module is used to determine the observation perspective, which is the perspective of the observer who triggered the voice question-and-answer command; The transformation module is used to perform an image affine transformation on the external environment image based on the observation viewpoint and the shooting viewpoint of the external environment image to obtain the transformed external environment image. The first processing unit is used to extract features from the transformed external environment image to obtain the external environment features.

14. The apparatus according to claim 8, characterized in that, The information acquisition module is used for: Upon receiving a voice command, the command type is identified. If the voice command is a question-and-answer command, it is determined that the voice question-and-answer command has been received, and the external environment information is obtained.

15. A vehicle-mounted terminal, characterized in that, The vehicle terminal includes a processor and a memory; the memory stores at least one instruction, which is executed by the processor to implement the voice question-and-answer method in the driving scenario as described in any one of claims 1 to 7.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one piece of program code, which is loaded and executed by a processor to implement the voice question-and-answer method in a driving scenario as described in any one of claims 1 to 7.

17. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the voice question-and-answer method in a driving scenario as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Picture searching method and device based on voice, storage medium and terminal

    CN109710796A

  • Voice application realization method, device and apparatus and computer readable storage medium

    CN109992248A

  • Vehicle-mounted questioning and answering method, vehicle-mounted questioning and answering system, vehicle and storage medium

    CN110459217A