Vehicle-mounted intelligent interaction method and device, computer readable storage medium and computer program product

By collecting occupants' visual and voice information in real time and utilizing a cross-modal attention matching fusion model, the problem of insufficient resource utilization in existing in-vehicle intelligent interaction assistants is solved, achieving more intelligent and accurate occupant interaction control and improving the intelligence and safety of the interaction.

CN121734432APending Publication Date: 2026-03-27SHENZHEN LONGHORN AUTOMOTIVE ELECTRONICS EQUIPCO
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing in-vehicle intelligent interactive assistants fail to fully utilize the visual and voice information resources of passengers, have limited level of interactive intelligence, and cannot effectively meet the real needs of passengers.

Method used

By collecting occupants' visual and voice information in real time, and using a cross-modal attention matching fusion model to perform feature fusion, ambiguity is eliminated, the occupants' precise interaction intentions and potential interaction needs are obtained, and function control commands are generated. Combined with confidence judgment and driving environment, safety thresholds are determined to achieve intelligent interaction.

Benefits of technology

It improves the intelligence level of in-vehicle intelligent interaction, generates function control commands that better meet the actual needs of passengers, enhances the accuracy and safety of interaction, and reduces the risk of misidentification and misoperation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121734432A_ABST
    Figure CN121734432A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a vehicle-mounted intelligent interaction method and device, a computer readable storage medium and a computer program product. The method comprises the following steps that predetermined visual information and voice information of passengers are correspondingly collected in real time; respectively analyzing the visual information and the voice information to correspondingly obtain and store visual attention area features and preliminary interaction intentions of the passengers; fusing the real-time visual attention area features and the preliminary interaction intention by adopting a cross-modal attention matching fusion model, and eliminating ambiguity to obtain the current accurate interaction intention of the passenger; fusing the historical and current visual attention area features of the passenger and the preliminary interaction intention to obtain the current potential interaction demand of the passenger; and generating a function control instruction based on the current accurate interaction intention and the potential interaction demand of the passenger, and sending the function control instruction to the vehicle equipment through the vehicle communication network. According to the embodiment, the potential demand of the passenger can be actively predicted.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the field of vehicle-mounted intelligent interaction, and in particular to a vehicle-mounted intelligent interaction method and device, a computer readable storage medium, and a computer program product. BACKGROUND

[0002] With the rapid development of artificial intelligence, autonomous driving, and information fusion technology, motor vehicles begin to be equipped with vehicle-mounted intelligent interaction assistants to facilitate the occupants (drivers and passengers) in the motor vehicle cabin to control the vehicle-mounted devices on the motor vehicle through voice or vision.

[0003] An existing vehicle-mounted intelligent interaction assistant is based on visual, voice, and other multi-modal information fusion technology, including feature layer fusion (such as visual and voice features being jointly encoded and then sent to a decision model), decision layer fusion (visual and voice information being modeled and then the decision results being fused), and mixed fusion (both feature and decision layer fusion being used) multi-strategy, and the specific technical implementation mainly includes: multi-modal Transformer model, multi-task joint training method, cross-modal attention mechanism, etc.

[0004] However, the inventors have found in specific implementation that the above-mentioned vehicle-mounted intelligent interaction assistant based on visual, voice, and other multi-modal information fusion technology is only used to passively respond to the explicit instructions issued by the occupants to control the vehicle-mounted devices, and can only simply and directly respond to the instructions issued by the occupants, and fails to fully and effectively utilize the performance of the related hardware and the collected resources such as predetermined visual information and voice information of the occupants to better analyze and meet the real needs of the occupants, and the intelligent degree of interaction with the occupants is relatively limited. SUMMARY

[0005] The technical problem to be solved by the embodiments of the present application is to provide a vehicle-mounted intelligent interaction method that can fully and reasonably utilize resources and more intelligently implement interaction with occupants.

[0006] The technical problem to be further solved by the embodiments of the present application is to provide a vehicle-mounted intelligent interaction device that can fully and reasonably utilize resources and more intelligently implement interaction with occupants.

[0007] The technical problem to be further solved by the embodiments of the present application is to provide a computer readable storage medium to store a computer program that can fully and reasonably utilize resources and more intelligently implement interaction with occupants.

[0008] The technical problem to be further solved by the embodiments of the present application is to provide a computer program product that can fully and reasonably utilize resources and more intelligently implement interaction with occupants.

[0009] To solve the above technical problems, the embodiment of the present application first provides the following technical solution: a vehicle-mounted intelligent interaction method, comprising the following steps: controlling a feature acquisition device installed in the cockpit to collect predetermined visual information and voice information of the occupant in real time, the predetermined visual information at least including facial expression information, gesture information, posture information and eye movement information; analyzing the predetermined visual information to correspondingly obtain and save the visual attention area features of the occupant, and also analyzing the voice information to correspondingly obtain and save the preliminary interaction intention of the occupant; using a cross-modal attention matching fusion model to fuse and eliminate ambiguity of the visual attention area features and the preliminary interaction intention of the occupant at present to obtain the accurate interaction intention of the occupant at present; using the cross-modal attention matching fusion model to fuse the visual attention area features and the preliminary interaction intention of the occupant in the past and at present to obtain the potential interaction demand of the occupant at present; and generating a function control instruction for controlling a corresponding vehicle-mounted device based on the accurate interaction intention and the potential interaction demand of the occupant at present, and sending the function control instruction to the vehicle-mounted device through a vehicle communication network.

[0010] Further, when the cross-modal attention matching fusion model is used to obtain the potential interaction demand of the occupant at present, the confidence of the potential interaction demand is also output; the generation of the function control instruction for controlling the corresponding vehicle-mounted device based on the accurate interaction intention and the potential interaction demand of the occupant at present comprises: judging whether the confidence of the potential interaction demand exceeds a preset threshold; if the confidence of the potential interaction demand exceeds the preset threshold, requesting the occupant to confirm the potential interaction demand, and based on the accurate interaction intention and the potential interaction demand of the occupant at present, respectively generating a corresponding function control instruction after receiving the demand confirmation instruction of the occupant, otherwise, based on the accurate interaction intention of the occupant at present, generating a corresponding function control instruction; and receiving the execution result feedback by the vehicle-mounted device and feeding back the execution result to the occupant.

[0011] Further, the method further comprises: an uncertainty index of the fusion result including the accurate interaction intention and the potential interaction demand of the occupant at present is calculated based on the consistency and confidence interval quantization method between the visual attention area features and the preliminary interaction intention fused by the cross-modal attention matching fusion model; a safety threshold of the uncertainty of the fusion result is determined based on the current driving scene and driving environment of the vehicle. determine a safety threshold of uncertainty of the fusion result based on a current driving scene and driving environment of the vehicle; determine whether the uncertainty index of the fusion result exceeds the safety threshold, if yes, request the passenger to confirm the fusion result, otherwise generate corresponding functional control instructions based on the current accurate interactive intention and the potential interactive demand of the passenger; and update parameters of the cross-modal attention matching fusion model in real time based on a confirmation result of the fusion result by the passenger.

[0012] Further, the visual attention area features and the preliminary interactive intention of the passenger in the past and at present are stored in the form of a cross-modal knowledge graph.

[0013] Further, the fusion of the visual attention area features and the preliminary interactive intention of the passenger at present by the cross-modal attention matching fusion model and the elimination of ambiguity to obtain the accurate interactive intention of the passenger at present specifically include: identify and analyze whether there is ambiguous referential information in the preliminary interactive intention; if there is the ambiguous referential information in the preliminary interactive intention, determine a referential target corresponding to the ambiguous referential information based on the visual attention area features of the passenger at present; and fuse the referential target in the visual attention area features of the passenger at present and the ambiguous referential information by the cross-modal attention matching fusion model to obtain the accurate interactive intention.

[0014] Further, the analysis of the predetermined visual information to correspondingly obtain and save the visual attention area features of the passenger specifically include: perform feature extraction on the predetermined visual information to determine a line-of-sight attention area, gesture action direction and facial posture of the passenger; and obtain and save the real-time visual attention area features based on the line-of-sight attention area, gesture action direction and facial posture of the passenger.

[0015] Further, the analysis of the speech information to correspondingly obtain and save the preliminary interactive intention of the passenger specifically include: perform data preprocessing on the speech information to obtain high-quality audio signals, the data preprocessing at least including noise reduction, echo cancellation and speech enhancement; convert the audio signals into text information by using real-time speech recognition technology; and identify the text information by using a natural language understanding algorithm model to obtain and save the preliminary interactive intention.

[0016] In another aspect, to solve the further technical problem above, the embodiment of the present application provides the following technical solution: a vehicle-mounted intelligent interaction device connected with a feature collection device installed in a cabin, the device comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, the processor implementing the vehicle-mounted intelligent interaction method according to any one of the above when executing the computer program.

[0017] In another aspect, to solve the further technical problem above, the embodiment of the present application provides the following technical solution: a computer-readable storage medium comprising a stored computer program, wherein the computer-readable storage medium controls a device where the computer-readable storage medium is located to execute the vehicle-mounted intelligent interaction method according to any one of the above when the computer program runs.

[0018] In another aspect, to solve the further technical problem above, the embodiment of the present application provides the following technical solution: a computer program product comprising a computer program, the computer program implementing the steps of the vehicle-mounted intelligent interaction method according to any one of the above when executed by a processor.

[0019] After the above technical solution, the embodiment of the present application has at least the following beneficial effects: the embodiment of the present application first pre-processes the predetermined visual information and the voice information of the occupant collected in real time to obtain the visual attention area feature and the preliminary interaction intention of the occupant at present, then obtains the accurate interaction intention of the occupant at present through fusion processing, and further fuses the visual attention area feature and the preliminary interaction intention of the occupant at present and in the past to predict the potential interaction demand of the occupant at present, finally, generates the function control instruction that can better meet the actual demand of the occupant based on the accurate interaction intention and the potential interaction demand, and sends the function control instruction to the vehicle equipment through the vehicle communication network to realize interaction control, the whole interaction process has higher intelligent degree, and the interaction result can better meet the actual demand of the occupant. BRIEF DESCRIPTION OF DRAWINGS

[0020] Figure 1 Flow chart of steps of an optional embodiment of the vehicle-mounted intelligent interaction method of the present application.

[0021] Figure 2 Flow chart of sub-steps of step S5 of an optional embodiment of the vehicle-mounted intelligent interaction method of the present application.

[0022] Figure 3 Flow chart of sub-steps of step S3 of an optional embodiment of the vehicle-mounted intelligent interaction method of the present application.

[0023] Figure 4 Principle block diagram of an optional embodiment of the vehicle-mounted intelligent interaction device of the present application.

[0024] Figure 5 Figure 1 is a functional module diagram of an optional embodiment of the vehicle-mounted intelligent interaction device of the present application. DETAILED DESCRIPTION

[0025] The present application will be further described below in conjunction with the drawings and specific embodiments. It should be understood that the following illustrative embodiments and descriptions are only intended to explain the present application and are not intended to limit the present application, and the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.

[0026] As shown in Figure 1 An optional embodiment of the present application provides a vehicle-mounted intelligent interaction method, comprising the following steps: S1: controlling a feature acquisition device 1 installed in a cockpit to collect predetermined visual information and voice information of an occupant in real time, wherein the predetermined visual information at least includes facial expression information, gesture information, posture information and eye movement information; S2: analyzing the predetermined visual information to correspondingly obtain and save the visual attention area features of the occupant, and also analyzing the voice information to correspondingly obtain and save the preliminary interaction intention of the occupant; S3: using a cross-modal attention matching fusion model to fuse and eliminate ambiguity of the visual attention area features and the preliminary interaction intention of the occupant at present to obtain the accurate interaction intention of the occupant at present; S4: using the cross-modal attention matching fusion model to fuse the visual attention area features and the preliminary interaction intention of the occupant at present and in the past to obtain the potential interaction demand of the occupant at present; and S5: generating a function control instruction for controlling a corresponding vehicle-mounted device based on the accurate interaction intention and the potential interaction demand of the occupant at present, and sending the function control instruction to the vehicle-mounted device through a vehicle communication network.

[0027] The embodiment of the present application first performs preprocessing after collecting the predetermined visual information and voice information of the occupant in real time to obtain the visual attention area features and the preliminary interaction intention of the occupant at present, and then obtains the accurate interaction intention of the occupant at present through fusion processing, and also fuses the visual attention area features and the preliminary interaction intention of the occupant at present and in the past to predict the potential interaction demand of the occupant at present, finally, generates a function control instruction that can better meet the actual demand of the occupant based on the accurate interaction intention and the potential interaction demand, and sends the function control instruction to the vehicle device through a vehicle communication network to realize interaction control, the whole interaction process has a higher degree of intelligence, and the interaction result can better meet the actual demand of the occupant.

[0028] In practice, the cross-modal attention matching fusion model can be a vision-language model (VLM) or a multi-modal Transformer model.

[0029] In an optional embodiment of the present application, the cross-modal attention matching fusion model is used to obtain the current potential interaction demand of the occupant, and the confidence of the potential interaction demand is also output. Figure 2 As shown, step S5 includes: S51: determining whether the confidence of the potential interaction demand exceeds a preset threshold; S52: if the confidence of the potential interaction demand exceeds the preset threshold, requesting the occupant to confirm the potential interaction demand, and generating corresponding function control instructions based on the accurate interaction intention and the potential interaction demand of the occupant at present after receiving the demand confirmation instruction of the occupant, otherwise, generating corresponding function control instructions based on the accurate interaction intention of the occupant at present; and S53: receiving the execution result fed back by the vehicle-mounted device and feeding back the execution result to the occupant.

[0030] In the present embodiment, when the cross-modal attention matching fusion model performs feature fusion, the confidence of the fusion result is also output correspondingly. When the current potential interaction demand of the occupant is predicted, it is first determined whether the confidence of the potential interaction demand exceeds a preset threshold. Only when the confidence of the potential interaction demand is large enough (exceeds the preset threshold), the demand confirmation request of whether to confirm the potential interaction demand is correspondingly fed back to the occupant, so that the occupant needs to actively confirm whether the potential interaction demand exists. Only when the occupant confirms that the potential interaction demand exists (the demand confirmation instruction fed back by the occupant based on the demand confirmation request), corresponding function control instructions are generated for the accurate interaction intention and the potential interaction demand. Otherwise, it indicates that the confidence of the potential interaction demand is low at this time, which may be misrecognition, and only corresponding function control instructions are generated based on the accurate interaction intention of the occupant at present, which is more intelligent. Moreover, finally, the execution result fed back by the vehicle-mounted device is fed back to the occupant, so that the occupant can accurately know the execution result.

[0031] In an optional embodiment of the present application, the method further includes: calculating the uncertainty index of the fusion result based on the consistency and the confidence interval quantization method between the visual attention region features and the preliminary interaction intention fused by the cross-modal attention matching fusion model, the fusion result including the accurate interaction intention and the potential interaction demand of the occupant at present; determining the safety threshold of the uncertainty of the fusion result based on the current driving scene and driving environment of the vehicle. determining whether the uncertainty index of the fusion result exceeds the safety threshold, if yes, requesting the passenger to confirm the fusion result, otherwise generating the corresponding function control instruction based on the accurate interaction intention and the potential interaction demand of the passenger at present; and updating the parameters of the cross-modal attention matching fusion model in real time based on the confirmation result of the passenger on the fusion result.

[0032] In the embodiment, the uncertainty index of the fusion result is also calculated, and the safety threshold of the uncertainty of the fusion result is determined based on the current driving scene and driving environment of the vehicle (for example, the complexity of the driving road surface, the driving speed, etc.), so as to realize dynamic adjustment of the safety threshold. By determining whether the uncertainty index exceeds the safety threshold, when the operation is performed, an additional reconfirmation request is actively initiated to further clarify the real intention of the passenger, so as to avoid the safety risk caused by misrecognition or misoperation. If it is not exceeded, the function control instruction corresponding to the passenger at present can be generated based on the accurate interaction intention and the potential interaction demand according to the steps of the foregoing embodiment. In addition, the parameters of the cross-modal attention matching fusion model are updated in real time based on the request feedback result fed back by the passenger according to the reconfirmation request, and the uncertainty in subsequent interaction is dynamically adjusted to continuously improve the interaction accuracy and safety.

[0033] In specific implementation, when the interaction result of the vehicle device is fed back to the passenger, the demand confirmation request of whether to confirm the potential interaction demand is fed back to the passenger, and the reconfirmation request of the fusion result is fed back to the passenger, a natural voice interaction feedback can be generated by using a voice synthesis algorithm, or a visual feedback device such as an instrument panel and a central control display screen can be used, of course, the combination of the two is better. The demand confirmation instruction fed back by the passenger based on the demand confirmation request and the request feedback result fed back by the passenger according to the reconfirmation request can be non-contact feedback such as voice and gesture action of the passenger, or can be contact feedback such as actual key input and touch input.

[0034] In addition, it can be understood that the uncertainty index is a normalized parameter value, and its actual meaning represents a probability confidence. The safety threshold is generally inversely proportional to the driving safety reflected by the current driving scene and driving environment, that is, the lower the driving safety reflected by the current driving scene and driving environment, the higher the safety threshold, and the higher the driving safety reflected by the current driving scene and driving environment, the lower the safety threshold, thereby ensuring safety. In addition, the uncertainty index of the fusion result is continuously compared with the safety threshold to determine whether the uncertainty index exceeds the safety threshold, so that the reconfirmation request is fed back as long as the uncertainty index exceeds the safety threshold. Therefore, the safety threshold decreases in a negative logarithmic trend with the number of additional reconfirmation requests (will not be lower than the convergence line), preventing frequent initiation of interactive confirmation and improving the driving experience.

[0035] In an optional embodiment of the present application, the historical and current visual attention area features of the occupant and the preliminary interactive intention are stored in the form of a cross-modal knowledge graph. In this embodiment, the historical visual attention area features and the preliminary interactive intention are stored in the form of a cross-modal knowledge graph, which supports vectorization representation (such as entity and relationship embedding) of multi-modal data such as natural language, images, and videos, and can effectively enhance cross-modal understanding ability. In an optional embodiment of the present application, as shown in Figure 3 The step S3 specifically includes: S31: identifying whether there is ambiguous referential information in the preliminary interactive intention; S32: if there is the ambiguous referential information in the preliminary interactive intention, determining a referential target corresponding to the ambiguous referential information based on the current visual attention area feature of the occupant; and S33: fusing the referential target in the current visual attention area feature of the occupant and the ambiguous referential information by using a cross-modal attention matching fusion model to obtain the accurate interactive intention.

[0036] In this embodiment, whether there is ambiguous referential information in the preliminary interactive intention is identified. When there is obviously ambiguous referential information, the referential target corresponding to the ambiguous referential information is determined in combination with the real-time visual attention area feature, so that the referential target in the real-time visual attention area feature and the ambiguous referential information can be fused to form the accurate interactive intention, and multi-modal fusion is realized.

[0037] In practice, it can be understood that the ambiguous reference information refers to unclear speech, for example, "turn down a little", and the reference target corresponding to the ambiguous reference information in the real-time visual attention area feature refers to the corresponding vehicle-mounted equipment in the cabin, for example, an air conditioner or a sound equipment in the vehicle; in combination with the foregoing example, the accurate interaction intention is "reduce the temperature of the air conditioner on the driver's side by 2 degrees Celsius".

[0038] In an optional embodiment of the present application, the analyzing the predetermined visual information to correspondingly obtain and save the visual attention area feature of the occupant specifically comprises: performing feature extraction on the predetermined visual information to determine a line-of-sight attention area, a gesture action direction and a facial posture of the occupant; and obtaining and saving the visual attention area feature based on the line-of-sight attention area, the gesture action direction and the facial posture of the occupant.

[0039] In the embodiment, the line-of-sight attention area, the gesture action direction and the facial posture of the occupant are determined by performing feature extraction on the predetermined visual information, and the above features can obviously reflect the object target that the occupant currently wants to interact, so that the visual attention area feature of the occupant is quickly obtained.

[0040] In an optional embodiment of the present application, the analyzing the voice information to correspondingly obtain and save the preliminary interaction intention of the occupant specifically comprises: performing data preprocessing on the voice information to obtain a high-quality audio signal, and the data preprocessing at least includes noise reduction, echo cancellation and voice enhancement; converting the audio signal into text information by using a real-time voice recognition technology; and identifying the text information by using a natural language understanding algorithm model to obtain and save the preliminary interaction intention.

[0041] In the embodiment, for the voice information, data preprocessing is first performed to improve the quality of the audio signal, then the voice information is converted into text information by voice recognition, and then the interaction intention of the occupant can be obtained by simply understanding the text information, which is simple and efficient.

[0042] An interaction example of the vehicle-mounted intelligent interaction method according to the embodiment of the present application is as follows: (1) the occupant only issues a vague instruction by voice: "turn up a little"; (2) the line-of-sight and gesture of the occupant are determined based on visual analysis, and it is determined that the occupant points to the "instrument panel backlight area"; (3) cross-modal reference disambiguation is performed, and it is confirmed that the interaction intention of the occupant is "turn up the brightness of the instrument panel"; (4) According to the historical cross-modal context knowledge graph fusion, the potential demand of the passenger is predicted to be: "improve the brightness of the center control screen", therefore, the passenger is actively fed back "whether the brightness of the center control screen also needs to be adjusted"; (5) The fusion uncertainty is calculated and the safety threshold is dynamically adjusted according to the current highway condition, if the uncertainty index exceeds the safety threshold, the confirmation is actively initiated again: "whether to determine to adjust the brightness of the center control screen"; (6) Finally, the feedback of the passenger is determined to send the function control instruction to the vehicle-mounted device and feedback the execution result.

[0043] Another interaction example of the vehicle-mounted intelligent interaction method of the embodiment of the application is as follows: (1) The driver's eye closure exceeds a predetermined threshold (for example: continuous eye closure exceeds 3 seconds), the head posture droop angle exceeds a preset 15-degree angle, and the voice response frequency significantly decreases; (2) It is identified that the driver has high fatigue or sleep risk, and an active intervention strategy is immediately generated; (3) The system actively reminds through the vehicle-mounted voice: "You seem to be very tired, please rest as soon as possible", and automatically adjusts the air conditioner temperature to decrease by 2 degrees Celsius, increases the volume of the sound or switches the music style to assist the driver to boost the spirit; (4) If the fatigue state continues to exceed the predetermined time, a nearby rest station is further actively recommended, and related information is displayed on the instrument panel or the center control screen, and the driver is actively guided to rest as soon as possible to ensure the safety of driving.

[0044] On the other hand, as shown in Figure 4 The embodiment of the application further provides a vehicle-mounted intelligent interaction device 3 connected with a feature acquisition device 1 installed in the cabin, comprising a processor 30, a memory 32, and a computer program stored in the memory 32 and configured to be executed by the processor 30, when the processor 30 executes the computer program, the vehicle-mounted intelligent interaction method as described in any of the above embodiments is realized.

[0045] In specific implementation, the feature acquisition device 1 is composed of a monitoring camera and a microphone array to realize the acquisition of passenger images and voice, wherein the monitoring camera can adopt a high-definition wide-angle infrared monitoring camera, so as to realize the visual information acquisition in a large visual angle and night scene; the microphone array is composed of multiple microphones.

[0046] For example, the computer program can be divided into one or more modules / units, which are stored in the memory 32 and executed by the processor to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program in the vehicle intelligent interaction device 3. For example, the computer program can be divided into Figure 5 The function modules in the vehicle intelligent interaction device 3, wherein the device control module 30, the information analysis module 32, the current intention determination module 34, the potential demand determination module 36 and the instruction output module 38 correspond to the above steps S1-S5, respectively.

[0047] The vehicle intelligent interaction device 3 can be a desktop computer, a notebook, a palm computer, a cloud server and the like. The vehicle intelligent interaction device 3 can include, but is not limited to, the processor 30 and the memory 32. Those skilled in the art can understand that the schematic diagram is only an example of the vehicle intelligent interaction device 3, and does not constitute a limitation on the vehicle intelligent interaction device 3, and can include more or fewer components than the diagram, or combine certain components, or different components, for example, the vehicle intelligent interaction device 3 can also include an input / output device, a network access device, a bus and the like.

[0048] The processor 30 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic components, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The processor 30 is the control center of the vehicle intelligent interaction device 3, and connects all parts of the vehicle intelligent interaction device 3 through various interfaces and lines.

[0049] The memory 32 can be used to store the computer programs and / or modules, and the processor 30 realizes various functions of the in-vehicle intelligent interaction device 3 by running or executing the computer programs and / or modules stored in the memory 32, and calling the data stored in the memory 32. The memory 32 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application program required by a function (such as a graphic recognition function, a graphic layering function, etc.), and the like; and the data storage area can store data (such as graphic data, etc.) created according to the use of the control device, and the like. In addition, the memory 32 can include a high-speed random access memory, and can also include a non-volatile memory, for example, a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0050] If the functions of the embodiments of the present application are realized in the form of software function modules or units and sold or used as independent products, they can be stored in a computer device readable storage medium. Based on such understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware, and the computer program can be stored in a computer readable storage medium. When the processor 30 executes the computer program, the steps of the above-mentioned various method embodiments can be realized. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the contents included in the computer readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, the computer readable medium does not include electrical carrier signals and telecommunication signals.

[0051] In still another aspect, the embodiments of the present application further provide a computer readable storage medium, which includes a stored computer program, wherein when the computer program runs, the device where the computer readable storage medium is located executes the in-vehicle intelligent interaction method according to any one of the above-mentioned embodiments.

[0052] In another aspect, an embodiment of the present application provides a computer program product comprising a computer program which, when executed by a processor, implements the steps of the vehicle intelligent interaction method according to any of the above embodiments.

[0053] The various embodiments are described in the specification in a progressive manner, each of which focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other.

[0054] The embodiments of the present application are described above in combination with the drawings, but the present application is not limited to the specific embodiments described above, and the specific embodiments described above are merely illustrative rather than limiting, and those of ordinary skill in the art can make many forms under the inspiration of the present application without departing from the purpose of the present application and the scope protected by the claims, which are all within the protection scope of the present application.

Claims

1. A vehicle-mounted intelligent interaction method, characterized in that, The method includes the following steps: The system controls the feature acquisition device installed in the cockpit to collect the occupants' predetermined visual and voice information in real time. The predetermined visual information includes at least: facial information, gesture information, posture information, and eye movement information. The predetermined visual information is analyzed to obtain and save the visual attention area features of the occupant, and the voice information is also analyzed to obtain and save the occupant's initial interaction intention. A cross-modal attention matching fusion model is used to fuse the current visual attention region features and the initial interaction intent of the occupant and eliminate ambiguity to obtain the occupant's current precise interaction intent; The cross-modal attention matching fusion model is used to fuse the occupant's historical and current visual attention region features with the initial interaction intent to obtain the occupant's current potential interaction needs; and Based on the occupant's current precise interaction intent and potential interaction needs, a function control command is generated to control the corresponding in-vehicle equipment, and the function control command is sent to the in-vehicle equipment through the vehicle communication network.

2. The in-vehicle intelligent interaction method as described in claim 1, characterized in that, When using the cross-modal attention matching fusion model to obtain the occupant's current potential interaction needs, the confidence level of the potential interaction needs is also output; the generation of function control commands for controlling the corresponding in-vehicle equipment based on the occupant's current precise interaction intent and potential interaction needs includes: Determine whether the confidence level of the potential interaction request exceeds a preset threshold; If the confidence level of the potential interaction request exceeds the preset threshold, then the occupant is requested to confirm the potential interaction request. Upon receiving the occupant's confirmation instruction, corresponding function control instructions are generated based on the occupant's current precise interaction intent and the potential interaction request, respectively. Otherwise, corresponding function control instructions are generated based on the occupant's current precise interaction intent. The system receives the execution result from the on-board equipment and then reports the execution result back to the passenger.

3. The in-vehicle intelligent interaction method as described in claim 2, characterized in that, The method further includes: The uncertainty index of the fusion result is calculated by the quantification method of the consistency and confidence interval between the visual attention region features and the initial interaction intention fused by the cross-modal attention matching fusion model. The fusion result includes the occupant's current precise interaction intention and potential interaction needs. A safety threshold for the uncertainty of the fusion result is determined based on the vehicle's current driving scenario and driving environment; Determine whether the uncertainty index of the fusion result exceeds the safety threshold. If so, request the occupant to confirm the fusion result; otherwise, generate the corresponding function control command based on the occupant's current precise interaction intention and potential interaction needs. Based on the occupants' confirmation of the fusion result, the parameters of the cross-modal attention matching fusion model are updated in real time.

4. The in-vehicle intelligent interaction method as described in claim 1, characterized in that, The occupant's historical and current visual attention region features and initial interaction intentions are stored in the form of a cross-modal knowledge graph.

5. The in-vehicle intelligent interaction method as described in claim 1, characterized in that, The step of employing a cross-modal attention matching fusion model to fuse the occupant's current visual attention region features and the initial interaction intent, and to eliminate ambiguity to obtain the occupant's current precise interaction intent, specifically includes: Identify and analyze whether there is ambiguous referential information in the initial interaction intent; If the preliminary interaction intent contains the ambiguous referential information, then the referential target corresponding to the ambiguous referential information is determined based on the current visual attention region features of the occupant; and A cross-modal attention matching fusion model is used to fuse the referential target and the fuzzy referential information in the current visual attention region features of the occupant to obtain the precise interaction intent.

6. The in-vehicle intelligent interaction method as described in claim 1, characterized in that, The analysis of the predetermined visual information to obtain and save the visual attention region features of the occupant specifically includes: Feature extraction is performed on the predetermined visual information to determine the occupant's gaze area, gesture direction, and facial posture; and The visual attention area features are obtained and saved based on the occupant's gaze area, gesture direction, and facial posture.

7. The in-vehicle intelligent interaction method as described in claim 1, characterized in that, The process of analyzing the voice information to obtain and save the occupant's initial interaction intent specifically includes: The speech information is preprocessed to obtain a high-quality audio signal. The preprocessing includes at least noise reduction, echo cancellation, and speech enhancement. The audio signal is converted into text information using real-time speech recognition technology; and The text information is identified using a natural language understanding algorithm model to obtain and save the initial interaction intent.

8. A vehicle-mounted intelligent interactive device, connected to a feature acquisition device installed in the cockpit, characterized in that, The device includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the in-vehicle intelligent interaction method as described in any one of claims 1-7.

9. A computer-readable storage medium comprising a stored computer program, wherein, When the computer program is running, it controls the device containing the computer-readable storage medium to perform the in-vehicle intelligent interaction method as described in any one of claims 1-7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the in-vehicle intelligent interaction method as described in any one of claims 1-7.