Vehicle-mounted voice interaction method and device based on environment perception, electronic equipment and storage medium

By collecting the perception data of the interior and exterior environment of the vehicle, identifying the body shapes of the driver and the occupant, and determining the voice interaction mode based on this information, the problem that the existing vehicle voice interaction system cannot adapt to the body shape changes is solved, and intelligent vehicle voice interaction is achieved, improving the interaction effect.

CN120089133APending Publication Date: 2025-06-03GAC HONDA AUTOMOBILE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510155333.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing vehicle voice interaction system cannot fully adapt to the changes in the body shape and driving environment of the driver or occupant, and lacks intelligence, and cannot meet the user's vehicle voice interaction needs.

Method used

By collecting the perception data of the interior and exterior environment of the vehicle, the body recognition results of the driver and the occupant are generated, and the voice interaction mode is determined based on these results to achieve intelligent vehicle voice interaction.

Benefits of technology

It realizes intelligent voice interaction in the vehicle, fully adapts to the changes in the body shape of the driver and the occupants, improves the effect of voice interaction in the vehicle, and meets the needs of users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089133A_ABST
    Figure CN120089133A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle-mounted voice interaction method and device based on environmental perception, electronic equipment and a storage medium. The method comprises the following steps: acquiring environmental perception data inside and outside a vehicle; a first posture recognition result and a second posture recognition result are generated according to the vehicle interior and exterior environment sensing data, the first posture recognition result corresponds to a driver, and the second posture recognition result corresponds to a passenger in the vehicle; and determining a voice interaction mode according to the first posture recognition result and the second posture recognition result, and performing vehicle-mounted voice interaction according to the voice interaction mode. The vehicle-mounted voice interaction method and device can recognize the postures of the driver and the passengers in the vehicle based on the environment sensing data inside and outside the vehicle, achieve intelligent vehicle-mounted voice interaction, fully adapt to the posture changes of the driver and the passengers in the vehicle, meet the vehicle-mounted voice interaction requirements of the user, improve the vehicle-mounted voice interaction effect, and can be widely applied to the technical field of vehicle-machine interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle-mounted voice interaction technology, and particularly relates to a vehicle-mounted voice interaction method, device, electronic device, and storage medium based on environmental perception. Background Art

[0002] With the continuous development of automotive technology, in-vehicle voice interaction systems have become one of the standard configurations of modern vehicles. These systems allow drivers or passengers to interact with the vehicle through voice, thereby improving driving safety and convenience. However, existing in-vehicle voice interaction systems usually only provide basic instructions and information feedback, cannot fully adapt to the changes in the postures of drivers or passengers and the driving environment, lack intelligence, and cannot meet the in-vehicle voice interaction needs of users. Summary of the Invention

[0003] The main purpose of the embodiments of this application is to propose a vehicle-mounted voice interaction method, device, electronic device, and storage medium based on environmental perception, which can achieve intelligent vehicle-mounted voice interaction and meet the in-vehicle voice interaction needs of users.

[0004] On the one hand, the embodiments of this application propose a vehicle-mounted voice interaction method based on environmental perception, and the method includes the following steps:

[0005] Collect environmental perception data inside and outside the vehicle;

[0006] Generate a first posture recognition result and a second posture recognition result according to the environmental perception data inside and outside the vehicle, where the first posture recognition result corresponds to the driver, and the second posture recognition result corresponds to the passengers inside the vehicle;

[0007] Determine a voice interaction mode according to the first posture recognition result and the second posture recognition result, and perform vehicle-mounted voice interaction according to the voice interaction mode.

[0008] In some embodiments, the collecting of the environmental perception data inside and outside the vehicle specifically includes:

[0009] Dynamically collect first posture perception data corresponding to the driver and second posture perception data corresponding to the passengers inside the vehicle. The first posture perception data includes a first sitting posture image, first seat pressure information, and first physiological data corresponding to the driver, and the second posture perception data includes a second sitting posture image, second seat pressure information, and second physiological data corresponding to the passengers inside the vehicle;

[0010] Dynamically collect the noise intensity inside and outside the vehicle and the light intensity inside the vehicle.

[0011] In some embodiments, the method further includes:

[0012] Determine an adaptive adjustment result based on the vehicle interior and exterior environment perception data, and perform vehicle machine adaptive adjustment according to the adaptive adjustment result.

[0013] In some embodiments, the determining an adaptive adjustment result based on the vehicle interior and exterior environment perception data and performing vehicle machine adaptive adjustment according to the adaptive adjustment result specifically includes:

[0014] Obtain the current vehicle machine sound intensity, and perform adaptive adjustment on the current vehicle machine sound intensity according to the vehicle interior and exterior noise intensity;

[0015] Obtain the current in-vehicle display screen brightness, and perform adaptive adjustment on the current in-vehicle display screen brightness according to the vehicle interior light intensity.

[0016] In some embodiments, the generating a first body gesture recognition result and a second body gesture recognition result based on the vehicle interior and exterior environment perception data specifically includes:

[0017] Obtain a body gesture recognition model;

[0018] Input the first body gesture perception data into the body gesture recognition model, perform body gesture recognition using the body gesture recognition model, and determine the corresponding current body gesture of the driver from the preset driver body gestures;

[0019] Input the second body gesture perception data into the body gesture recognition model, perform body gesture recognition using the body gesture recognition model, and determine the corresponding current body gesture of the passenger from the preset passenger body gestures.

[0020] In some embodiments, the determining a voice interaction mode based on the first body gesture recognition result and the second body gesture recognition result and performing in-vehicle voice interaction according to the voice interaction mode specifically includes:

[0021] Obtain the first in-vehicle voice control area corresponding to the driver, determine a first target voice interaction mode from multiple preset driver voice interaction modes according to the current body gesture of the driver, and perform voice interaction control on the first in-vehicle voice control area according to the first target voice interaction mode, where the driver voice interaction mode includes voice prompt frequency, voice prompt content, and voice volume;

[0022] Obtain the second in-vehicle voice control area corresponding to the in-vehicle passenger, determine a second target voice interaction mode from multiple preset passenger voice interaction modes according to the current body gesture of the passenger, and perform voice interaction control on the second in-vehicle voice control area according to the second target voice interaction mode, where the passenger preset voice interaction mode includes voice prompt frequency, voice prompt content, and voice volume.

[0023] In some embodiments, the method further includes:

[0024] Obtaining a sample data set, where the sample data set includes a plurality of in-vehicle and out-of-vehicle environment perception data with postural recognition labels;

[0025] Constructing the postural recognition model and training the postural recognition model using the sample data set.

[0026] On the other hand, an embodiment of the present application proposes an in-vehicle voice interaction device based on environment perception, and the device includes:

[0027] A first module for collecting in-vehicle and out-of-vehicle environment perception data;

[0028] A second module for generating a first postural recognition result and a second postural recognition result according to the in-vehicle and out-of-vehicle environment perception data, where the first postural recognition result corresponds to the driver and the second postural recognition result corresponds to the in-vehicle occupants;

[0029] A third module for determining a voice interaction mode according to the first postural recognition result and the second postural recognition result, and performing in-vehicle voice interaction according to the voice interaction mode.

[0030] On the other hand, an embodiment of the present application proposes an electronic device, and the electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the foregoing in-vehicle voice interaction method is implemented.

[0031] On the other hand, an embodiment of the present application proposes a computer-readable storage medium, and the computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the foregoing in-vehicle voice interaction method is implemented.

[0032] The embodiments of the present application at least include the following beneficial effects: An in-vehicle voice interaction method, device, electronic device, and storage medium based on environment perception provided by the present application collect in-vehicle and out-of-vehicle environment perception data, generate a first postural recognition result and a second postural recognition result according to the in-vehicle and out-of-vehicle environment perception data, where the first postural recognition result corresponds to the driver and the second postural recognition result corresponds to the in-vehicle occupants, determine a voice interaction mode according to the first postural recognition result and the second postural recognition result, and perform in-vehicle voice interaction according to the voice interaction mode. The present application can recognize the postures of the driver and the in-vehicle occupants based on the in-vehicle and out-of-vehicle environment perception data, realize intelligent in-vehicle voice interaction, fully adapt to the postural changes of the driver and the in-vehicle occupants, meet the in-vehicle voice interaction needs of users, and improve the in-vehicle voice interaction effect. Description of the Drawings

[0033] Figure 1It is a flowchart of a vehicle-mounted voice interaction method based on environmental perception provided by an embodiment of the present application;

[0034] Figure 2 It is a schematic diagram of the division of the in-vehicle voice control area in an embodiment of the present application;

[0035] Figure 3 It is a schematic structural diagram of a vehicle-mounted voice interaction device based on environmental perception provided by an embodiment of the present application;

[0036] Figure 4 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. Specific embodiments

[0037] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application detailed in the appended claims.

[0038] It can be understood that the terms "first", "second", etc. used in the present application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if", "when" as used herein may be interpreted as "when...", "when...", or "in response to determining".

[0039] The terms "at least one", "a plurality", "each", "any one", etc. used in the present application, at least one includes one, two or more, a plurality includes two or more, each refers to each of the corresponding plurality, and any one refers to any one of the plurality.

[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0041] It should be noted that in each specific embodiment of the present application, when it comes to relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of the present application will be obtained.

[0042] Referring to Figure 1 , Figure 1 is an optional flowchart of a vehicle-mounted voice interaction method based on environmental perception provided by an embodiment of the present application. The method may include but is not limited to steps S101 to S103:

[0043] Step S101, collect vehicle interior and exterior environmental perception data;

[0044] Step S102, generate a first body posture recognition result and a second body posture recognition result according to the vehicle interior and exterior environmental perception data, wherein the first body posture recognition result corresponds to the driver, and the second body posture recognition result corresponds to the vehicle occupants;

[0045] Step S103, determine the voice interaction mode according to the first body posture recognition result and the second body posture recognition result, and perform vehicle-mounted voice interaction according to the voice interaction mode.

[0046] In some embodiments, by connecting a variety of vehicle-mounted sensors, the environmental perception data inside and outside the vehicle is dynamically collected, such as light intensity, noise level, and the body postures of the people inside the vehicle. According to the vehicle interior and exterior environmental perception data, the voice interaction mode is determined, and vehicle-mounted voice interaction is performed according to the voice interaction mode to achieve intelligent vehicle-mounted voice interaction.

[0047] In some embodiments, step S101 may include but is not limited to steps S201 to S202:

[0048] Step S201, dynamically collect the first body posture perception data corresponding to the driver and the second body posture perception data corresponding to the vehicle occupants. The first body posture perception data includes the first sitting posture image corresponding to the driver, the first seat pressure information, and the first physiological data. The second body posture perception data includes the second sitting posture image corresponding to the vehicle occupants, the second seat pressure information, and the second physiological data;

[0049] Step S202, dynamically collect the noise intensity inside and outside the vehicle and the light intensity inside the vehicle.

[0050] In some embodiments, in-vehicle sensors may include, but are not limited to, biosensors, seat pressure sensors, light sensors, noise sensors, and cameras. Specifically, the biosensor is installed on the seat belt to collect physiological data corresponding to the driver and the vehicle occupants, i.e., the above-mentioned first physiological data and second physiological data. The physiological data may include, but is not limited to, heart sound, blood pressure, pulse, blood flow, body temperature, respiratory rate, blood glucose level, and blood components. Seat pressure sensors are installed on each seat in the vehicle, and the seat pressure corresponding to the driver and the vehicle occupants is detected through the seat pressure sensors on each seat in the vehicle to obtain the above-mentioned first seat pressure information and second seat pressure information. A camera is installed in the vehicle, and the driver and the vehicle occupants are photographed through the camera to collect images of the driver and the vehicle occupants when sitting on the seats in the vehicle, i.e., the above-mentioned first sitting posture images and second sitting posture images. The light sensor is used to monitor the brightness change inside and outside the vehicle and collect the corresponding light intensity inside the vehicle. The noise sensor is used to detect the noise level inside and outside the vehicle to obtain the corresponding noise intensity inside and outside the vehicle.

[0051] In some embodiments, step S102 may include, but is not limited to, steps S301 to S303:

[0052] Step S301, obtain a body posture recognition model;

[0053] Step S302, input the first body posture perception data into the body posture recognition model, use the body posture recognition model to perform body posture recognition, and determine the corresponding current body posture of the driver from the preset driver body postures;

[0054] Step S303, input the second body posture perception data into the body posture recognition model, use the body posture recognition model to perform body posture recognition, and determine the corresponding current body posture of the occupant from the preset occupant body postures.

[0055] In some embodiments, optionally, a body posture recognition model is constructed. The body posture recognition model is a classification model, such as a random forest model, a decision tree model, or a support vector machine model, etc. Then, a sample data set is obtained. The sample data set includes multiple in-vehicle and out-of-vehicle environment perception data with body posture recognition labels. The body posture recognition model is trained using the sample data set. The in-vehicle and out-of-vehicle environment perception data includes corresponding sitting posture image data, seat pressure sensor data, and biosensor data, etc.

[0056] The first body posture perception data and the second body posture perception data are respectively input into the trained body posture recognition model to perform body posture recognition, and the corresponding current body posture of the occupant and the current body posture of the driver are determined.

[0057] Exemplarily, the preset driver body postures may include, but are not limited to, a forward-leaning and focused body posture, a distracted driving body posture, and a fatigued driving body posture, specifically as follows:

[0058] Focused driving posture: manifested as the body leaning forward significantly, the head stretching forward, and the eyes looking intently ahead of the vehicle or in a specific direction;

[0059] Distracted driving posture: manifested as the head turning to other directions inside the vehicle. For example, when looking at a mobile phone or talking to a passenger in the vehicle, the head will tilt to one side, the body will twist slightly with the rotation of the head, and one hand will leave the steering wheel to operate the mobile phone or do other things;

[0060] Drowsy driving posture: manifested as the head drooping or tilting to one side involuntarily, the eyes half - open and half - closed, and even experiencing short - term eye - closing. The body will tilt to one side of the seat, the back will no longer be straight, and the grip strength of both hands on the steering wheel may become loose;

[0061] The preset occupant postures may include but are not limited to normal posture, relaxed reclined posture, and sound asleep posture, as follows:

[0062] Normal posture: manifested as the passenger's body being straight, the back fitting against the seat backrest, the head remaining upright, the feet flat on the floor, and both hands naturally placed on the legs or holding the armrests, maintaining a normal sitting posture;

[0063] Relaxed reclined posture: manifested as the body tilting to one side of the seat, the head leaning on the seat headrest or the window, the limbs being in a relatively relaxed state, and / or crossing one leg over the other;

[0064] Sound asleep posture: manifested as the head tilting to one side, leaning on the seat, the window, or the adjacent passenger, the mouth slightly open, the body swaying slightly with the movement of the vehicle, and the breathing being even and steady.

[0065] In some embodiments, step S103 may include but is not limited to steps S401 to S402:

[0066] Step S401, obtain the first in - vehicle voice control area corresponding to the driver, determine the first target voice interaction mode from multiple driver - preset voice interaction modes according to the driver's current posture, and perform voice interaction control on the first in - vehicle voice control area according to the first target voice interaction mode. The driver voice interaction mode includes voice prompt frequency, voice prompt content, and voice volume;

[0067] Step S402, obtain the second in - vehicle voice control area corresponding to the vehicle occupant, determine the second target voice interaction mode from multiple occupant - preset voice interaction modes according to the occupant's current posture, and perform voice interaction control on the second in - vehicle voice control area according to the second target voice interaction mode. The occupant preset voice interaction mode includes voice prompt frequency, voice prompt content, and voice volume.

[0068] In some embodiments, referring toFigure 2 , Figure 2 This is an optional schematic diagram of the in-vehicle voice control area in an embodiment of the present application. Assuming that the maximum number of passengers in the vehicle is 5, including in-vehicle voice control areas A to E. In-vehicle voice control area A is the in-vehicle voice control area corresponding to the driver. In-vehicle voice control area B is the in-vehicle voice control area corresponding to the front-row passenger. In-vehicle voice control area C is the in-vehicle voice control area corresponding to the rear-row passenger 1 directly behind the driver. In-vehicle voice control area E is the in-vehicle voice control area corresponding to the rear-row passenger 3 directly behind the front-row passenger. In-vehicle voice control area D is the in-vehicle voice control area corresponding to the rear-row passenger 2 between the rear-row passenger 3 and the rear-row passenger 1. And so on. The in-vehicle voice control area can be divided according to different vehicle models and the installation positions of in-vehicle voice devices without limitation.

[0069] In some embodiments, different current postures of the driver correspond to different voice interaction modes. Exemplarily, when the driver's current posture is a focused driving posture, the corresponding first target voice interaction mode is: adjusting the voice prompt frequency to a first preset frequency to decrease and avoid disturbing the driver, adjusting the voice volume to a preset volume, and at the same time, giving a voice prompt when encountering important road condition information or safety warnings; when the driver's current posture is a focused driving posture, the corresponding first target voice interaction mode is: increasing the voice prompt frequency to a preset second frequency, increasing the voice volume to a preset second volume, and giving a timely reminder to the driver, and the voice prompt content is "Please pay attention to driving safety and avoid fatigue driving", and so on. Similarly, different current postures of the passengers also correspond to different voice interaction modes. For example, when the passenger's current posture is a relaxed reclining posture or a sleeping posture, the corresponding second target voice interaction mode is: decreasing the voice prompt frequency to a preset third frequency, decreasing the voice volume to a preset third volume, and avoiding disturbing the passenger's rest. For information irrelevant to the passenger, such as prompts related to other routes in the navigation, no voice prompt is given temporarily.

[0070] In some embodiments, the above vehicle-mounted voice interaction method may further include step S501:

[0071] Step S501: Determine an adaptive adjustment result according to the vehicle interior and exterior environment perception data, and perform vehicle machine adaptive adjustment according to the adaptive adjustment result.

[0072] In some embodiments, step S501 may include but is not limited to steps S601 to S602:

[0073] Step S601: Obtain the current vehicle machine sound intensity, and perform adaptive adjustment on the current vehicle machine sound intensity according to the vehicle interior and exterior noise intensity;

[0074] Step S602: Obtain the current brightness of the in-vehicle display screen, and adaptively adjust the current brightness of the in-vehicle display screen according to the light intensity inside the vehicle.

[0075] In some embodiments, the current brightness of the in-vehicle display screen is automatically adjusted according to the light intensity inside the vehicle, so as to provide a clear visual experience under various lighting conditions. For example, a first mapping relationship table between the light intensity inside the vehicle and the brightness of the in-vehicle display screen is established. According to the light intensity inside the vehicle, the target brightness of the in-vehicle display screen corresponding to the light intensity inside the vehicle is determined from the first mapping relationship table, and the current brightness of the in-vehicle display screen is gradually adjusted to the target brightness of the in-vehicle display screen.

[0076] In some embodiments, the current sound intensity of the in-vehicle head unit is automatically adjusted according to the noise intensity inside and outside the vehicle to avoid the sound being too loud or too small and affecting the passenger experience. For example, a second mapping relationship table between the noise intensity outside the vehicle and the sound intensity of the in-vehicle head unit is established, and a third mapping relationship table between the noise intensity inside the vehicle and the sound intensity of the in-vehicle head unit is established. When the noise intensity outside the vehicle is greater than the noise intensity inside the vehicle, the target sound intensity of the in-vehicle head unit corresponding to the noise intensity outside the vehicle is determined from the second mapping relationship table, and the current sound intensity of the in-vehicle head unit is gradually adjusted to the target sound intensity of the in-vehicle head unit. When the noise intensity inside the vehicle is greater than the noise intensity outside the vehicle, the target sound intensity of the in-vehicle head unit corresponding to the noise intensity inside the vehicle is determined from the third mapping relationship table, and the current sound intensity of the in-vehicle head unit is gradually adjusted to the target sound intensity of the in-vehicle head unit.

[0077] Refer to Figure 3 , Figure 3 FIG.

[0078] is an optional structural schematic diagram of an in-vehicle voice interaction device based on environmental perception provided by an embodiment of the present application. The device is used to implement the above-mentioned in-vehicle voice interaction method, and the device may include:

[0079] The first module is used to collect environmental perception data inside and outside the vehicle;

[0080] The second module is used to generate a first body gesture recognition result and a second body gesture recognition result according to the environmental perception data inside and outside the vehicle. The first body gesture recognition result corresponds to the driver, and the second body gesture recognition result corresponds to the passengers inside the vehicle;

[0081] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present application. The functions specifically implemented by the device embodiments of the present application are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0082] An embodiment of the present application further provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned vehicle-mounted voice interaction method is implemented. The electronic device can be any intelligent terminal including a tablet computer, etc.

[0083] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present application. The functions specifically implemented by the device embodiments of the present application are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0084] Please refer to Figure 4 , Figure 4 , which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0085] A processor 901, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit, central processor), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0086] A memory 902, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 902, and the processor 901 is used to call and execute the vehicle-mounted voice interaction method of the embodiments of the present application;

[0087] An input / output interface 903, which is used to implement information input and output;

[0088] A communication interface 904, which is used to implement communication and interaction between the device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0089] A bus 905, which transmits information between various components of the device (such as the processor 901, the memory 902, the input / output interface 903, and the communication interface 904);

[0090] Among them, the processor 901, the memory 902, the input / output interface 903, and the communication interface 904 are communicatively connected to each other inside the device through the bus 905.

[0091] The embodiment of the present application also provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the above-mentioned in-vehicle voice interaction method.

[0092] It can be understood that the content in the above method embodiments is applicable to this storage medium embodiment. The functions specifically implemented by this storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0093] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0094] The embodiment of the present application also provides a vehicle, which includes the electric drive assembly of the above in-vehicle voice interaction device or electronic device. Specifically, the vehicle can be a private car, such as a sedan, an SUV, an MPV, or a pickup truck, etc. The vehicle can also be an operating vehicle, such as a minibus, a bus, a small truck, or a large trailer, etc. The vehicle can be a fuel vehicle or a new energy vehicle. When the vehicle is a new energy vehicle, it can be a hybrid vehicle or a pure electric vehicle.

[0095] An in-vehicle voice interaction method, device, electronic device, and storage medium based on environmental perception provided by the embodiment of the present application can identify the postures of the driver and vehicle occupants based on in-vehicle and out-of-vehicle environmental perception data, implement intelligent in-vehicle voice interaction, fully adapt to the posture changes of the driver and vehicle occupants, meet the in-vehicle voice interaction needs of users, and improve the in-vehicle voice interaction effect.

[0096] The embodiments described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0097] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.

[0098] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0099] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0100] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of this application and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0101] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one)" or similar expressions below refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0102] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0103] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0104] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0105] It should be recognized that embodiments of the present invention can be implemented or carried out by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable memory. The method can be implemented in a computer program using standard programming techniques including a non-transitory computer-readable storage medium configured with the computer program, wherein the storage medium so configured causes the computer to operate in a specific and predefined manner - according to the methods and drawings described in the specific embodiments. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if desired, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. In addition, for this purpose the program is capable of running on a special-purpose integrated circuit programmed for this purpose.

[0106] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0107] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings, which does not limit the scope of rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall be within the scope of rights of the embodiments of this application.

Claims

1. A vehicle-mounted voice interaction method based on environment perception, characterized in that: The method comprises the following steps: Collecting environmental perception data inside and outside the vehicle; Generate a first body recognition result and a second body recognition result according to the vehicle interior and exterior environment perception data, wherein the first body recognition result corresponds to the driver, and the second body recognition result corresponds to the vehicle occupant; A voice interaction mode is determined according to the first body posture recognition result and the second body posture recognition result, and in-vehicle voice interaction is performed according to the voice interaction mode.

2. The vehicle-mounted voice interaction method according to claim 1, characterized in that: The collecting of vehicle internal and external environment perception data specifically includes: Dynamically collecting first body posture perception data corresponding to the driver and second body posture perception data corresponding to the vehicle occupant, wherein the first body posture perception data includes a first sitting posture image corresponding to the driver, first seat pressure information, and first physiological data, and the second body posture perception data includes a second sitting posture image corresponding to the vehicle occupant, second seat pressure information, and second physiological data; Dynamically collect the noise intensity inside and outside the car and the light intensity inside the car.

3. The vehicle-mounted voice interaction method according to claim 2, characterized in that: The method further comprises: An adaptive adjustment result is determined according to the vehicle interior and exterior environment perception data, and a vehicle computer adaptive adjustment is performed according to the adaptive adjustment result.

4. The vehicle-mounted voice interaction method according to claim 3, characterized in that: The determining of the adaptive adjustment result according to the vehicle interior and exterior environment perception data, and performing vehicle computer adaptive adjustment according to the adaptive adjustment result, specifically includes: Acquire the current vehicle sound intensity, and adaptively adjust the current vehicle sound intensity according to the noise intensity inside and outside the vehicle; The current brightness of the vehicle display screen is obtained, and the brightness of the current vehicle display screen is adaptively adjusted according to the light intensity inside the vehicle.

5. The vehicle-mounted voice interaction method according to claim 2, characterized in that: The generating a first body recognition result and a second body recognition result according to the vehicle interior and exterior environment perception data specifically includes: Obtaining a body posture recognition model; Inputting the first posture sensing data into the posture recognition model, performing posture recognition using the posture recognition model, and determining the corresponding current posture of the driver from preset driver postures; The second body posture sensing data is input into the body posture recognition model, body posture recognition is performed using the body posture recognition model, and the corresponding current body posture of the occupant is determined from preset occupant body postures.

6. The vehicle-mounted voice interaction method according to claim 5, characterized in that: The determining of a voice interaction mode according to the first body recognition result and the second body recognition result, and performing in-vehicle voice interaction according to the voice interaction mode specifically includes: Acquire a first in-vehicle voice control area corresponding to the driver, determine a first target voice interaction mode from a plurality of driver preset voice interaction modes according to the current body posture of the driver, and perform voice interaction control on the first in-vehicle voice control area according to the first target voice interaction mode, wherein the driver voice interaction mode includes voice prompt frequency, voice prompt content, and voice volume; Obtain a second in-vehicle voice control area corresponding to the in-vehicle occupant, determine a second target voice interaction mode from a plurality of occupant preset voice interaction modes according to the occupant's current body posture, and perform voice interaction control on the second in-vehicle voice control area according to the second target voice interaction mode, wherein the occupant preset voice interaction mode includes voice prompt frequency, voice prompt content, and voice volume.

7. The vehicle-mounted voice interaction method according to claim 5, characterized in that: The method further comprises: Acquire a sample data set, wherein the sample data set includes a plurality of vehicle interior and exterior environment perception data with body recognition tags; The posture recognition model is constructed, and the posture recognition model is trained using the sample data set.

8. A vehicle-mounted voice interaction device based on environment perception, characterized in that: The device comprises: The first module is used to collect environmental perception data inside and outside the vehicle; A second module is used to generate a first body recognition result and a second body recognition result according to the vehicle interior and exterior environment perception data, wherein the first body recognition result corresponds to the driver, and the second body recognition result corresponds to the vehicle occupant; The third module is used to determine a voice interaction mode according to the first posture recognition result and the second posture recognition result, and perform in-vehicle voice interaction according to the voice interaction mode.

9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the in-vehicle voice interaction method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the in-vehicle voice interaction method according to any one of claims 1 to 7 is implemented.