Multi-screen interaction method, device and equipment and computer readable storage medium

By identifying audio signals and seat sensors to determine the user's position, the system controls the in-vehicle equipment in the target sound zone to interact with the user, solving the problems of inconvenient interaction for rear-seat users and interference for front-seat users, and achieving a more efficient multi-screen interactive experience.

CN120921906APending Publication Date: 2025-11-11BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510806379.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-03-25
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

When rear-seat users interact with the vehicle's infotainment system via voice commands, the central control screen's response is inconvenient and interferes with the driving of front-seat users.

Method used

By identifying audio signals to determine the user's location, the in-vehicle device controlling the target sound zone interacts with the user, distinguishes between the user's vocal sound zone and the seated sound zone, and takes into account the brain-buttock separation sitting posture to improve the accuracy of target sound zone judgment.

Benefits of technology

It improves the user interaction experience, reduces interference with users in other audio zones, and ensures that users can effectively interact with the in-vehicle equipment in their respective audio zones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120921906A_ABST
    Figure CN120921906A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-screen interaction method, device and equipment and a computer readable storage medium, and the method comprises the steps: determining a seating sound region as a target sound region of a user if a sound production sound region corresponding to a user generating an audio signal in a vehicle is not consistent with the seating sound region of the user; and controlling the vehicle-mounted equipment corresponding to the target sound area to interact with the user. When the target sound area is determined, the pronunciation sound area and the seating sound area of the user are distinguished, the conditions that the user is in a brain-hip separation sitting posture in the sound area dimension and the like are fully considered, the accuracy of target sound area judgment is improved, the vehicle-mounted equipment corresponding to the target sound area is controlled to interact with the user, and the user experience is improved. Therefore, the user can interact with the vehicle-mounted device corresponding to the sound area where the user is located, the interaction experience of the user is improved, and meanwhile interference to users in other sound areas is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese Patent No. 202210306313.4, filed on March 25, 2022. Technical Field

[0002] This disclosure relates to the field of human-computer interaction technology, and in particular to a multi-screen interaction method, apparatus, device, and computer-readable storage medium. Background Technology

[0003] With the continuous development of the automotive industry, multiple screens are installed in the car for user convenience; for example, a typical car has two screens installed, in addition to the central control screen corresponding to the driver's seat, a rear screen is also installed.

[0004] Typically, when users interact with the vehicle's voice assistant via voice, the system is usually located in a fixed location within the car, such as integrated with the central control screen corresponding to the driver's seat. When users farther away from the system, such as rear-seat passengers, interact via voice, the central control screen responds to their interaction requests, which can be inconvenient for rear-seat passengers. For example, when a rear-seat passenger activates the voice assistant, the assistant's image is displayed on the central control screen, not only depriving rear-seat passengers of an interactive experience but also potentially interfering with the driving of front-seat passengers. Summary of the Invention

[0005] To address, or at least partially address, the aforementioned technical problems, this disclosure provides a multi-screen interaction method, apparatus, device, and computer-readable storage medium, enabling users to interact with in-vehicle devices corresponding to their location, thereby improving the user's interactive experience and reducing interference with users in other audio zones.

[0006] In a first aspect, embodiments of this disclosure provide a multi-screen interaction method, including:

[0007] If the vocal register of the user generating the audio signal inside the vehicle is inconsistent with the seating register of the user, then the seating register is determined as the target vocal register of the user.

[0008] Control the in-vehicle device corresponding to the target audio region to interact with the user.

[0009] Secondly, embodiments of this disclosure provide a multi-screen interaction device, including:

[0010] The determination module is used to determine the target sound zone of the user if the vocal sound zone corresponding to the user who generates the audio signal in the vehicle is inconsistent with the seat sound zone where the user is located.

[0011] The control module is used to control the interaction between the in-vehicle device corresponding to the target sound zone and the user.

[0012] Thirdly, embodiments of this disclosure provide an in-vehicle device, including:

[0013] Memory;

[0014] Processor; and

[0015] Computer programs;

[0016] The computer program is stored in the memory and configured to be executed by the processor to implement the method as described in the first aspect.

[0017] Fourthly, embodiments of this disclosure provide a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the method described in the first aspect.

[0018] Fifthly, embodiments of this disclosure also provide a computer program product, which includes a computer program or instructions that, when executed by a processor, implement the multi-screen interaction method as described above.

[0019] The multi-screen interaction method, apparatus, device, and computer-readable storage medium provided in this disclosure, when the user's vocal register and the user's seated vocal register are inconsistent, determines the seated vocal register as the user's target vocal register and controls the in-vehicle device corresponding to the target vocal register to interact with the user. When determining the target vocal register, the method distinguishes between the user's vocal register and the seated vocal register, fully considering factors such as the user's brain-buttock separation sitting posture in the vocal register dimension, thereby improving the accuracy of the target vocal register judgment. This allows the user to interact with the in-vehicle device corresponding to the user's vocal register, improving the user's interactive experience and reducing interference to users in other vocal registers. Attached Figure Description

[0020] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0021] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 A flowchart of a multi-screen interaction method provided in this disclosure embodiment;

[0023] Figure 2 This is a vehicle hardware architecture diagram corresponding to a multi-screen interaction method provided in an embodiment of this disclosure;

[0024] Figure 3 A flowchart of another multi-screen interaction method provided in this disclosure embodiment;

[0025] Figure 4 A vehicle hardware architecture diagram corresponding to another multi-screen interaction method provided in this embodiment of the disclosure;

[0026] Figure 5 A schematic diagram illustrating the implementation principle of another multi-screen interaction method provided in this embodiment of the disclosure;

[0027] Figure 6 A flowchart of another multi-screen interaction method provided in this disclosure embodiment;

[0028] Figure 7 This is a schematic diagram of the structure of the multi-screen interaction device provided in the embodiments of this disclosure;

[0029] Figure 8 This is a schematic diagram of the structure of the vehicle-mounted device provided in an embodiment of this disclosure. Detailed Implementation

[0030] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0031] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0032] Typically, users interact with the in-car voice interaction system via voice commands, but the system's responses are displayed on the central control screen. For example, when a rear-seat user wakes up the voice assistant, the assistant's image is displayed on the central control screen. This not only deprives rear-seat users of an interactive experience but also interferes with the driving experience of front-seat users.

[0033] To address this issue, this disclosure provides a multi-screen interaction method, which will be described below with reference to specific embodiments.

[0034] Figure 1 A flowchart of a multi-screen interaction method provided in this embodiment of the disclosure; Figure 2This is a vehicle hardware architecture diagram provided for embodiments of this disclosure. Before introducing the multi-screen interaction method, let's first discuss... Figure 2 The vehicle hardware architecture diagram shown is presented below.

[0035] The overall hardware architecture of this vehicle is as follows: the car is divided into different audio zones, and varying numbers of screens, audio acquisition devices, and audio playback devices are installed in different locations within the vehicle as needed. A voice assistant is also installed inside the vehicle. The audio acquisition devices can be microphones; the audio playback devices can be speakers, horns, etc. Figure 2 As shown, the hardware architecture diagram includes a complete vehicle 20, a vehicle infotainment system 21, audio zones 22, 23, 24, and 25, microphones 26 and 27 corresponding to audio zone 22, microphones 28 and 29 corresponding to audio zone 23, microphones 210 and 211 corresponding to audio zone 24, microphones 212 and 213 corresponding to audio zone 25, a central control screen 214 corresponding to the driver's seat, a screen 215 corresponding to the passenger's seat, and a screen 216 corresponding to the rear seats. All the microphones, speakers, and screens are connected to the vehicle infotainment system 21, and a voice assistant (not shown in the diagram) is integrated into the vehicle infotainment system.

[0036] The following combination Figure 2 The vehicle hardware architecture diagram shown is for Figure 1 The multi-screen interaction method shown here is introduced. The method mainly includes the following steps:

[0037] S101, Acquire audio signal.

[0038] Specifically, in this embodiment, different areas inside the vehicle will be divided into different sound zones, for example... Figure 2 The diagram shows registers 22, 23, 24, and 25; microphones and speakers will be installed in each register, for example... Figure 2 The diagram shows the microphones and speakers corresponding to each audio zone. The microphones capture the user's audio signal, and the speakers play the user's audio signal. When the user makes a sound, one or more microphones can capture the user's audio signal, play it through the corresponding speaker, and send the audio signal to the vehicle's infotainment system.

[0039] S102. By identifying the audio signal, determine the target location of the user who emitted the audio signal.

[0040] Specifically, audio signal recognition involves converting an audio signal into corresponding text information. Optionally, the process of converting an audio signal into corresponding text information mainly includes feature extraction, acoustic modeling, language modeling, and dictionary and decoding processes.

[0041] Specifically, before feature extraction, in order to more effectively highlight features, it is often necessary to perform preprocessing work such as filtering and framing on the acquired audio signal to extract the acquired audio signal from the original signal, and then perform feature extraction to convert the sound signal from the time domain to the frequency domain, providing suitable feature vectors for the acoustic model; in the acoustic model, the score of each feature vector on the acoustic feature is calculated based on the acoustic features; then, the language model calculates the probability of the possible word sequence corresponding to the audio signal according to relevant linguistic theories; finally, the word sequence is decoded according to the existing dictionary to obtain the final possible text representation.

[0042] In this embodiment, when the text content included in the audio signal contains wake-up words such as "Hello, please turn on" or "hello, hello," the location of the user who emitted the audio signal is determined as the target location. Here, the wake-up word can be understood as the word used to activate the voice interaction system of a smart device; specifically, in this embodiment, the wake-up word is the word used to activate the in-vehicle voice assistant.

[0043] For example, the target location can be determined based on the signal strength and orientation of the target audio signal collected by each microphone. For instance... Figure 2 In the architecture diagram shown, when the user sitting in the front passenger seat speaks, multiple microphones may be able to collect the audio signal. Therefore, the multiple microphones will transmit the audio signal they have collected to the vehicle system for processing. The vehicle system determines that the audio signal collected by microphone 28 has the greatest intensity, and at the same time, the direction of the audio signal is from the right half of the vehicle. Therefore, it determines that the user's target position is the front passenger seat.

[0044] S103. Control the vehicle-mounted equipment corresponding to the target location to interact with the user via voice.

[0045] Optionally, in this embodiment, controlling the in-vehicle device corresponding to the target location to interact with the user via voice includes: controlling the in-vehicle device corresponding to the target location to display a virtual image, the virtual image being used to interact with the user via voice.

[0046] Specifically, in this embodiment, only one voice assistant is set up in the entire vehicle. After determining the target location of the user who wakes up the voice assistant, the virtual image of the voice assistant is displayed on the screen corresponding to the target location. For example, when the target location of the user who wakes up the voice assistant is in the back seat of the vehicle, the virtual image is displayed on the rear screen 216 corresponding to the rear seat; when the target location of the user who wakes up the voice assistant is in the driver's seat, the virtual image is displayed on the central control screen 214 corresponding to the driver's seat; and when the target location of the user who wakes up the voice assistant is in the passenger seat, the virtual image is displayed on the passenger screen 215 corresponding to the passenger seat.

[0047] Optionally, in this embodiment, after the screen displaying the virtual avatar is determined, the audio playback device corresponding to the in-vehicle equipment is controlled to play the voice broadcast by the virtual avatar.

[0048] Specifically, in this embodiment, after determining the display screen for the virtual avatar, the audio playback device corresponding to the virtual avatar display screen is controlled to play the voice announcement of the virtual avatar. For example, when the virtual avatar is displayed on the rear screen 216, the audio playback device 212 corresponding to the rear screen (e.g., the user emitting the audio signal is located on the left side of the rear seat) is controlled to play the voice announcement of the virtual avatar; when the virtual avatar is displayed on the central control screen 214, the audio playback device 27 corresponding to the central control screen is controlled to play the voice announcement of the virtual avatar; and when the virtual avatar is displayed on the passenger-side screen 215, the audio playback device 29 corresponding to the passenger-side screen is controlled to play the voice announcement of the virtual avatar.

[0049] It is understandable that in this embodiment, since there is only one voice assistant, when a user in another location outside the target location wakes up the voice assistant, that is, when a new target location appears, the virtual image of the voice assistant will be displayed on the screen corresponding to the new target location, and the voice broadcast by the virtual image will be played on the audio playback device corresponding to the screen corresponding to the new target location.

[0050] This embodiment of the disclosure identifies the target location of the user who sent the audio signal by recognizing the acquired audio signal, and controls the in-vehicle device corresponding to the target location to conduct voice interaction with the user. This allows the user to conduct voice interaction with the in-vehicle device corresponding to the user's location, improving the user's interactive experience and reducing interference to users in other locations.

[0051] Figure 3 This is another multi-screen control method provided in the embodiments of this disclosure. Figure 4 This is another vehicle hardware architecture diagram provided in this disclosure. Figure 5 This is a schematic diagram illustrating the implementation principle of another multi-screen control method provided in this embodiment. Before introducing the multi-screen control method provided in this embodiment, let's first discuss... Figure 4 The vehicle hardware architecture diagram shown is presented below.

[0052] exist Figure 2 Based on the hardware architecture diagram shown, Figure 4 The vehicle hardware architecture diagram shown also includes: cameras and car seat sensors. The car seat sensor is a thin-film contact sensor with contacts evenly distributed on the pressure-bearing surface of the seat. When the seat is subjected to external pressure, it generates a trigger signal. Figure 4 As shown, in Figure 2Based on this, the vehicle hardware architecture diagram also includes a camera 41 corresponding to the driver's seat, a camera 42 corresponding to the passenger seat, a camera 43 corresponding to the rear seats, a seat pressure sensor 44 corresponding to the driver's seat, a seat back angle sensor 45 corresponding to the driver's seat, a seat fore-and-aft position sensor 46 corresponding to the driver's seat, a seat pressure sensor 47 corresponding to the passenger seat, a seat back angle sensor 48 corresponding to the passenger seat, a seat fore-and-aft position sensor 49 corresponding to the passenger seat, a seat pressure sensor 410 corresponding to the left rear seat, a seat back angle sensor 411 corresponding to the left rear seat, a seat fore-and-aft position sensor 412 corresponding to the left rear seat, a seat pressure sensor 413 corresponding to the right rear seat, a seat back angle sensor 414 corresponding to the right rear seat, and a seat fore-and-aft position sensor 415 corresponding to the right rear seat.

[0053] The following is combined with Figure 4 and Figure 5 right Figure 3 Another multi-screen control method will be introduced, which includes the following steps:

[0054] S301, Acquire audio signal.

[0055] In this embodiment, this step is the same as step S101, and will not be described again here.

[0056] S302. By recognizing the audio signal, determine the pronunciation position of the audio signal.

[0057] Specifically, in this embodiment, when the text content included in the audio signal includes wake-up words such as "Hello, please turn on" or "hello, hello", the location of the user who emitted the audio signal is determined as the pronunciation location.

[0058] For example, the location of the sound can be determined by analyzing the signal strength and orientation of the target audio signal collected by each microphone. Figure 5 The pronunciation location is shown.

[0059] S303. Based on the posture information of the seat corresponding to the sound source and the sound source position, determine the target location of the user who emitted the audio signal.

[0060] Specifically, after determining the user's pronunciation location, the target location of the user is determined based on the seat's posture information corresponding to that pronunciation location. This seat posture information can specifically include seat backrest information and seat position information.

[0061] Optionally, the seat posture information corresponding to the sound output position can be determined based on the seat back angle sensor and the seat fore-and-aft position sensor corresponding to the sound output position, i.e., using... Figure 5 The seat back sensor and seat fore-and-aft position sensor shown are determined.

[0062] Optionally, based on the posture information of the seat corresponding to the sound-emitting position and the sound-emitting position, the target position of the user emitting the audio signal can be determined. Specifically, based on the posture information of the seat corresponding to the sound-emitting position and the sound-emitting position, the position of the user's buttocks and head can be determined.

[0063] Specifically, the seat back angle sensor is used to determine the tilt angle of the seat back and the pressure at various positions on the seat back. Based on the tilt angle of the seat back and the pressure at various positions on the seat back, the position of the user's buttocks and head can be determined. The seat fore-and-aft position sensor is used to determine the fore-and-aft position of the seat. Based on the fore-and-aft position of the seat, combined with the seat back angle sensor, the position of the user's buttocks and head can be further determined.

[0064] For example, when the sound source is in the passenger seat and the passenger has not adjusted the seat back angle or the fore-and-aft position, it can be based on... Figure 4 The seat back angle sensor 48 and seat fore-and-aft position sensor 49 indicate that the user's current sitting posture is approximately upright. When the passenger in the front seat adjusts the seat back angle to approximately the same horizontal plane as the seat and simultaneously adjusts the front passenger seat to approximately move closer to the rear seats, it can be determined that... Figure 4 The seat back angle sensor 48 and seat fore-and-aft position sensor 49 indicate that the user's current sitting posture is approximately lying flat. In this embodiment, the user's approximately lying flat sitting posture is referred to as brain-hip separation, which can be specifically understood as the user's hips and head not being in the same vocal range.

[0065] In this embodiment, after determining the position of the user's buttocks and head based on the posture information of the seat corresponding to the sound position and the sound position, if the position of the user's buttocks and the position of the head are the same, then the position of the user's buttocks or the position of the head is taken as the target position of the user who emitted the audio signal.

[0066] Specifically, when the vocal register corresponding to the user's hip position and the vocal register corresponding to the user's head position are in the same vocal register—for example, as described above when the user is sitting upright in the front passenger seat—the position of the user's hips and the position of their head are in the same vocal register. Figure 4The sound range 23 is shown. Therefore, the position of the user's hips or head is taken as the target position of the user emitting the audio signal.

[0067] If the user's hips are not in the same position as their head, then the position of the user's hips will be taken as the target position of the user emitting the audio signal.

[0068] Specifically, when the vocal register corresponding to the user's hip position is not in the same vocal register as the vocal register corresponding to the user's head position, such as the user described above in the front passenger seat with a head-hip separation posture, the user's hips and head are not in the same vocal register, meaning the user's hip position is in... Figure 4 The user's head is positioned in the indicated audio range 23. Figure 4 The sound range 25 is shown. When the user is in a seated position with the head and buttocks separated, the sound range where the user's buttocks are located is taken as the target sound range, and the position of the buttocks is taken as the target position of the user emitting the audio signal.

[0069] Optionally, in other embodiments, it may be possible to first determine whether there is a person on the seat corresponding to the sound source based on at least one of the visual sensor signal and the pressure sensor corresponding to the sound source; if there is a person on the seat corresponding to the sound source, then determine the target location of the user who emitted the audio signal based on the posture information of the seat corresponding to the sound source and the sound source location.

[0070] Specifically, the visual sensor signal corresponding to the sound production location can be image information captured by a camera at that sound production location, for example... Figure 5 The visual positioning shown; the pressure sensor corresponding to the sound-producing position can specifically be the pressure value of the seat at the sound-producing position collected by the pressure sensor corresponding to that sound-producing position, for example... Figure 5 The seat pressure sensor shown determines whether someone is in the seat based on at least one of the image information captured by a camera at the sound source location and the pressure value of the seat at the sound source location.

[0071] Once it is determined that there is someone in the seat corresponding to the sound source, the position of the user's buttocks and head is determined based on the posture information of the seat corresponding to the sound source and the sound source position. If the position of the user's buttocks and head are not the same, the position of the user's buttocks is taken as the target position of the user who emitted the audio signal. If the position of the user's buttocks and head are the same, the position of the user's buttocks or head is taken as the target position of the user who emitted the audio signal.

[0072] Optionally, in some embodiments, the aforementioned sound localization, visual localization, seat pressure sensor, seat back angle sensor, and seat fore-and-aft position sensor can be integrated into a single unit. Figure 5 The spatial positioning model shown is used to determine the position of the user's hips and head, and then to determine the target location.

[0073] Furthermore, if the position of the user's buttocks and the position of the head are not the same, then the audio signal collected by the audio acquisition device corresponding to the position of the user's buttocks and the audio signal collected by the audio acquisition device corresponding to the position of the user's head are received.

[0074] Specifically, when a user's hips and head are in different audio registers, all audio signals collected by the hip and head audio acquisition devices are sent to the vehicle's infotainment system. For example, as described above, when the user is in a head-hip separation sitting posture in the passenger seat, the system will... Figure 4 The audio signals collected by microphone 28 and microphone 212 are sent to the vehicle's infotainment system. The system receives these signals and sends them to the voice assistant for further interaction. Audio signals collected by microphones at other locations are sent to the infotainment system, but the system does not receive these signals for further interaction.

[0075] It is understandable that when a user's hips and head are in the same sound zone, the audio acquisition devices corresponding to the hips and head are the same. Therefore, the signal collected by the audio acquisition device is sent to the vehicle's infotainment system, which receives the audio signal and sends it to the voice assistant for subsequent interaction. Audio signals collected by microphones at other locations are sent to the vehicle's infotainment system, but the system will not receive the audio signal for subsequent interaction.

[0076] Optionally, in some embodiments, when a user at the target location interacts with the virtual avatar of the voice assistant, and the audio signal emitted by a user at another location contains a wake-up word, the microphone at the other location collects the audio signal and sends it to the vehicle system. After receiving the audio signal, the vehicle system executes steps S302-S303 to determine a new target location.

[0077] S304. Control the on-board equipment corresponding to the target location to interact with the user via voice.

[0078] In this embodiment, this step is the same as step S103, and will not be described again here.

[0079] In this embodiment, when determining the target location, the location of the user's buttocks and head is distinguished, fully considering the user's sitting posture with the head and buttocks separated, thus improving the accuracy of target location judgment. Furthermore, when the user is in a sitting posture with the head and buttocks separated, audio signals collected by audio acquisition devices corresponding to the user's buttocks and head locations are received. This ensures that the user can interact with the virtual avatar of the voice assistant regardless of whether they are adjusting from a near-lying sitting posture to an upright sitting posture or vice versa, improving the user's interactive experience.

[0080] Figure 6 Another multi-screen interaction method provided in this disclosure includes the following steps:

[0081] S601, Acquire audio signal.

[0082] In this embodiment, the method is the same as step S101, and will not be described again here.

[0083] S602. By identifying the audio signal, determine the target location of the user who emitted the audio signal.

[0084] In this embodiment, this step is the same as step S303, and will not be described again here.

[0085] S603. If the in-vehicle device corresponding to the target location is in the open state, control the in-vehicle device corresponding to the target location to interact with the user via voice.

[0086] Specifically, in this embodiment, when the vehicle-mounted device corresponding to the target location interacts with the user via voice, it is necessary to determine whether the screen corresponding to the target location is on.

[0087] For example, if the target location is a location other than the driver's seat and the passenger seat, and the in-vehicle equipment corresponding to the other location is in a closed and / or folded state, then the central control device will interact with the user via voice.

[0088] Specifically, in some vehicles, the screens corresponding to the driver's seat and the passenger's seat are automatically turned on when the vehicle is in use, and these screens cannot be physically closed while the vehicle is in use; however, the screens for the rear seats can be physically closed or folded. Therefore, when the user sending the audio signal is located in the rear seats (e.g., the left rear seat) and the screen for the rear seats (e.g., the left rear seat) is closed or / or folded, preventing the virtual avatar of the voice assistant from being displayed on the rear seat screen, the virtual avatar will be displayed on the central control screen corresponding to the driver's seat, and the audio playback device will be the same as the audio playback device for the central control screen corresponding to the driver's seat.

[0089] Because the target location of the user emitting the audio signal is in the rear seat (e.g., the left side of the rear seat), the vehicle system receives the audio signal collected by the audio acquisition device corresponding to the rear seat location (e.g., the left side of the rear seat).

[0090] Optionally, in other vehicles, when the target location of the user emitting the audio signal is a rear seat position (e.g., the left rear seat) and the screen corresponding to the rear seat position (e.g., the left rear seat) is in a closed or / or folded state, the virtual image can also be displayed in the front passenger seat position. Therefore, the audio playback device is the audio playback device corresponding to the screen corresponding to the front passenger seat position. Because the target location of the user emitting the audio signal is a rear seat position (e.g., the left rear seat), the vehicle system receives the audio signal collected by the audio acquisition device corresponding to the rear seat position (e.g., the left rear seat).

[0091] Optionally, in other vehicles, when the target location of the user emitting the audio signal is a rear seat (e.g., the left rear seat) and the screen corresponding to that rear seat (e.g., the left rear seat) is closed or / or folded, a virtual avatar can be displayed on a screen closer to the target location (e.g., the left rear seat). For example, in a vehicle with three rows of seats, there may be screens corresponding to the driver's seat, the front passenger seat, the second row, and the third row. When a third-row user (e.g., the left third row) emits an audio signal, and the screen corresponding to the third row (e.g., the left third row) is closed and / or folded, the virtual avatar is displayed on the screen corresponding to the second row; therefore, the audio playback device is the audio playback device corresponding to the screen of the second row. Because the target location of the user emitting the audio signal is in the third row (e.g., the left third row), the vehicle's infotainment system receives the audio signal collected by the audio acquisition device corresponding to the third row (e.g., the left third row).

[0092] It should be noted that when the screen corresponding to the target location is in a closed state and / or a folded state, the processing logic of the virtual image, audio playback device, and audio acquisition device provided in this disclosure is only a part of the feasible technical methods. This disclosure does not limit the specific processing logic of the virtual image, audio playback device, and audio acquisition device. Other variations of the above feasible technical solutions are all within the protection scope of this embodiment.

[0093] Understandably, while the central control screens corresponding to the driver's and passenger's seats cannot be physically turned off, they are in a sleep state. Therefore, when a virtual avatar needs to be displayed on these screens, but they are in a sleep state, the screens will be turned on. Similarly, screens in other locations are not turned off and / or folded, but are in a sleep state; when a virtual avatar needs to be displayed on other screens, those screens will also be turned on.

[0094] This embodiment of the disclosure determines the display position of the virtual image based on the screen state, so that when the screen corresponding to the target position is in a closed and / or folded state, the virtual image can be displayed on other screens, ensuring the requirement that there is a response when the image is woken up, and further improving the user's interactive experience.

[0095] In some embodiments, the multi-screen interaction method includes: if the vocal register corresponding to the user generating the audio signal in the vehicle is inconsistent with the seating register of the user, then the seating register is determined as the user's target register; and the in-vehicle device corresponding to the target register is controlled to interact with the user.

[0096] In this embodiment, the user can be a user who initiates interaction with the vehicle. The pronunciation zone can be the sound zone where the pronunciation is located. Pronunciation location can be understood as obtaining specific location coordinates through sound source location. The sound source can be the sound source that triggers the generation of the audio signal. This embodiment does not limit the method of pronunciation location; for example, the pronunciation location can be the position of the user's head or the position of the user's mouth. Optionally, the pronunciation zone can be the sound zone where the user's head is located. The sound zone can be used for location positioning within the vehicle with the sound zone as the smallest spatial granularity. This embodiment does not limit the method of determining the pronunciation zone. The seating sound zone can be the sound zone where the user sits within the vehicle. Optionally, the seating sound zone can be the sound zone where the user's buttocks are located. This embodiment does not limit the method of determining the seating sound zone. The target sound zone can be the sound zone where the vehicle believes the user is located when interacting with the user. The interaction can be human-computer interaction between the in-vehicle device and the user. This embodiment does not limit the type of interaction; for example, the interaction can include one or more of voice interaction, text interaction, and virtual avatar action interaction.

[0097] In this embodiment of the disclosure, different areas inside the vehicle are divided into different sound zones, for example... Figure 2 The diagram shows registers 22, 23, 24, and 25; microphones and speakers will be installed in each register, for example... Figure 2 The diagram shows the microphones and speakers corresponding to each sound zone. The microphones can collect user audio signals; when a user makes a sound, one or more microphones can capture the user's audio signal. If the multi-screen interaction device receives an audio signal containing a wake-up word, it indicates that the user emitting the audio signal wants to interact with the vehicle. The multi-screen interaction device can determine whether the user's vocal sound zone matches the user's seating sound zone within the vehicle. If they do not match, the seating sound zone is used as the target sound zone for subsequent interaction. Furthermore, the multi-screen interaction device can control the in-vehicle equipment corresponding to the target sound zone to interact with the user.

[0098] The multi-screen interaction method provided in this disclosure determines the user's target audio zone when the user's vocal range and the user's seated audio zone are inconsistent. It then controls the in-vehicle device corresponding to the target audio zone to interact with the user. In determining the target audio zone, the method distinguishes between the user's vocal range and seated audio zone, fully considering factors such as the user's posture (brain-buttock separation) in the audio zone dimension. This improves the accuracy of target audio zone judgment, allowing the user to interact with the in-vehicle device corresponding to their audio zone, enhancing the user's interactive experience, and reducing interference with users in other audio zones.

[0099] In this embodiment of the disclosure, if the pronunciation area and the seating area are determined, it can be determined whether the pronunciation area and the seating area are consistent by comparing the pronunciation area and the seating area; or if the pronunciation area and the seating area are not determined, it can be determined whether the pronunciation area and the seating area are consistent based on the posture information of the seat.

[0100] In some embodiments of this disclosure, the vocal register corresponding to the user generating the audio signal inside the vehicle is inconsistent with the seating register of the user, including: determining the vocal register and the seating register; if the vocal register and the seating register are different registers inside the vehicle, then determining that the vocal register and the seating register are inconsistent.

[0101] In this embodiment, when the text content included in the audio signal contains wake-up words such as "Hello, please turn on" or "hello, hello," the multi-screen interaction device can determine the sound region where the sound source emitting the audio signal is located within the vehicle as the pronunciation sound region. Furthermore, the sound region where the user is seated within the vehicle is determined as the seating sound region. Optionally, the user can be determined through pronunciation location. Further, the pronunciation sound region and the seating sound region are compared; if the pronunciation sound region and the seating sound region are different sound regions within the vehicle, then it is determined that the pronunciation sound region and the seating sound region are inconsistent.

[0102] In some embodiments of this disclosure, the vocal range is determined by identifying an audio signal, or by identifying image information collected by a visual sensor inside the vehicle.

[0103] Among these methods, determining the pronunciation region through audio signal recognition can involve determining the signal strength and orientation of multiple audio signals collected from multiple microphones, for example... Figure 5 The pronunciation localization is shown. Specifically, determining the pronunciation region can involve head recognition, which can be either target recognition of the head within an image or target recognition of related parts of the head within an image. These related parts include, but are not limited to, parts of the head such as the mouth.

[0104] In this embodiment, the multi-screen interaction device can acquire audio signals and perform text recognition on the audio signals, converting a segment of audio signal into corresponding text information. Optionally, the process of converting audio signals into corresponding text information mainly includes feature extraction, acoustic modeling, language modeling, and dictionary and decoding processes. Specifically, before feature extraction, in order to more effectively highlight features, it is often necessary to perform preprocessing work such as filtering and framing on the acquired audio signals to extract the acquired audio signals from the original signals, and then perform feature extraction to convert the sound signals from the time domain to the frequency domain, providing suitable feature vectors for the acoustic model; in the acoustic model, the score of each feature vector on the acoustic features is calculated based on the acoustic features; then, the language model calculates the probability of possible word sequences corresponding to the audio signal based on relevant linguistic theories; finally, based on the existing dictionary, the word sequence is decoded to obtain the final possible text representation. In this embodiment, when the text content included in the audio signal includes wake words such as "Hello, please turn on" or "hello, hello", the user who issued the above audio signal is identified as the user interacting with the vehicle. The wake-up word can be understood as a word used to wake up the interactive system of a smart device. In this embodiment, the wake-up word is the word used to wake up the in-vehicle voice assistant.

[0105] In this embodiment, the vocal range can be determined based on the signal strength and orientation of the audio signals collected by each microphone. For example, Figure 2 In the architecture diagram shown, when the user sitting in the front passenger seat speaks, multiple microphones in the vehicle can collect the audio signal. Therefore, the multiple microphones transmit the audio signals they collect to the vehicle system for processing. The vehicle system determines that the audio signal collected by microphone 28 has the greatest intensity, and at the same time, the direction of the audio signal originates from the right half of the vehicle. Therefore, it determines that the user's corresponding vocal range is the front passenger seat.

[0106] In this embodiment, with user authorization, the vehicle's visual sensors can collect the user's image information. The multi-screen interaction device can perform head target recognition on the image information of the user currently interacting with the vehicle, and determine the vocal region where the recognition result is located as the vocal region. This head target recognition includes, but is not limited to, target recognition of the entire head and / or target recognition of other parts of the head such as the mouth.

[0107] In some embodiments of this disclosure, the seating sound zone is determined by a seat pressure sensor, or the seating sound zone is determined by identifying image information collected by a vision sensor inside the vehicle.

[0108] Among them, Figure 2 Based on the hardware architecture diagram shown, Figure 4The vehicle hardware architecture diagram shown also includes: cameras and car seat sensors. Figure 4 As shown, in Figure 2 Based on this, the vehicle hardware architecture diagram also includes a camera 41 corresponding to the driver's seat, a camera 42 corresponding to the passenger's seat, a camera 43 corresponding to the rear seats, a seat pressure sensor 44 corresponding to the driver's seat, a seat back angle sensor 45 corresponding to the driver's seat, a seat fore-and-aft position sensor 46 corresponding to the driver's seat, a seat pressure sensor 47 corresponding to the passenger's seat, a seat back angle sensor 48 corresponding to the passenger's seat, a seat fore-and-aft position sensor 49 corresponding to the passenger's seat, a seat pressure sensor 410 corresponding to the left rear seat, a seat back angle sensor 411 corresponding to the left rear seat, a seat fore-and-aft position sensor 412 corresponding to the left rear seat, a seat pressure sensor 413 corresponding to the right rear seat, a seat back angle sensor 414 corresponding to the right rear seat, and a seat fore-and-aft position sensor 415 corresponding to the right rear seat. The seat pressure sensor can be a thin-film contact sensor, with the sensor's contacts distributed on the pressure-bearing surface of the seat. When the seat is subjected to external pressure, a trigger signal is generated.

[0109] In this embodiment, the multi-screen interaction device can determine the seat where the user interacting with the vehicle is located based on voice localization, and determine the corresponding pressure value according to the seat pressure sensor of that seat. If the pressure value is greater than the pressure threshold of the seat pressure sensor, it indicates that the user is sitting on that seat, and the sound zone where that seat is located is determined as the sitting sound zone.

[0110] In this embodiment, the identification of the seating sound zone can specifically be hip recognition. With user authorization, the vehicle's visual sensors can collect the user's image information. The multi-screen interaction device can perform hip target recognition on the image information of the user currently interacting with the vehicle, and determine the sound zone where the recognition result is located as the seating sound zone. This hip target recognition includes, but is not limited to, the recognition of the hip as a whole and / or the recognition of related parts of the hip. These related parts include, but are not limited to, the waist, legs, and other parts that have a positional relationship with the hip.

[0111] In some embodiments of this disclosure, the pronunciation zone corresponding to the user generating the audio signal inside the vehicle is inconsistent with the seating zone of the user. This includes determining that the pronunciation zone and the seating zone are inconsistent when the posture information of the user's seat meets the target posture conditions. In some embodiments of this disclosure, the posture information includes one or more of the following: the tilt angle of the seat back, the pressure of the seat back, and the fore-and-aft position of the seat.

[0112] The target posture condition can be a pre-defined representation of the seat posture when the user's vocal range and seat position are different. This target posture condition can correspond to the type of posture information. For example, if the posture information includes the seat back tilt angle, the target posture condition can include a tilt angle greater than a preset angle threshold; if the posture information includes seat back pressure, the target posture condition can include seat back pressure greater than a seat back pressure threshold; if the posture information includes seat fore-and-aft position, the target posture condition can include seat fore-and-aft movement greater than a preset movement distance. If the posture information meets one or more of the following conditions, it is determined that the vocal range and seat position are inconsistent: the tilt angle is greater than a preset angle threshold, the seat back pressure is greater than a seat back pressure threshold, and the seat fore-and-aft movement is greater than a preset movement distance. For example, a multi-screen interactive device can determine whether the user's seat position (where the buttocks are located) and the vocal range (where the head is located) are consistent based on the seat back tilt angle and the pressure at various positions on the seat back.

[0113] In this embodiment, the multi-screen interaction device can determine whether the posture information of the seat where the user is sitting meets the target posture conditions based on the posture information of the seat. If so, it determines that the pronunciation range and the sitting range are inconsistent, and then determines the target range where the user is located based on the determination result. The posture information of the seat can specifically include information related to the seat back and information related to the seat position.

[0114] Optionally, the user's posture information can be determined based on the seat back angle sensor and the seat fore-and-aft position sensor, i.e., using methods such as... Figure 5 The seat back sensor and seat fore-and-aft position sensor shown are determined.

[0115] Optionally, based on the user's posture information in the seat, it can be determined whether the vocal range and the seated range are consistent, and then based on the consistency result, the target range of the user emitting the audio signal can be determined. Specifically, it can be the seated range or vocal range of the user emitting the audio signal determined based on the user's posture information in the seat.

[0116] Specifically, the seat back angle sensor is used to determine the tilt angle of the seat back and the pressure at various points on the seat back. Based on the tilt angle and / or the pressure at various points on the seat back, it can be determined whether the user's sitting vocal range and vocal range are in the same vocal range. The seat fore-and-aft position sensor is used to determine the position of the seat's fore-and-aft movement. Based on the position of the seat's fore-and-aft movement, combined with the seat back angle sensor, it can further determine whether the user's sitting vocal range and vocal range belong to the same vocal range.

[0117] For example, if the user interacting with the vehicle is the passenger in the front seat, when the passenger adjusts the seat back angle to approximately the same horizontal plane as the seat, and simultaneously adjusts the front seat backward to approximately move closer to the rear seats, it can be based on... Figure 4 The seat back angle sensor 48 and seat fore-and-aft position sensor 49 indicate that the user's current sitting posture is approximately lying flat. In this case, the user's vocal range and sitting range are different vocal ranges. In this embodiment, the user's approximately lying flat sitting posture is referred to as brain-hip separation in the vocal range dimension, which can be specifically understood as the user's hips and head not being in the same vocal range.

[0118] If a user's sitting vocal register and vocal register are inconsistent, the user's sitting vocal register can be used as the target vocal register for the user interacting with the vehicle. Specifically, when a user's sitting vocal register and vocal register are not in the same register—for example, as described above, when a user is in a head-buttock separation sitting posture in the passenger seat—the user's sitting vocal register and vocal register are not in the same register, meaning the user's sitting vocal register is in… Figure 4 The indicated register 23 is where the user's vocal register is located. Figure 4 The indicated vocal range is 25. When the user is in a seated posture with the head and buttocks separated in the vocal range dimension, the user's seated vocal range is taken as the target vocal range.

[0119] In some embodiments of this disclosure, if there is no one in the seat within the sound production area, the user's seat is the seat adjacent to the seat within the sound production area.

[0120] The seats adjacent to the seats within the vocal range can include one or more of the following: the seat in front of the user within the vocal range, the seat to the left of the user within the vocal range, and the seat to the right of the user within the vocal range. The seat in front corresponds to the user lying down, the seat to the left corresponds to the user leaning to the right to communicate across the vocal range, and the seat to the right corresponds to the user leaning to the left to communicate across the vocal range.

[0121] In this embodiment, the multi-screen interaction device can determine the user's seating area and vocal area based on the posture information of the seat in front of the vocal area, even when no one is seated in the vocal area where the vocalization is located. It can then determine that the seating area and vocal area are inconsistent. Specifically, with user authorization, the multi-screen interaction device can determine the user's vocalization location using voiceprint localization, facial recognition, or other methods, and identify the vocal area as the vocal area. Further, it uses seat pressure sensors to determine whether a person is seated in the vocal area. If no one is present, the user's actual seating position is likely in an adjacent seat, such as the seat in front of the user. For example, the multi-screen interaction device can identify the seat in front of the user as the user's seat. Furthermore, the positional information of the seat in front is obtained. If the positional information meets the target positional condition, it means that the seat where the user is located is in a flat or nearly flat position. In this state, the user's vocal range is the vocal range where the vocalization is located, and the user's sitting range is the vocal range of the row in front of the vocal range. The vocal range and the sitting range are not the same.

[0122] Optionally, in some embodiments, the aforementioned sound localization, visual localization, seat pressure sensor, seat back angle sensor, and seat fore-and-aft position sensor can be integrated into a single unit. Figure 5 The spatial positioning model shown is used to determine the user's sitting and vocal registers, and then to determine the target vocal register.

[0123] In some embodiments of this disclosure, the multi-screen interaction method further includes: if the vocalization zone corresponding to the user generating the audio signal in the vehicle is inconsistent with the seating zone of the user, then interaction is performed based on the audio signal collected by the audio acquisition device corresponding to the seating zone and the audio signal collected by the audio acquisition device corresponding to the vocalization zone.

[0124] In this embodiment, when the user's seating area and vocalization area are in different areas, all audio signals collected by the audio acquisition devices in both areas are sent to the vehicle's infotainment system. For example, as described above, when the user is in a head-buttock separation sitting posture in the front passenger seat, the system will... Figure 4 The audio signals collected by microphone 28 and microphone 212 are sent to the vehicle system. The vehicle system receives the audio signals and sends them to the voice assistant for subsequent interaction. The audio signals collected by microphones at other locations are not used for subsequent user interaction with the vehicle system.

[0125] It is understandable that when the user's seating area and vocalization area are in the same area, the audio acquisition devices corresponding to the seating area and vocalization area are the same. Therefore, the signal acquired by the audio acquisition device is sent to the vehicle system, which receives the audio signal and sends it to the voice assistant for subsequent interaction. The audio signals acquired by microphones at other locations are sent to the vehicle system, but the vehicle system will not perform subsequent interaction based on the audio signals from those other locations.

[0126] Optionally, in some embodiments, during the interaction between a user in the target audio region and the virtual avatar of the voice assistant, if the audio signal emitted by a user in another audio region contains a wake-up word, the microphones at other locations collect the audio signal of that other audio region and send it to the vehicle system. After receiving the audio signal of that other audio region, the vehicle system determines the new target audio region corresponding to that other audio region.

[0127] In the above solution, when the user is in a sitting posture with brain-buttock separation in the vocal range dimension, the audio signal is collected by the audio acquisition device corresponding to the user's sitting vocal range and vocal range. This ensures that the user can interact with the virtual image of the voice assistant whether adjusting from a near-lying sitting posture to an upright sitting posture or vice versa, thus improving the user's interactive experience.

[0128] In some embodiments of this disclosure, the multi-screen interaction method further includes: if the pronunciation zone of the user generating the audio signal in the vehicle is consistent with the seating zone of the user, then the seating zone or the pronunciation zone is determined as the user's target zone.

[0129] In this embodiment, the multi-screen interaction device can determine whether the pronunciation area and the seating area are consistent by comparing them, provided that the pronunciation area and the seating area are determined. Alternatively, the multi-screen interaction device can determine whether the pronunciation area and the seating area are consistent based on the seat's posture information if the pronunciation area and the seating area are not determined. Specifically, the multi-screen interaction device can determine whether the posture information of the seat in which the user is interacting meets the target posture conditions. If the target posture conditions are not met, the pronunciation area and the seating area are determined to be consistent, and the seating area or the pronunciation area is then determined as the target sound area where the user is located.

[0130] For example, if the sound source is in the passenger seat and the passenger has not adjusted the seat back angle or the seat's fore-and-aft position, it can be determined based on... Figure 4The seat back angle sensor 48 and seat fore-and-aft position sensor 49 indicate that the user's current sitting posture is approximately upright, thus determining that the vocalization range and the seating range are consistent. When the user's seating range and vocalization range are within the same range, such as when the user is sitting upright in the front passenger seat as described above, the seating range where the user's buttocks are located and the vocalization range where the head is located are in the same range, meaning that the user's seating range and vocalization range are both within the same range. Figure 4 The indicated sound region is 23. Therefore, the user's vocal range or seating range is taken as the target sound region where the user emitting the audio signal is located.

[0131] In some embodiments of this disclosure, controlling the in-vehicle device corresponding to the target audio zone to interact with the user includes: controlling the in-vehicle device corresponding to the target audio zone to display a virtual avatar, the virtual avatar being used to interact with the user.

[0132] The in-vehicle device corresponding to the target sound zone can be a pre-associated in-vehicle device. For example, the in-vehicle device corresponding to the rear sound zone can be the rear screen, the in-vehicle device corresponding to the driver's sound zone can be the central control screen, and the in-vehicle device corresponding to the passenger sound zone can be the passenger screen.

[0133] In this embodiment, a voice assistant can be set up throughout the vehicle. After determining the target audio zone of the user who wakes up the voice assistant, a virtual image of the voice assistant is displayed on the screen corresponding to the target audio zone. For example, when the target audio zone of the user who wakes up the voice assistant is the rear audio zone of the vehicle, the virtual image is displayed on the rear screen 216 corresponding to the rear audio zone; when the target audio zone of the user who wakes up the voice assistant is the driver's audio zone, the virtual image is displayed on the central control screen 214 corresponding to the driver's audio zone; and when the target audio zone of the user who wakes up the voice assistant is the passenger's audio zone, the virtual image is displayed on the passenger screen 215 corresponding to the passenger's audio zone. Specifically, in this embodiment, when controlling the in-vehicle device corresponding to the target audio zone to interact with the user, it is necessary to determine whether the screen corresponding to the target audio zone is open. If the screen corresponding to the target audio zone is open, the in-vehicle device corresponding to the target audio zone is controlled to display the virtual image.

[0134] In some embodiments of this disclosure, after determining the screen on which the virtual avatar is displayed, the multi-screen interaction method further includes: controlling the audio playback device corresponding to the in-vehicle equipment to play the voice broadcast by the virtual avatar.

[0135] In this embodiment, after the display screen of the virtual avatar is determined, the audio playback device corresponding to the virtual avatar display screen is controlled to play the voice broadcast by the virtual avatar. For example, such as... Figure 2As shown, when the virtual avatar is displayed on the rear screen 216, the audio playback device 212 corresponding to the rear screen (e.g., the user emitting the audio signal is in the left rear seat) is controlled to play the voice broadcast by the virtual avatar; when the virtual avatar is displayed on the central control screen 214, the audio playback device 27 corresponding to the central control screen is controlled to play the voice broadcast by the virtual avatar; when the virtual avatar is displayed on the passenger side screen 215, the audio playback device 29 corresponding to the passenger side screen is controlled to play the voice broadcast by the virtual avatar.

[0136] It is understandable that in this embodiment, since there is only one voice assistant, when a user in another voice region outside the target voice region wakes up the voice assistant, that is, when a new target voice region appears, the virtual image of the voice assistant will be displayed on the screen corresponding to the new target voice region, and the voice broadcast by the virtual image will be played on the audio playback device corresponding to the screen corresponding to the new target voice region.

[0137] This embodiment of the disclosure identifies the target audio zone of the user sending the audio signal by recognizing the acquired audio signal, and controls the in-vehicle device corresponding to the target audio zone to interact with the user. This allows the user to interact with the in-vehicle device corresponding to the user's audio zone, improving the user's interactive experience and reducing interference to users in other audio zones.

[0138] In some embodiments of this disclosure, controlling the in-vehicle device corresponding to the target audio zone to interact with the user includes: if the target audio zone is another audio zone other than the driver's audio zone and the passenger's audio zone, and the in-vehicle device corresponding to the target audio zone is in a closed state and / or a folded state, then controlling a backup device to interact with the user; wherein, the backup device includes at least one of the central control device, the screen corresponding to the passenger's seat, and the available screen closest to the target audio zone.

[0139] The backup device can be another in-vehicle device selected when the in-vehicle device corresponding to the target audio zone is unable to interact. The available screen can be a screen that is currently on; specifically, the available screen can be a screen that is currently on and lit, or a screen that is currently on and off.

[0140] For example, if the target audio zone is a zone other than the driver's and passenger's audio zones, and the in-vehicle devices corresponding to these other zones are in a closed and / or folded state, the multi-screen interaction device can control the central control device to interact with the user. Specifically, in some vehicles, the screens corresponding to the driver's and passenger's audio zones are automatically turned on when the vehicle is in use, and these screens cannot be physically closed while the vehicle is in use; however, the screens corresponding to the rear audio zones can be physically closed or folded. Therefore, when the target audio zone of the user emitting the audio signal is the rear audio zone (e.g., the left rear seat) and the screen corresponding to the rear audio zone (e.g., the left rear seat) is in a closed and / or folded state, causing the virtual image of the voice assistant to be unable to be displayed on the screen corresponding to the rear audio zone (e.g., the left rear seat), the multi-screen interaction device can control the display of the virtual image on the central control screen corresponding to the driver's audio zone, while the audio playback device is the audio playback device corresponding to the central control screen corresponding to the driver's audio zone.

[0141] Because the target sound zone of the user emitting the audio signal is in the rear sound zone (e.g., the left rear seat), the vehicle system receives the audio signal collected by the audio acquisition device corresponding to the rear sound zone (e.g., the left rear seat).

[0142] Optionally, in some embodiments, when the target audio zone of the user emitting the audio signal is the rear audio zone (e.g., the left rear seat) and the screen corresponding to the rear audio zone (e.g., the left rear seat) is in a closed state or / or folded state, the virtual image can also be displayed in the passenger-side audio zone. Therefore, the audio playback device is the audio playback device corresponding to the screen corresponding to the passenger-side audio zone. Because the target audio zone of the user emitting the audio signal is in the rear audio zone (e.g., the left rear seat), the vehicle system receives the audio signal collected by the audio acquisition device corresponding to the rear audio zone (e.g., the left rear seat).

[0143] Optionally, in some embodiments, when the target location of the user emitting the audio signal is the rear audio zone (e.g., the left rear seat) and the screen corresponding to the rear audio zone (e.g., the left rear seat) is in a closed or / or folded state, a virtual image can also be displayed on the available screen closest to the target audio zone (e.g., the left rear seat). For example, when the vehicle has three rows of seats, there may be screens corresponding to the driver's seat, the front passenger seat, the second row, and the third row. When a third-row user (e.g., the left side of the third row) emits an audio signal, and the screen corresponding to the third-row audio zone (e.g., the left side of the third row) is in a closed and / or folded state, a virtual image is displayed on the screen corresponding to the second-row audio zone. Therefore, the audio playback device is the audio playback device corresponding to the screen of the second-row audio zone. Because the target audio zone of the user emitting the audio signal is in the third-row audio zone (e.g., the left side of the third row), the vehicle system receives the audio signal collected by the audio acquisition device corresponding to the third-row audio zone (e.g., the left side of the third row).

[0144] It should be noted that when the screen corresponding to the target audio zone is in a closed state and / or a folded state, the processing logic of the virtual image, audio playback device, and audio acquisition device provided in this disclosure is only a part of the feasible technical methods. This disclosure does not limit the specific processing logic of the virtual image, audio playback device, and audio acquisition device. Other variations of the above feasible technical solutions are all within the protection scope of this embodiment.

[0145] Understandably, while the central control screens corresponding to the driver's and passenger's audio zones cannot be physically turned off, they are in a sleep-mode state. Therefore, when a virtual avatar needs to be displayed on the screens corresponding to the driver's and passenger's audio zones, but these screens are in a sleep-mode state, they will be illuminated. Similarly, the screens in other audio zones are not turned off and / or folded, but are in a sleep-mode state; when a virtual avatar needs to be displayed on the screens in other audio zones, those screens will also be illuminated.

[0146] This embodiment of the disclosure determines the in-vehicle device for displaying the virtual image based on the screen state, so that when the screen corresponding to the target audio zone is in a closed and / or folded state, the virtual image can be displayed on the screen corresponding to other audio zones, ensuring that there is a response when the device is woken up, and further improving the user's interactive experience.

[0147] In some embodiments, the multi-screen interaction method includes: if the user's sitting posture in the vehicle is a brain-buttock separation sitting posture, then the sound zone where the user's buttocks are located is taken as the target sound zone; and controlling the in-vehicle device corresponding to the target sound zone to interact with the user.

[0148] Among them, the user's near-lying sitting posture is called the brain-hip separation sitting posture, which can be specifically understood as the user's buttocks and head not being in the same vocal range.

[0149] Specifically, when the vocal register of a user's buttocks is not in the same vocal register as the vocal register of their head, such as when a user is sitting in the front passenger seat in a head-buttock separation posture as described above, the user's buttocks and head are not in the same vocal register, meaning the vocal register of the user's buttocks is in a different vocal register. Figure 4 The indicated sound register 23 is the sound register where the user's head is located. Figure 4 The indicated sound zone is 25. When the user is in a head-buttock separation sitting posture, the sound zone where the user's buttocks are located is taken as the target sound zone. Furthermore, the system interacts with the user on the in-vehicle device corresponding to this target sound zone.

[0150] In some embodiments, the user's sitting posture inside the vehicle is a brain-buttock separation sitting posture, including: the vocal range where the user's buttocks are located is inconsistent with the vocal range where the user's head is located.

[0151] The vocal range where the user's buttocks are located can be understood as the sitting vocal range, and the vocal range where the user's head is located can be understood as the vocal range.

[0152] In this embodiment, the multi-screen interaction device can use the sound region where the sound source emitting the above-mentioned audio signal is located as the sound region where the user's head is located, and determine the sound region where the user's buttocks are located in the vehicle as the seating sound region. Further, the sound region comparison between the sound region and the seating sound region is performed. If the sound region and the seating sound region are different sound regions in the vehicle, it is determined that the sound region and the seating sound region are inconsistent.

[0153] In some embodiments of this disclosure, the sound zone where the user's head is located is determined by identifying the audio signal or by identifying the image information collected by the visual sensor inside the vehicle.

[0154] Among these methods, determining the acoustic region where the user's head is located through audio signal recognition can involve determining the signal strength and orientation of multiple audio signals collected from multiple microphones, for example... Figure 5The pronunciation localization is shown. Determining the vocal region where the user's head is located can specifically involve head recognition. This head recognition can be target identification of the head within an image, or target identification of related parts of the head within an image. These related parts include, but are not limited to, the mouth and other parts of the head. In this embodiment, with user authorization, the vehicle's visual sensors can collect the user's image information, and the multi-screen interaction device can perform head target recognition on the image information of the user currently interacting with the vehicle, determining the vocal region where the recognition result is located as the vocal region where the head is located.

[0155] In some embodiments of this disclosure, the sound zone where the user's buttocks are located is determined by a seat pressure sensor, or the sound zone where the user sits is determined by identifying image information collected by a visual sensor inside the vehicle.

[0156] In this embodiment, the multi-screen interaction device can determine the seat where the user interacting with the vehicle is located based on voice localization, and determine the corresponding pressure value according to the seat pressure sensor of that seat. If the pressure value is greater than the pressure threshold of the seat pressure sensor, it indicates that the user is sitting on that seat, and the sound zone where that seat is located is determined as the sound zone where the user's buttocks are located.

[0157] In this embodiment, with user authorization, the vehicle's visual sensors can collect the user's image information. The multi-screen interaction device can perform target recognition of the buttocks on the image information of the user currently interacting with the vehicle, and determine the vocal range where the recognition result is located as the vocal range of the user's buttocks. This target recognition of the buttocks includes, but is not limited to, target recognition of the buttocks as a whole and / or recognition of related parts of the buttocks. These related parts include, but are not limited to, parts such as the waist and legs that have a positional relationship with the buttocks.

[0158] In some embodiments, the user's sitting posture inside the vehicle is a head-buttock separation posture, including: the user's seat posture information meets the target posture conditions. In some embodiments, the posture information includes one or more of the following: the seat back tilt angle, the seat back pressure, and the seat fore-aft position.

[0159] In some embodiments, if the sound source positioning seat is unoccupied within the sound region where the user's head is located, the user's seat is the seat adjacent to the sound source positioning seat. The sound source positioning seat can be a seat within the sound region determined by the sound location.

[0160] In this embodiment, the multi-screen interaction device can identify the seat within the sound source positioning area as the sound source positioning seat. When no one is seated in this sound source positioning seat, it determines the sound region of the user's buttocks and head based on the posture information of the seat in front of the sound source positioning area, and determines that the sound region of the buttocks and head are inconsistent. Specifically, the multi-screen interaction device can determine the user's sound location through sound source positioning, head image recognition, etc., and identify the sound region where the sound location is located as the sound region of the user's head. Further, it uses seat pressure sensors to determine whether there is a person in the sound source positioning seat within the sound source positioning area. If no one is present, it indicates that the user's actual seating position is likely in the seat in front of this one. Therefore, the seat in front is identified as the user's seat. Furthermore, the posture information of the seat in front is obtained. If the posture information meets the target posture conditions, it means that the seat where the user is located is in a flat or nearly flat position. In this state, the sound zone where the user's head is located is the sound zone where the pronunciation is located, and the sound zone where the user's buttocks are located is the sound zone of the row in front of the sound zone where the head is located. The pronunciation sound zone and the seat position sound zone are not the same.

[0161] In some embodiments, if the user's sitting posture inside the vehicle is a head-buttock separation posture, the interaction between the in-vehicle device and the user is based on the audio signals collected by the audio acquisition device corresponding to the sound zone of the buttocks and the audio acquisition device corresponding to the sound zone of the head. Therefore, the user can interact with the virtual avatar of the voice assistant even when adjusting from a near-lying sitting posture to an upright sitting posture, thus improving the user's interactive experience.

[0162] In some embodiments, controlling the in-vehicle device corresponding to the target audio zone to interact with the user includes: controlling the in-vehicle device corresponding to the target audio zone to display a virtual avatar, the virtual avatar being used to interact with the user.

[0163] In some embodiments, controlling the in-vehicle device corresponding to the target sound zone to interact with the user includes: if the target sound zone is a sound zone other than the driver's sound zone and the passenger's sound zone, and the in-vehicle device corresponding to the target sound zone is in a closed state and / or a folded state, then controlling a backup device to interact with the user; wherein the backup device includes at least one of the central control device, the screen corresponding to the passenger's seat, and the available screen closest to the target sound zone.

[0164] In this embodiment, when determining the target location, the audio region corresponding to the user's buttocks and the head are distinguished, fully considering the user's posture with brain-buttock separation, thus improving the accuracy of target audio region judgment. When the user is in a brain-buttock separation posture, audio signals collected by audio acquisition devices corresponding to the user's buttocks and head audio regions are received. This allows the user to interact with the virtual avatar of the voice assistant even when adjusting from a near-lying posture to an upright sitting posture, improving the user's interactive experience.

[0165] Figure 7 This is a schematic diagram of the structure of a multi-screen interaction device provided in an embodiment of this disclosure. The multi-screen interaction device provided in this embodiment can execute the processing flow provided in the embodiment of the screen interaction method, such as… Figure 7 As shown, the multi-screen interactive device 70 includes:

[0166] Acquisition module 71 is used to acquire audio signals.

[0167] The determination module 72 is used to determine the target location of the user who emitted the audio signal by recognizing the audio signal.

[0168] The control module 73 is used to control the vehicle-mounted device corresponding to the target location to interact with the user via voice.

[0169] Optionally, when determining the target location of the user who emitted the audio signal by recognizing the audio signal, the determining module 72 is specifically used to: determine the pronunciation location of the audio signal by recognizing the audio signal; and determine the target location of the user who emitted the audio signal based on the posture information of the seat corresponding to the pronunciation location and the pronunciation location.

[0170] Optionally, the posture information of the seat corresponding to the sound-producing position is determined based on the seat back angle sensor and the seat fore-and-aft position sensor corresponding to the sound-producing position.

[0171] Optionally, when determining the target location of the user who emitted the audio signal based on the posture information of the seat corresponding to the pronunciation location and the pronunciation location, the determining module 72 is specifically used to: determine whether there is a person on the seat corresponding to the pronunciation location based on at least one of the visual sensor signal and the pressure sensor corresponding to the pronunciation location; if there is a person on the seat corresponding to the pronunciation location, then determine the target location of the user who emitted the audio signal based on the posture information of the seat corresponding to the pronunciation location and the pronunciation location.

[0172] Optionally, when determining the target location of the user emitting the audio signal based on the posture information of the seat corresponding to the pronunciation location and the pronunciation location, the determining module 72 is specifically used to: determine the position of the user's buttocks and the position of the user's head based on the posture information of the seat corresponding to the pronunciation location and the pronunciation location; if the position of the user's buttocks and the position of the user's head are inconsistent, then the position of the user's buttocks is taken as the target location of the user emitting the audio signal; if the position of the user's buttocks and the position of the user's head are consistent, then the position of the user's buttocks or the position of the user's head is taken as the target location of the user emitting the audio signal.

[0173] Optionally, the multi-screen interactive device further includes a receiving module 74, which is used to receive audio signals collected by an audio acquisition device corresponding to the position of the user's buttocks and the position of the user's head if the position of the user's buttocks and the position of the user's head are not the same.

[0174] Optionally, when the control module 73 controls the vehicle-mounted device corresponding to the target location to interact with the user via voice, it is specifically used to: control the vehicle-mounted device corresponding to the target location to display a virtual image, which is used to interact with the user via voice.

[0175] Optionally, the control module 73 is also used to: control an audio playback device around the vehicle-mounted equipment to play the voice broadcast by the virtual avatar.

[0176] Optionally, when the control module 73 controls the in-vehicle device corresponding to the target location to interact with the user via voice, it is further specifically used to: if the in-vehicle device corresponding to the target location is in an open state, control the in-vehicle device corresponding to the target location to interact with the user via voice.

[0177] Figure 7 The multi-screen interaction device shown in the embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effect are similar, and will not be repeated here.

[0178] In other embodiments, the screen projection control device includes: a determining module, configured to determine the sitting sound zone as the user's target sound zone if the vocal sound zone corresponding to the user generating the audio signal in the vehicle is inconsistent with the sitting sound zone where the user is located; and a control module, configured to control the in-vehicle device corresponding to the target sound zone to interact with the user.

[0179] Optionally, the pronunciation zone corresponding to the user generating the audio signal in the vehicle is inconsistent with the seating zone of the user, including: determining the pronunciation zone and the seating zone; if the pronunciation zone and the seating zone are different zones in the vehicle, then determining that the pronunciation zone and the seating zone are inconsistent.

[0180] Optionally, the pronunciation range is determined by identifying the audio signal, or the pronunciation range is determined by identifying the image information collected by the visual sensor inside the vehicle.

[0181] Optionally, the seating sound zone is determined by a seat pressure sensor, or the seating sound zone is determined by identifying image information collected by a vision sensor inside the vehicle.

[0182] Optionally, the pronunciation zone corresponding to the user generating the audio signal in the vehicle is inconsistent with the seating zone of the user, including: determining that the pronunciation zone and the seating zone are inconsistent when the posture information of the user's seat meets the target posture conditions.

[0183] Optionally, the posture information includes one or more of the following: the tilt angle of the seat back, the pressure of the seat back, and the fore-and-aft position of the seat.

[0184] Optionally, if the vocal range corresponding to the user generating the audio signal in the vehicle is inconsistent with the seating range of the user, then the seating range is determined as the target vocal range of the user. This includes: if there is someone in the seat corresponding to the vocal position, then determining whether the vocal range and the seating range are consistent; if the vocal range corresponding to the user generating the audio signal in the vehicle is inconsistent with the seating range of the user, then the seating range is determined as the target vocal range of the user.

[0185] Optionally, the multi-screen interactive device further includes a receiving module: the receiving module is used to receive the audio signal collected by the audio acquisition device corresponding to the seating area and the audio signal collected by the audio acquisition device corresponding to the vocal area if the vocal area corresponding to the user generating the audio signal in the vehicle is inconsistent with the seating area where the user is located.

[0186] Optionally, the determining module is further configured to: if the vocal register corresponding to the user generating the audio signal in the vehicle is consistent with the seating register of the user, then determine the seating register or the vocal register as the user's target register.

[0187] Optionally, the control module is specifically used to: control the in-vehicle device corresponding to the target audio zone to display a virtual image, the virtual image being used to interact with the user.

[0188] Optionally, the control module is further configured to: control the audio playback device corresponding to the vehicle-mounted equipment to play the voice broadcast by the virtual avatar.

[0189] Optionally, the control module is specifically used to: if the target sound zone is another sound zone other than the driver's sound zone and the passenger's sound zone, and the in-vehicle device corresponding to the target sound zone is in a closed state and / or a folded state, then control the backup device to interact with the user; wherein, the backup device includes at least one of the central control device, the screen corresponding to the passenger's seat, and the available screen closest to the target sound zone.

[0190] In other embodiments, the screen projection control device includes: a determination module, used to determine the sound zone where the user's buttocks are located as the target sound zone if the user's sitting posture in the vehicle is a brain-buttock separation sitting posture; and a control module, used to control the in-vehicle device corresponding to the target sound zone to interact with the user.

[0191] Optionally, the user's sitting posture inside the vehicle is a brain-buttock separation sitting posture, including: the sound zone where the user's buttocks are located is inconsistent with the sound zone where the user's head is located; or, the posture information of the user's seat meets the target posture conditions.

[0192] Optionally, the posture information includes one or more of the following: the tilt angle of the seat back, the pressure of the seat back, and the fore-and-aft position of the seat.

[0193] Optionally, if there is no one on the sound source positioning seat in the sound zone where the user's head is located, then the seat where the user is located is the seat adjacent to the sound source positioning seat.

[0194] Optionally, controlling the in-vehicle device corresponding to the target audio zone to interact with the user includes: controlling the in-vehicle device corresponding to the target audio zone to display a virtual avatar, the virtual avatar being used to interact with the user.

[0195] Optionally, controlling the in-vehicle device corresponding to the target sound zone to interact with the user includes: if the target sound zone is a sound zone other than the driver's sound zone and the passenger's sound zone, and the in-vehicle device corresponding to the target sound zone is in a closed state and / or a folded state, then controlling a backup device to interact with the user; wherein, the backup device includes at least one of the central control device, the screen corresponding to the passenger's seat, and the available screen closest to the target sound zone.

[0196] Figure 8 This is a schematic diagram of the structure of an in-vehicle device provided in an embodiment of this disclosure. The in-vehicle device provided in this embodiment of the disclosure can execute the processing flow provided in the multi-screen interaction method embodiment, such as... Figure 8As shown, the vehicle-mounted device 80 includes: a memory 81, a processor 82, a computer program, and a communication interface 83; wherein the computer program is stored in the memory 81 and configured to be executed by the processor 82 using the multi-screen interaction method described above.

[0197] In addition, this disclosure also provides a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the multi-screen interaction method described in the above embodiments.

[0198] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0199] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0200] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0201] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0202] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.

[0203] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0204] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0205] Furthermore, this disclosure also provides a computer program product comprising a computer program or instructions that, when executed by a processor, implement the battery remaining life estimation method as described above.

[0206] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0207] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A multi-screen interaction method, characterized in that, The method includes: If the vocal register of the user generating the audio signal inside the vehicle is inconsistent with the seating register of the user, then the seating register is determined as the target vocal register of the user. Control the in-vehicle device corresponding to the target audio region to interact with the user.

2. The method according to claim 1, characterized in that, The vocal register corresponding to the user generating the audio signal inside the vehicle is inconsistent with the vocal register of the user's seat location, including: Determine the pronunciation range and the seating range; If the vocalization area and the seating area are different vocalization areas within the vehicle, then it is determined that the vocalization area and the seating area are inconsistent.

3. The method according to claim 1 or 2, characterized in that, The pronunciation region is determined by identifying the audio signal, and / or the pronunciation region is determined by identifying the image information collected by the visual sensor inside the vehicle.

4. The method according to any one of claims 1-3, characterized in that, The seating sound zone is determined by a seat pressure sensor, and / or the seating sound zone is determined by identifying image information collected by a vision sensor inside the vehicle.

5. The method according to claims 1-4, characterized in that, The vocal register corresponding to the user generating the audio signal inside the vehicle is inconsistent with the vocal register of the user's seat location, including: If the user's seat posture information meets the target posture conditions, it is determined that the vocalization region and the seating region are inconsistent.

6. The method according to claim 5, characterized in that, The posture information includes one or more of the following: the tilt angle of the seat back, the pressure of the seat back, and the fore-and-aft position of the seat.

7. The method according to claim 5, characterized in that: If no one is seated in the vocal range, the user is seated in the seat adjacent to the seat in the vocal range.

8. The method according to any one of claims 1-7, characterized in that, The method further includes: If the vocal register of the user generating the audio signal inside the vehicle is consistent with the seating register of the user, then the seating register or the vocal register is determined as the user's target vocal register.

9. The method according to any one of claims 1-8, characterized in that, The control of the in-vehicle device corresponding to the target audio region to interact with the user includes: The vehicle-mounted device corresponding to the target audio region is controlled to display a virtual avatar, which is used to interact with the user.

10. The method according to claim 9, characterized in that, The method further includes: Control the audio playback device corresponding to the vehicle-mounted equipment to play the voice broadcast by the virtual avatar.

11. The method according to any one of claims 1-10, characterized in that, The control of the in-vehicle device corresponding to the target sound zone to interact with the user includes: If the target sound zone is a sound zone other than the driver's sound zone and the passenger's sound zone, and the in-vehicle device corresponding to the target sound zone is in a closed state and / or folded state, then the backup device is controlled to interact with the user; wherein, the backup device includes at least one of the central control device, the screen corresponding to the passenger, and the available screen closest to the target sound zone.

12. A multi-screen interaction method, characterized in that, The method includes: If the user's sitting posture inside the vehicle is a brain-buttock separation posture, then the vocal range where the user's buttocks are located will be used as the target vocal range. Control the in-vehicle device corresponding to the target audio region to interact with the user.

13. The method according to claim 12, characterized in that, The user's sitting posture inside the vehicle is a brain-buttock separation sitting posture, including: The sound zone where the user's buttocks are located is inconsistent with the sound zone where the user's head is located; or, the posture information of the user's seat meets the target posture conditions.

14. The method according to claim 13, characterized in that, The posture information includes one or more of the following: the tilt angle of the seat back, the pressure of the seat back, and the fore-and-aft position of the seat.

15. The method according to claim 13, characterized in that, If there is no one in the sound source positioning seat within the sound zone where the user's head is located, then the user's seat is the seat adjacent to the sound source positioning seat.

16. The method according to any one of claims 12-15, characterized in that, The control of the in-vehicle device corresponding to the target sound zone to interact with the user includes: The vehicle-mounted device corresponding to the target audio region is controlled to display a virtual avatar, which is used to interact with the user.

17. The method according to any one of claims 12-16, characterized in that, The control of the in-vehicle device corresponding to the target sound zone to interact with the user includes: If the target sound zone is a sound zone other than the driver's sound zone and the passenger's sound zone, and the in-vehicle device corresponding to the target sound zone is in a closed state and / or folded state, then the backup device is controlled to interact with the user; wherein, the backup device includes at least one of the central control device, the screen corresponding to the passenger, and the available screen closest to the target sound zone.

18. A multi-screen interactive device, characterized in that, The device includes: The determination module is used to determine the target sound zone of the user if the vocal sound zone corresponding to the user who generates the audio signal in the vehicle is inconsistent with the seat sound zone where the user is located. The control module is used to control the interaction between the in-vehicle device corresponding to the target sound zone and the user.

19. A vehicle-mounted device, characterized in that, include Memory; processor; as well as Computer programs; The computer program is stored in the memory and configured to be executed by the processor to implement the method as described in any one of claims 1-17.

20. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-17.