Multi-screen Interaction Method, Device, Equipment and Computer Readable Storage Medium
The user's location is determined through audio signal recognition and the on-board equipment that controls the target location to perform voice interaction, solving the problem of inconvenience between the rear user and the voice assistant in the car, improving the user experience and reducing interference to users in other locations.
Patent Information
- Application Number
- CN202210306313.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-25
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-03-25
AI Technical Summary
When the rear user interacts with the voice assistant in the car through voice mode, the response will be displayed on the central control screen, resulting in poor interaction experience and interfering with the driving of the front user.
The user's location is determined through audio signal recognition, and the on-board equipment corresponding to the target position is controlled to perform voice interaction, display virtual images and broadcast voice.
It improves the interactive experience of the rear row users, reduces interference to the front row users, and enhances the convenience of multi-screen interaction.
Smart Images

Figure CN115431762B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of human - computer interaction technologies, and particularly to a multi - screen interaction method, apparatus, device, and computer - readable storage medium. Background Art
[0002] With the continuous development of the automotive industry, for the convenience of users, multiple screens are installed in the vehicle; for example, generally, 2 screens are installed in the vehicle. In addition to the central control screen corresponding to the driver's seat, a rear - row screen is also installed.
[0003] Generally, when a user interacts with a voice assistant supported by a vehicle computer through voice, the vehicle computer is usually set in a fixed place in the vehicle. For example, the vehicle computer is integrated with the central control screen corresponding to the driver's seat. When a user far from the vehicle computer, such as a user in the rear row of the vehicle, interacts through voice, the central control screen responds to the user's interaction request, which brings inconvenience to the rear - row user. For example, when a rear - row user wakes up the voice assistant through voice, the image of the voice assistant is displayed on the central control screen, which not only results in no interaction experience for the rear - row user, but also interferes with the driving of the front - row user, etc. Summary of the Invention
[0004] To solve the above - mentioned technical problems or at least partially solve the above - mentioned technical problems, the present disclosure provides a multi - screen interaction method, apparatus, device, and computer - readable storage medium, enabling a user to perform voice interaction with an in - vehicle device corresponding to the user's location, improving the user's interaction experience, and reducing interference with users in other locations at the same time.
[0005] In a first aspect, an embodiment of the present disclosure provides a multi - screen interaction method, including:
[0006] Obtain an audio signal;
[0007] Determine a target location where the user who emits the audio signal is located by identifying the audio signal;
[0008] Control the in - vehicle device corresponding to the target location to perform voice interaction with the user.
[0009] In a second aspect, an embodiment of the present disclosure provides a multi - screen interaction apparatus, including:
[0010] An obtaining module, configured to obtain an audio signal;
[0011] A determining module, configured to determine a target location where the user who emits the audio signal is located by identifying the audio signal;
[0012] A control module, configured to control the in - vehicle device corresponding to the target location to perform voice interaction with the user.
[0013] In a third aspect, an embodiment of the present disclosure provides a vehicle-mounted device, including:
[0014] a memory;
[0015] a processor; and
[0016] a computer program;
[0017] wherein, the computer program is stored in the memory and is configured to be executed by the processor to implement the method as described in the first aspect.
[0018] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to implement the method as described in the first aspect.
[0019] In a fifth aspect, an embodiment of the present disclosure further provides a computer program product, which includes a computer program or instruction, and when the computer program or instruction is executed by a processor, it implements the multi-screen interaction method as described above.
[0020] The multi-screen interaction method, device, equipment, and computer-readable storage medium provided by the embodiments of the present disclosure identify the acquired audio signal to determine the target position where the user sending the audio signal is located, and control the vehicle-mounted device corresponding to the target position to perform voice interaction with the user, so that the user can perform voice interaction with the vehicle-mounted device corresponding to the user's location, improving the user's interaction experience and reducing interference to users in other positions at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.
[0022] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0023] Figure 1 It is a flowchart of a multi-screen interaction method provided by an embodiment of the present disclosure;
[0024] Figure 2 It is a vehicle hardware architecture diagram corresponding to a multi-screen interaction method provided by an embodiment of the present disclosure;
[0025] Figure 3 It is another flowchart of a multi-screen interaction method provided by an embodiment of the present disclosure;
[0026] Figure 4 Another vehicle hardware architecture diagram corresponding to the multi-screen interaction method provided by the embodiments of the present disclosure;
[0027] Figure 5 Schematic diagram of the implementation principle of another multi-screen interaction method provided by the embodiments of the present disclosure;
[0028] Figure 6 Flowchart of another multi-screen interaction method provided by the embodiments of the present disclosure;
[0029] Figure 7 Schematic diagram of the structure of the multi-screen interaction device provided by the embodiments of the present disclosure;
[0030] Figure 8 Schematic diagram of the structure of the in-vehicle device provided by the embodiments of the present disclosure. Detailed implementation manners
[0031] In order to more clearly understand the above objects, features and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other.
[0032] In the following description, many specific details are set forth to facilitate a thorough understanding of the present disclosure, but the present disclosure may be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all the embodiments.
[0033] Generally, users will interact with the voice interaction system supported in the vehicle by voice, but the responses of the voice interaction system are all displayed on the center control screen. For example, when a user in the back row wakes up the voice assistant by voice, but the image of the voice assistant is displayed on the center control screen, which not only results in no interaction experience for the user in the back row, but also causes interference to the driver in the front row and the like.
[0034] To solve this problem, the embodiments of the present disclosure provide a multi-screen interaction method, which will be introduced below in combination with specific embodiments.
[0035] Figure 1 Flowchart of the multi-screen interaction method provided by the embodiments of the present disclosure; Figure 2 Vehicle hardware architecture diagram provided by the embodiments of the present disclosure. Before introducing the multi-screen interaction method, first introduce the Figure 2 vehicle hardware architecture diagram shown.
[0036] The overall architecture of the vehicle hardware is as follows: The car is divided into different sound zones, and different numbers of screens, audio collection devices, and audio playback devices are installed at different positions in the car as needed. A voice assistant is installed in the car. Among them, the audio collection device can be a microphone; the audio playback device can be a speaker, a horn, etc. As Figure 2 shown, the hardware architecture diagram includes the entire vehicle 20, the car machine 21, sound zone 22, sound zone 23, sound zone 24, sound zone 25, microphone 26 and speaker 27 corresponding to sound zone 22, microphone 28 and speaker 29 corresponding to sound zone 23, microphone 210 and speaker 211 corresponding to sound zone 24, microphone 212 and speaker 213 corresponding to sound zone 25, the central control screen 214 corresponding to the driver's seat position, the screen 215 corresponding to the passenger seat position, and the screen 216 corresponding to the rear row position. The above microphones, speakers, and screens are all connected to the car machine 21, and the voice assistant (not shown in the figure) is integrated into the car machine.
[0037] Next, in combination with Figure 2 the vehicle hardware architecture diagram shown, Figure 1 the multi-screen interaction method shown will be introduced. This method mainly includes the following steps:
[0038] S101. Obtain an audio signal.
[0039] Specifically, in this embodiment, different areas in the vehicle are divided into different sound zones, such as Figure 2 sound zone 22, sound zone 23, sound zone 24, sound zone 25 shown; microphones and speakers are installed in different sound zones, such as Figure 2 the microphones and speakers corresponding to each sound zone shown. Among them, the microphone can collect the user's audio signal, and the speaker can play the user's audio signal. When the user makes a sound, one or more microphones can collect the user's audio signal, play it through the corresponding speaker, and send the audio signal to the car machine.
[0040] S102. Determine the target position of the user who emits the audio signal by identifying the audio signal.
[0041] Specifically, identifying the audio signal is to convert a section of audio signal into corresponding text information. Optionally, the process of converting the audio signal into corresponding text information mainly includes several processes such as feature extraction, acoustic model, language model, and dictionary and decoding.
[0042] Specifically, before feature extraction, in order to more effectively highlight features, it is often necessary to perform preprocessing operations such as filtering and framing on the collected audio signal, extract the collected audio signal from the original signal, and then perform feature extraction to convert the sound signal from the time domain to the frequency domain to provide a suitable feature vector for the acoustic model; in the acoustic model, calculate the score of each feature vector on the acoustic features according to the acoustic features; then, the language model calculates the probability of the possible phrase sequence corresponding to the audio signal according to relevant linguistic theories; finally, according to the existing dictionary, decode the phrase sequence to obtain the final possible text representation.
[0043] In this embodiment, when the text content included in the audio signal includes wake-up words such as "Hello, please turn on the machine", "hello, hello", etc., the position where the user who emits the above audio signal is located is determined as the target position. Among them, the wake-up word can be understood as the word that wakes up the voice interaction system of the intelligent device. Specifically in this embodiment, the wake-up word is the word that wakes up the in-vehicle voice assistant.
[0044] Exemplarily, the target position can be determined according to the signal intensity and azimuth of the target audio signal collected by each microphone. For example, in the architecture diagram shown, when a user sitting in the co-pilot speaks, it is possible that multiple microphones can collect the audio signal. Therefore, multiple microphones transmit the audio signals they collect to the in-vehicle computer for processing. The in-vehicle computer determines that the intensity of the audio signal collected by microphone 28 is the largest, and at the same time the direction source of the audio signal is the right half of the vehicle. Therefore, the target position where the user is located is determined to be the co-pilot. Figure 2 As shown in the architecture diagram, when a user sitting in the co-pilot speaks, it is possible that multiple microphones can collect the audio signal. Therefore, multiple microphones transmit the audio signals they collect to the in-vehicle computer for processing. The in-vehicle computer determines that the intensity of the audio signal collected by microphone 28 is the largest, and at the same time the direction source of the audio signal is the right half of the vehicle. Therefore, the target position where the user is located is determined to be the co-pilot.
[0045] S103. Control the in-vehicle device corresponding to the target position to perform a voice interaction with the user.
[0046] Optionally, in this embodiment, controlling the in-vehicle device corresponding to the target position to perform a voice interaction with the user includes: controlling the in-vehicle device corresponding to the target position to display a virtual image, and the virtual image is used to perform a voice interaction with the user.
[0047] Specifically, in this embodiment, only one voice assistant is set in the whole vehicle. After determining the target position of the user who wakes up the voice assistant, the virtual image of the voice assistant is displayed on the screen corresponding to the target position. Exemplarily, when the target position of the user who wakes up the voice assistant is the rear row position of the vehicle, the virtual image is displayed on the rear row screen 216 corresponding to the rear row position; when the target position of the user who wakes up the voice assistant is the driver's seat position, the virtual image is displayed on the central control screen 214 corresponding to the driver's seat position; when the target position of the user who wakes up the voice assistant is the co-pilot position, the virtual image is displayed on the co-pilot screen 215 corresponding to the co-pilot position.
[0048] Optionally, in this embodiment, after determining the screen for virtual avatar display, control the audio playback device corresponding to the vehicle-mounted device to play the voice broadcast by the virtual avatar.
[0049] Specifically, in this embodiment, after determining the display screen of the virtual avatar, control the audio playback device corresponding to the virtual avatar display screen to play the voice broadcast by the virtual avatar. Exemplarily, when the virtual avatar is displayed on the rear row screen 216, control the audio playback device 212 corresponding to the rear row screen (for example, the user who emits the audio signal is in the left rear position) to play the voice broadcast by the virtual avatar; when the virtual avatar is displayed on the center control screen 214, control the audio playback device 27 corresponding to the center control screen to play the voice broadcast by the virtual avatar; when the virtual avatar is displayed on the co-pilot screen 215, control the audio playback device 29 corresponding to the co-pilot screen to play the voice broadcast by the virtual avatar.
[0050] It can be understood that in this embodiment, since there is only one voice assistant, when a user at a position other than the target position wakes up the voice assistant, that is, when a new target position appears, the virtual avatar of the voice assistant will be displayed on the screen corresponding to the new target position, and at the same time, the voice broadcast by the virtual avatar will also be played on the audio playback device corresponding to the screen corresponding to the new target position.
[0051] In the embodiment of the present disclosure, by identifying the acquired audio signal, determine the target position where the user sending the audio signal is located, and control the vehicle-mounted device corresponding to the target position to perform voice interaction with the user, so that the user can perform voice interaction with the vehicle-mounted device corresponding to the user's location, improving the user's interaction experience and reducing interference to users in other positions at the same time.
[0052] Figure 3 This is another multi-screen control method provided by the embodiment of the present disclosure. Figure 4 This is another vehicle hardware architecture diagram provided by the embodiment of the present disclosure. Figure 5 This is the implementation principle diagram of another multi-screen control method provided by the embodiment of the present disclosure. Before introducing the multi-screen control method provided in this embodiment, first introduce Figure 4 the vehicle hardware architecture diagram shown.
[0053] In Figure 2 on the basis of the hardware architecture diagram shown, Figure 4 the vehicle hardware architecture diagram shown further includes: a camera, an automotive seat sensor. Among them, the automotive seat sensor is a thin-film type contact sensor, and the contacts of the sensor are evenly distributed on the force-bearing surface of the seat, and a trigger signal is generated when the seat is subjected to external pressure. As Figure 4 shown, in Figure 2Based on this, the vehicle hardware architecture diagram further includes a camera 41 corresponding to the driver's seat, a camera 42 corresponding to the co-driver's seat, a camera 43 corresponding to the rear seats, a seat pressure sensor 44 corresponding to the driver's seat, a seat backrest angle sensor 45 corresponding to the driver's seat, a seat front-back position sensor 46 corresponding to the driver's seat, a seat pressure sensor 47 corresponding to the co-driver's seat, a seat backrest angle sensor 48 corresponding to the co-driver's seat, a seat front-back position sensor 49 corresponding to the co-driver's seat, a seat pressure sensor 410 corresponding to the left rear seat, a seat backrest angle sensor 411 corresponding to the left rear seat, a seat front-back position sensor 412 corresponding to the left rear seat, a seat pressure sensor 413 corresponding to the right rear seat, a seat backrest angle sensor 414 corresponding to the right rear seat, and a seat front-back position sensor 415 corresponding to the right rear seat.
[0054] Next, in combination with Figure 4 and Figure 5 for Figure 3 Another multi-screen control method shown in the figure will be introduced. This method includes the following steps:
[0055] S301. Obtain an audio signal.
[0056] In this embodiment, this step is the same as step S101 and will not be elaborated here.
[0057] S302. Determine the pronunciation position of the audio signal by identifying the audio signal.
[0058] Specifically, in this embodiment, when the text content included in the audio signal includes wake-up words such as "Hello, please turn on the machine", "Hello, hello", etc., the position where the user who emits the above audio signal is located is determined as the pronunciation position.
[0059] Exemplarily, to determine the pronunciation position, the target position can be determined based on the signal intensity and phase of the target audio signal collected by each microphone. For example Figure 5 the pronunciation positioning shown.
[0060] S303. Determine the target position where the user who emits the audio signal is located based on the pose information of the seat corresponding to the pronunciation position and the pronunciation position.
[0061] Specifically, after determining the pronunciation position of the user, the target position where the user is located is determined based on the pose information of the seat corresponding to the pronunciation position. Among them, the pose information of the seat can specifically be the seat backrest information and the position information of the seat.
[0062] Optionally, the pose information of the seat corresponding to the pronunciation position can be determined according to the seat back angle sensor and the seat front-back position sensor corresponding to the pronunciation position, that is, determined by using the seat back sensor and the seat front-back position sensor as shown in Figure 5 shown.
[0063] Optionally, according to the pose information of the seat corresponding to the pronunciation position and the pronunciation position, determining the target position where the user who emits the audio signal is located can specifically be determining the position where the buttocks of the user who emits the audio signal are located and the position where the head is located according to the pose information of the seat corresponding to the pronunciation position and the pronunciation position.
[0064] Specifically, the seat back angle sensor is used to judge the tilt angle of the seat back and the pressure at each position of the seat back. According to the tilt angle of the seat back and the pressure at each position of the seat back, the positions of the user's buttocks and head can be judged; the seat front-back position sensor is used to judge the front-back position of the seat. According to the front-back position of the seat and in combination with the seat back angle sensor, the positions of the user's buttocks and head can be further judged.
[0065] Exemplarily, when the sound generation position is the co-pilot position and the co-pilot position user has not adjusted the angle of the seat back and the front-back position of the seat, it can be judged according to the Figure 4 shown seat back angle sensor 48 and seat front-back position sensor 49 that the user's current sitting posture is approximately sitting upright. When the co-pilot position user adjusts the angle of the seat back and the seat is approximately on the same horizontal plane and at the same time adjusts the co-pilot seat backward approximately close to the rear seat, it can be judged according to the Figure 4 shown seat back angle sensor 48 and seat front-back position sensor 49 that the user's current sitting posture is approximately lying flat. In this embodiment, the user's approximately lying flat sitting posture is called brain-hip separation, which can specifically be understood as that the user's buttocks and the user's head are not in the same sound area.
[0066] In this embodiment, after determining the position where the buttocks of the user who emits the audio signal are located and the position where the head is located according to the pose information of the seat corresponding to the pronunciation position and the pronunciation position, if the position where the buttocks of the user are located and the position where the head is located are the same, then the position where the buttocks of the user are located or the position where the head is located is used as the target position where the user who emits the audio signal is located.
[0067] Specifically, when the sound area corresponding to the position where the buttocks of the user are located and the sound area corresponding to the position where the head of the user is located are in the same sound area, for example, when the user sits upright in the co-pilot position as described above, the position where the buttocks of the user are located and the position where the head is located are in the same sound area, that is, both the position where the buttocks of the user are located and the position where the head is located are in Figure 4The pitch range 23 shown. Therefore, the position where the user's hip or head is located is used as the target position of the user who emits the audio signal.
[0068] If the position where the user's hip is located is different from the position where the user's head is located, then the position where the user's hip is located is used as the target position of the user who emits the audio signal.
[0069] Specifically, when the pitch range corresponding to the position where the user's hip is located and the pitch range corresponding to the position where the user's head is located are not in the same pitch range, for example, when the user is in the co-pilot position and has a separated brain-hip sitting posture as described above, the user's hip and the user's head are not in the same pitch range, that is, the position where the user's hip is located is in Figure 4 the pitch range 23 shown, and the position where the user's head is located is in Figure 4 the pitch range 25 shown. When the user is in a separated brain-hip sitting posture, the pitch range where the user's hip is located is used as the target pitch range, and the position where the hip is located is used as the target position of the user who emits the audio signal.
[0070] Optionally, in some other embodiments, it may also be to first determine whether there is someone on the seat corresponding to the pronunciation position according to at least one of the visual sensor signal corresponding to the pronunciation position and the pressure sensor corresponding to the pronunciation position; if there is someone on the seat corresponding to the pronunciation position, then determine the target position of the user who emits the audio signal according to the pose information of the seat corresponding to the pronunciation position and the pronunciation position.
[0071] Specifically, the visual sensor signal corresponding to the pronunciation position may specifically be the image information captured by the camera at the pronunciation position, for example Figure 5 the visual positioning shown; the pressure sensor corresponding to the pronunciation position may specifically be the pressure value of the seat at the pronunciation position collected by the pressure sensor corresponding to the pronunciation position, for example Figure 5 the seat pressure sensor shown. Determine whether there is someone on the seat according to at least one of the image information captured by the camera at the pronunciation position and the pressure value of the seat at the pronunciation position.
[0072] After it is determined that there is someone on the seat corresponding to the pronunciation position, determine the position where the user's hip is located and the position where the user's head is located according to the pose information of the seat corresponding to the pronunciation position and the pronunciation position; if the position where the user's hip is located is different from the position where the user's head is located, then the position where the user's hip is located is used as the target position of the user who emits the audio signal; if the position where the user's hip is located is the same as the position where the user's head is located, then the position where the user's hip is located or the position where the user's head is located is used as the target position of the user who emits the audio signal.
[0073] Optionally, in some embodiments, the above-mentioned pronunciation positioning, visual positioning, seat pressure sensor, seat back angle sensor, and seat front-back position sensor can be integrated into a spatial positioning model as shown in Figure 5 to determine the positions of the user's buttocks and head, and further determine the target position.
[0074] Furthermore, if the positions of the user's buttocks and head are inconsistent, the audio signal collected by the audio collection device corresponding to the position of the user's buttocks and the audio signal collected by the audio collection device corresponding to the position of the user's head are received.
[0075] Specifically, when the user's buttocks and head are in different sound zones, all the audio signals collected by the buttocks and head audio collection devices are sent to the vehicle head unit. For example, when the user in the co-pilot position is in a brain-buttocks separated sitting posture as described above, the Figure 4 audio signal collected by the microphone 28 and the audio signal collected by the microphone 212 shown are sent to the vehicle head unit, and the vehicle head unit receives the audio signal and sends it to the voice assistant for subsequent interaction; the audio signals collected by the microphones corresponding to other positions are sent to the vehicle head unit, and the vehicle head unit does not receive the audio signal for subsequent interaction.
[0076] It can be understood that when the user's buttocks and head are in the same sound zone, the audio collection devices corresponding to the buttocks and head are the same. Therefore, the signal collected by this audio collection device is sent to the vehicle head unit, and the vehicle head unit receives the audio signal and sends it to the voice assistant for subsequent interaction; the audio signals collected by the microphones corresponding to other positions are sent to the vehicle head unit, and the vehicle head unit does not receive the audio signal for subsequent interaction.
[0077] Optionally, in some embodiments, during the interaction between the user at the target position and the virtual image of the voice assistant, when the audio signal emitted by the user at other positions contains a wake-up word, after the audio signal is collected by the microphone corresponding to the other position and sent to the vehicle head unit, the vehicle head unit receives the audio signal and executes steps S302 - S303 to determine a new target position.
[0078] S304. Control the in-vehicle device corresponding to the target position to perform a voice interaction with the user.
[0079] In this embodiment, this step is the same as step S103 and will not be elaborated here.
[0080] In the embodiment of the present disclosure, when determining the target position, the positions of the user's buttocks and head are distinguished, and the sitting posture with the user's brain and buttocks separated is fully considered, improving the accuracy of target position judgment. Further, when the user is in the sitting posture with the brain and buttocks separated, the audio signals collected by the audio collection devices corresponding to the positions of the user's buttocks and head are received, ensuring that the user can interact with the virtual image of the voice assistant whether adjusting from an approximately lying-down sitting posture to a sitting-up straight posture or from a sitting-up straight posture to an approximately lying-down sitting posture, improving the user's interaction experience.
[0081] Figure 6 Another multi-screen interaction method provided by the embodiment of the present disclosure, the method includes the following steps:
[0082] S601. Obtain an audio signal.
[0083] In this embodiment, this method is the same as the step S101, and will not be described in detail here.
[0084] S602. Determine the target position where the user who emits the audio signal is located by identifying the audio signal.
[0085] In this embodiment, this step is the same as the step S303, and will not be described in detail here.
[0086] S603. If the in-vehicle device corresponding to the target position is in an open state, control the in-vehicle device corresponding to the target position to perform a voice interaction with the user.
[0087] Specifically, in this embodiment, when controlling the in-vehicle device corresponding to the target position to perform a voice interaction with the user, it is necessary to determine whether the screen corresponding to the target position is in an open state.
[0088] Exemplarily, if the target position is a position other than the driver's seat and the passenger seat, and the in-vehicle device corresponding to the other position is in a closed state and / or a folded state, control the central control device to perform a voice interaction with the user.
[0089] Specifically, in some vehicles, when the vehicle is in use, the screens corresponding to the driver's seat and the front passenger seat will be automatically turned on, and as long as the vehicle is in use, the screens corresponding to the driver's seat and the front passenger seat cannot be physically turned off; however, the screen corresponding to the rear seat can be physically turned off or folded. Therefore, when the target position of the user who emits the audio signal is the rear seat (e.g., the left rear seat) and the screen corresponding to the rear seat (e.g., the left rear seat) is in the off state and / or the folded state, resulting in the virtual image of the voice assistant being unable to be displayed on the screen corresponding to the rear seat (e.g., the left rear seat), the virtual image will be controlled to be displayed on the center control screen corresponding to the driver's seat, and at the same time, the audio playback device is the audio playback device corresponding to the center control screen of the driver's seat.
[0090] Since the target position of the user who emits the audio signal is in the rear seat (e.g., the left rear seat), the vehicle computer receives the audio signal collected by the audio collection device corresponding to the rear seat (e.g., the left rear seat).
[0091] Optionally, in other vehicles, when the target position of the user who emits the audio signal is the rear seat (e.g., the left rear seat) and the screen corresponding to the rear seat (e.g., the left rear seat) is in the off state and / or the folded state, the virtual image can also be displayed on the front passenger seat. Therefore, the audio playback device is the audio playback device corresponding to the screen of the front passenger seat; since the target position of the user who emits the audio signal is in the rear seat (e.g., the left rear seat), the vehicle computer receives the audio signal collected by the audio collection device corresponding to the rear seat (e.g., the left rear seat).
[0092] Optionally, in some other vehicles, when the target position of the user who emits the audio signal is the rear seat (e.g., the left rear seat) and the screen corresponding to the rear seat (e.g., the left rear seat) is in the off state and / or the folded state, the virtual image can also be displayed on the screen closer to the target position (e.g., the left rear seat). Exemplarily, when the vehicle has three rows of seats, there may be a screen corresponding to the driver's seat, a screen corresponding to the front passenger seat, a screen corresponding to the second row of seats, a screen corresponding to the third row of seats, etc. When a user in the third row (e.g., the left side of the third row) emits an audio signal and the screen corresponding to the third row (e.g., the left side of the third row) is in the off state and / or the folded state, the virtual image is displayed on the screen corresponding to the second row of seats. Therefore, the audio playback device is the audio playback device corresponding to the screen of the second row of seats; since the target position of the user who emits the audio signal is in the third row (e.g., the left side of the third row), the vehicle computer receives the audio signal collected by the audio collection device corresponding to the third row (e.g., the left side of the third row).
[0093] It should be noted that when the screen corresponding to the target position is in the off state and / or the folded state, the processing logics of the virtual image, the audio playback device, and the audio collection device provided by the embodiments of the present disclosure are only partially feasible technical solutions. The embodiments of the present disclosure do not limit the specific processing logics of the virtual image, the audio playback device, and the audio collection device. Other variant solutions for the above feasible technical solutions are within the protection scope of this embodiment.
[0094] It can be understood that although the central control screen corresponding to the driver's seat and the screen corresponding to the passenger seat cannot be physically turned off, the central control screen corresponding to the driver's seat and the screen corresponding to the passenger seat will have a screen-off state. Therefore, when the virtual image needs to be displayed on the central control screen corresponding to the driver's seat and the screen corresponding to the passenger seat, but the central control screen corresponding to the driver's seat and the screen corresponding to the passenger seat are in the screen-off state, the central control screen corresponding to the driver's seat and the screen corresponding to the passenger seat will be lit. Similarly, except for the screens corresponding to the driver's seat and the passenger seat, the states of the screens at other positions are not in the off and / or folded states, but in the screen-off state. When the virtual image needs to be displayed on the screens corresponding to other positions, the screens corresponding to other positions will also be lit.
[0095] The embodiments of the present disclosure determine the virtual image display position according to the screen state, so that when the screen corresponding to the target position is in the off and / or folded state, the virtual image can be displayed on other screens, ensuring the requirement that there must be a response when awakened, and further improving the user's interaction experience.
[0096] Figure 7 It is a schematic structural diagram of a multi-screen interaction device provided by the embodiments of the present disclosure. The multi-screen interaction device provided by the embodiments of the present disclosure can execute the processing flow provided by the embodiments of the screen interaction method, such as Figure 7 shown, the multi-screen interaction device 70 includes:
[0097] An acquisition module 71, configured to acquire an audio signal.
[0098] A determination module 72, configured to determine the target position where the user who emits the audio signal is located by identifying the audio signal.
[0099] A control module 73, configured to control the in-vehicle device corresponding to the target position to perform voice interaction with the user.
[0100] Optionally, when the determination module 72 is used to determine the target position where the user who emits the audio signal is located by identifying the audio signal, it is specifically used for: identifying the audio signal to determine the pronunciation position of the audio signal; determining the target position where the user who emits the audio signal is located according to the pose information of the seat corresponding to the pronunciation position and the pronunciation position.
[0101] Optionally, the pose information of the seat corresponding to the pronunciation position is determined according to the seat back angle sensor and the seat front-back position sensor corresponding to the pronunciation position.
[0102] Optionally, when the determination module 72 is used to determine the target position where the user who emits the audio signal is located according to the pose information of the seat corresponding to the pronunciation position and the pronunciation position, it is specifically used for: determining whether there is someone on the seat corresponding to the pronunciation position according to at least one of the visual sensor signal corresponding to the pronunciation position and the pressure sensor corresponding to the pronunciation position; if there is someone on the seat corresponding to the pronunciation position, determining the target position where the user who emits the audio signal is located according to the pose information of the seat corresponding to the pronunciation position and the pronunciation position.
[0103] Optionally, when the determination module 72 is used to determine the target position where the user who emits the audio signal is located according to the pose information of the seat corresponding to the pronunciation position and the pronunciation position, it is specifically used for: determining the position where the buttocks of the user who emits the audio signal are located and the position where the head is located according to the pose information of the seat corresponding to the pronunciation position and the pronunciation position; if the position where the buttocks of the user are located and the position where the head is located are inconsistent, using the position where the buttocks of the user are located as the target position where the user who emits the audio signal is located; if the position where the buttocks of the user are located and the position where the head is located are consistent, using the position where the buttocks of the user are located or the position where the head is located as the target position where the user who emits the audio signal is located.
[0104] Optionally, the multi-screen interaction device further includes a receiving module 74, and the receiving module 74 is used for, if the position where the buttocks of the user are located and the position where the head is located are inconsistent, receiving the audio signal collected by the audio collection device corresponding to the position where the buttocks of the user are located, and receiving the audio signal collected by the audio collection device corresponding to the position where the head of the user is located.
[0105] Optionally, when the control module 73 is used to control the in-vehicle device corresponding to the target position to perform a voice interaction with the user, it is specifically used for: controlling the in-vehicle device corresponding to the target position to display a virtual image, and the virtual image is used to perform a voice interaction with the user.
[0106] Optionally, the control module 73 is further configured to: control an audio playback device around the vehicle-mounted device to play the voice broadcast by the virtual avatar.
[0107] Optionally, when the control module 73 is configured to control the vehicle-mounted device corresponding to the target location to perform a voice interaction with the user, it is further specifically configured to: if the vehicle-mounted device corresponding to the target location is in an open state, control the vehicle-mounted device corresponding to the target location to perform a voice interaction with the user.
[0108] Figure 7 The multi-screen interaction device in the illustrated embodiment can be used to execute the technical solutions of the above method embodiments, and its implementation principle and technical effects are similar, which will not be elaborated here.
[0109] Figure 8 This is a schematic structural diagram of a vehicle-mounted device provided by an embodiment of the present disclosure. The vehicle-mounted device provided by the embodiment of the present disclosure can execute the processing flow provided by the multi-screen interaction method embodiment, such as Figure 8 As shown, the vehicle-mounted device 80 includes: a memory 81, a processor 82, a computer program, and a communication interface 83; wherein, the computer program is stored in the memory 81 and is configured to be executed by the processor 82 to perform the multi-screen interaction method as described above.
[0110] In addition, an embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to implement the multi-screen interaction method described in the above embodiments.
[0111] It should be noted that the above-mentioned computer-readable medium in the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0112] In some embodiments, the client and the server may communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0113] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0114] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0115] The units involved in the embodiments described in the present disclosure may be implemented in software or in hardware. Wherein, the name of the unit does not constitute a limitation on the unit itself in some cases.
[0116] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and the like.
[0117] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0118] In addition, an embodiment of the present disclosure further provides a computer program product, which includes a computer program or instruction that, when executed by a processor, implements the battery remaining life estimation method as described above.
[0119] It should be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0120] The above are only specific embodiments of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to the embodiments described herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multi-screen interaction method, characterized in that, The method includes: Obtain an audio signal; Determine the target position where the user who emits the audio signal is located by identifying the audio signal; Control the in-vehicle device corresponding to the target position to perform a voice interaction with the user; Determining the target position where the user who emits the audio signal is located includes: Determine the position where the user's buttocks are located and the position where the user's head is located; If the position where the user's buttocks are located is different from the position where the user's head is located, then use the position where the user's buttocks are located as the target position where the user who emits the audio signal is located; If the position where the user's buttocks are located is the same as the position where the user's head is located, then use the position where the user's buttocks are located or the position where the user's head is located as the target position where the user who emits the audio signal is located.
2. The method according to claim 1, wherein Determining the target position where the user who emits the audio signal is located by identifying the audio signal includes: Determine the pronunciation position of the audio signal by identifying the audio signal; Determine the target position where the user who emits the audio signal is located according to the pose information of the seat corresponding to the pronunciation position and the pronunciation position.
3. The method according to claim 2, wherein The pose information of the seat corresponding to the pronunciation position is determined according to the seat back angle sensor and the seat front-back position sensor corresponding to the pronunciation position.
4. The method according to claim 2, wherein Determining the target position where the user who emits the audio signal is located according to the pose information of the seat corresponding to the pronunciation position and the pronunciation position includes: Determine whether there is someone on the seat corresponding to the pronunciation position according to at least one of the visual sensor signal corresponding to the pronunciation position and the pressure sensor corresponding to the pronunciation position; If there is someone on the seat corresponding to the pronunciation position, then determine the target position where the user who emits the audio signal is located according to the pose information of the seat corresponding to the pronunciation position and the pronunciation position.
5. The method according to claim 1, wherein The method further includes: If the position where the user's buttocks are located is different from the position where the user's head is located, then receive the audio signal collected by the audio collection device corresponding to the position where the user's buttocks are located, and receive the audio signal collected by the audio collection device corresponding to the position where the user's head is located.
6. The method according to claim 1, wherein Controlling the in-vehicle device corresponding to the target position to perform a voice interaction with the user includes: Control the in-vehicle device corresponding to the target position to display a virtual image, and the virtual image is used to perform a voice interaction with the user.
7. The method according to claim 6, characterized in that, The method further includes: Control the audio playback device corresponding to the in-vehicle device to play the voice broadcast by the virtual image.
8. The method according to claim 1, wherein Controlling the in-vehicle device corresponding to the target position to perform a voice interaction with the user further includes: If the in-vehicle device corresponding to the target position is in an open state, then control the in-vehicle device corresponding to the target position to perform a voice interaction with the user.
9. A multi-screen interaction device, characterized in that, The device includes: An acquisition module for acquiring an audio signal; A determination module for determining the target position where the user who emits the audio signal is located by identifying the audio signal; A control module for controlling the in-vehicle device corresponding to the target position to perform a voice interaction with the user; The determining module is specifically configured to determine the position of the user's hip and the position of the user's head where the audio signal is emitted; If the position of the user's hip and the position of the user's head are inconsistent, then the position of the user's hip is used as the target position where the user who emits the audio signal is located; If the position of the user's hip and the position of the user's head are consistent, then either the position of the user's hip or the position of the user's head is used as the target position where the user who emits the audio signal is located.
10. A vehicle-mounted device, characterized in that, Comprising A memory; A processor; And A computer program; Wherein, the computer program is stored in the memory and is configured to be executed by the processor to implement the method according to any one of claims 1-8.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1-8.
Citation Information
Patent Citations
Vehicle-mounted screen control method and system, vehicle and storage medium
CN110231866A
Voice interaction method and device, electronic equipment and storage medium
CN111694433A