Method and device for determining wake-up sound area, medium and equipment
By identifying the image in the vehicle cockpit, determining the driver and passenger information and sound area, and determining the target sound area with the wake-up command, the voice interaction problem caused by the fixed sound area of the vehicle is solved, and accurate voice control and interaction are achieved.
Patent Information
- Application Number
- CN202510749331.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-08-26
AI Technical Summary
In the prior art, the wake-up sound area of the vehicle is usually fixed, resulting in passengers in non- wake-up sound area such as passengers such as passengers cannot effectively wake up the voice assistant function of the vehicle and cannot realize human-vehicle interaction.
By acquiring the images in the vehicle cockpit, identifying the driver and passenger information and their respective corresponding sound areas, and determining the target sound areas in combination with the wake-up command to ensure the reception of the voice control command.
In the case of manually waking up the vehicle's voice assistant function, the wake-up sound area is accurately determined to ensure that the vehicle executes voice control instructions as expected by the user, and improves the voice interaction experience between the user and the vehicle.
Smart Images

Figure CN120544568A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of speech recognition technology, and in particular to a method, device, medium, and equipment for determining a wake-up sound zone. Background Art
[0002] With the development of vehicle technology, voice interaction has become more common in the process of human-vehicle interaction. Typically, when a user interacts with a vehicle through voice, the user needs to wake up the vehicle's voice assistant function. Users in the wake-up sound zone inside the vehicle can then issue voice commands to the vehicle.
[0003] Currently, users can activate the vehicle's voice assistant function using either a wake-up word or manual wake-up method. However, when using the manual wake-up method, the vehicle sets a fixed sound zone as the wake-up sound zone. For example, if the front passenger clicks the voice assistant icon on the vehicle screen to trigger the voice assistant function, the vehicle will not receive the front passenger's voice command when the front passenger issues a voice command, resulting in the failure to achieve the desired interaction effect. Therefore, when manually activating the vehicle's voice assistant function, how to accurately determine the vehicle's wake-up sound zone becomes an urgent problem that needs to be solved. Summary of the Invention
[0004] In order to solve the above technical problems, the present disclosure provides a method, device, medium and equipment for determining a wake-up sound zone to solve the problem that the expected human-vehicle interaction effect cannot be achieved due to setting a fixed sound zone.
[0005] In one aspect, a method for determining a wake-up sound zone is provided, comprising:
[0006] In response to receiving the wake-up command, acquiring an image of the vehicle cabin captured by an image sensor; wherein the image includes an image of the vehicle cabin at a time corresponding to the moment when the wake-up command is received and an image of the vehicle cabin within a preset time period before the wake-up command is received;
[0007] determining, based on the image, information of the driver and passenger in the vehicle and a first sound zone corresponding to each of the driver and passenger;
[0008] A target sound zone is determined according to the wake-up instruction, the driver and passenger information, and the first sound zone corresponding to each of the driver and passenger, so as to receive a voice control instruction of the driver and passenger corresponding to the target sound zone.
[0009] In another aspect, a device for determining a wake-up sound zone is provided, comprising:
[0010] a first acquisition module, configured to acquire, in response to receiving a wake-up command, an image of the vehicle cabin captured by an image sensor; wherein the image includes an image of the vehicle cabin at a time corresponding to the moment the wake-up command is received and an image of the vehicle cabin within a preset time period before the wake-up command is received;
[0011] a first determining module, configured to determine, based on the image, information of the driver and passenger in the vehicle and a first sound zone corresponding to each of the driver and passenger;
[0012] The second determining module is used to determine a target sound zone according to the wake-up instruction, the driver and passenger information, and the first sound zone corresponding to each of the driver and passenger, so as to receive a voice control instruction of the driver and passenger corresponding to the target sound zone.
[0013] On the other hand, a computer program product is proposed. When an instruction processor in the computer program product executes, the method for determining the wake-up sound zone proposed in the embodiment of the first aspect of the present disclosure is performed.
[0014] On the other hand, an electronic device is proposed, which includes: a processor; a memory for storing executable instructions of the processor; the processor is used to read the executable instructions from the memory and execute the instructions to implement the method for determining the wake-up sound zone described in the first aspect above.
[0015] The method for determining the wake-up sound zone provided by the embodiment of the present disclosure is that when a user wakes up the voice assistant function through a wake-up command, since the vehicle's driver and passenger information and the first sound zone corresponding to each driver and passenger can be determined by obtaining an image of the vehicle's cabin, the target sound zone can be determined based on the wake-up command, the driver and passenger information, and the first sound zone corresponding to each driver and passenger, so that the voice control command of the driver and passenger corresponding to the target sound zone can be received. Since, when the wake-up command is for the user to manually wake up the vehicle's voice assistant function, the target sound zone can be determined in the first sound zone corresponding to the driver and passenger who has performed the wake-up behavior by combining the wake-up command, the driver and passenger information, and the first sound zone corresponding to each driver and passenger, the wake-up sound zone can be accurately determined when the vehicle's voice assistant function is manually woken up, thereby ensuring that the vehicle executes the voice control command of the user in the wake-up sound zone as expected by the user, that is, improving the voice interaction experience between the user and the vehicle. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 3 is a schematic diagram of a scenario of a method for determining a wake-up sound zone provided by an exemplary embodiment of the present disclosure.
[0017] Figure 2 4 is a flowchart of a method for determining a wake-up sound zone provided by an exemplary embodiment of the present disclosure.
[0018] Figure 3 3 is a flowchart of a method for determining a wake-up sound zone provided by another exemplary embodiment of the present disclosure.
[0019] Figure 4 Still another exemplary embodiment of the present disclosure provides a flowchart of a method for determining a wake-up sound zone.
[0020] Figure 5 A schematic structural diagram of a device for determining a wake-up sound zone provided by yet another exemplary embodiment of the present disclosure.
[0021] Figure 6 is a structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0022] To explain the present disclosure, example embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments. It should be understood that the present disclosure is not limited to the example embodiments.
[0023] It should be noted that the relative arrangement of components and steps, the numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present disclosure unless specifically stated otherwise.
[0024] Application Overview
[0025] In related technologies, after waking up the vehicle with voice control, if the user wants to use the vehicle's voice assistant function, they must first determine the vehicle's wake-up sound zone. After determining the wake-up sound zone, the driver or passenger in the wake-up sound zone can issue voice commands to the vehicle to achieve human-vehicle interaction.
[0026] However, when using touch to activate the vehicle's voice assistant function, the vehicle's wake-up sound zone is typically set to a fixed zone. For example, if the driver's seat is set as the wake-up sound zone, and the front passenger clicks the voice assistant icon on the vehicle screen to trigger the voice assistant function, if the front passenger wants to issue a voice command to the vehicle, the front passenger's position is not in the wake-up sound zone, resulting in the vehicle being unable to receive the front passenger's voice command, and thus, human-vehicle interaction cannot be achieved. Therefore, when manually activating the vehicle's voice assistant function, how to accurately determine the vehicle's wake-up sound zone becomes an urgent problem that needs to be solved.
[0027] Based on the above technical problems, the method for determining the wake-up sound zone provided in the embodiment of the present disclosure is that when the user wakes up the voice assistant function through a wake-up command, since the vehicle's driver and passenger information and the first sound zone corresponding to each driver and passenger can be determined by obtaining an image inside the vehicle cabin, the target sound zone can be determined based on the wake-up command, the driver and passenger information, and the first sound zone corresponding to each driver and passenger, so that the voice control command of the driver and passenger corresponding to the target sound zone can be received. Since the target sound zone can be determined in the first sound zone corresponding to the driver and passenger who has the wake-up behavior by combining the wake-up command, the driver and passenger information, and the first sound zone corresponding to each driver and passenger, the wake-up sound zone can be accurately determined when the voice assistant function of the vehicle is manually woken up, thereby ensuring that the vehicle executes the voice control command of the user in the wake-up sound zone as expected by the user, that is, improving the voice interaction experience between the user and the vehicle.
[0028] Exemplary Systems
[0029] Figure 1 FIG2 is a schematic diagram of a scenario of a method for determining a wake-up sound zone provided by an exemplary embodiment of the present disclosure. The method can be executed by a vehicle, which includes a vehicle screen 01 and a camera 02.
[0030] For example, Figure 1 As shown, camera 02 continuously captures images of the vehicle cabin in real time. When the front passenger passenger needs to interact with the vehicle via voice, they can click the voice assistant control 03 on the vehicle screen 01. In response to this click, the vehicle captures images of the vehicle cabin at the time of the click, as well as images of the vehicle cabin within 3 seconds prior to receiving the click. Based on these images, it can be determined that the driver and passenger information indicates that the front passenger passenger has performed a touch operation on the vehicle screen 01. The first audio zone corresponding to the front passenger passenger is the front passenger seat, and the first audio zone corresponding to the driver is the driver seat. Based on the wake-up command being a touch operation, the driver and passenger information, and the first audio zones corresponding to each driver and passenger, the target audio zone is determined to be the front passenger seat. The vehicle can then receive voice control commands from the driver and passenger in the front passenger seat, thereby enabling voice interaction between the vehicle and the user when the vehicle's voice assistant function is manually activated.
[0031] The method for determining a wake-up sound zone provided by the embodiments of the present disclosure is such that when a user wakes up the voice assistant function through a wake-up command, the vehicle's driver and passenger information and the first sound zone corresponding to each driver and passenger can be determined by obtaining an image of the vehicle's cabin. Therefore, the target sound zone can be determined based on the wake-up command, the driver and passenger information, and the first sound zone corresponding to each driver and passenger, thereby receiving the voice control command of the driver and passenger corresponding to the target sound zone. Since the target sound zone can be determined in the first sound zone corresponding to the driver and passenger who has performed the wake-up action by combining the wake-up command, the driver and passenger information, and the first sound zone corresponding to each driver and passenger, the wake-up sound zone can be accurately determined when the vehicle's voice assistant function is manually awakened, thereby ensuring that the vehicle executes the voice control command of the user in the wake-up sound zone as expected by the user, thereby improving the voice interaction experience between the user and the vehicle.
[0032] Exemplary Methods
[0033] Figure 2 4 is a flowchart of a method for determining a wake-up sound zone provided by an exemplary embodiment of the present disclosure.
[0034] For example, the above method can be executed by a vehicle, or by a chip on a vehicle. Figure 2 As shown, the above method may include the following steps:
[0035] Step 201 : In response to receiving a wake-up instruction, an image of the interior of a vehicle cabin captured by an image sensor is acquired.
[0036] The images include images of the vehicle cabin at the time corresponding to when the wake-up command is received and images of the vehicle cabin within a preset time period before the wake-up command is received.
[0037] In some embodiments, the wake-up command is a command for waking up the voice assistant function of the vehicle. The wake-up command can be a user touch operation on the voice assistant area on the vehicle display screen, or a user gesture command when the user is detected to be looking at the voice assistant area on the vehicle display screen.
[0038] In some embodiments, an image sensor can be used to capture at least one image in real time from within the vehicle cabin. The image sensor can include at least one camera located within the vehicle cabin, which can be positioned at various locations within the vehicle. For example, the at least one camera can include a camera mounted above the vehicle's display screen, or a surround-view camera mounted in the center of the vehicle's roof.
[0039] Exemplarily, after the vehicle is powered on, at least one camera in the vehicle continuously captures the interior of the vehicle cabin at 30 fps to obtain at least one image.
[0040] In other embodiments, the at least one image may be an image captured by the same camera within a preset time period, or may be an image captured by multiple cameras at the same time (for example, the time corresponding to the wake-up instruction).
[0041] In some examples, images of the vehicle cabin captured by the image sensor at the time corresponding to the wake-up instruction can be obtained in real time, and images of the vehicle cabin captured by the image sensor within a preset time period before the wake-up instruction can be obtained from the storage space.
[0042] For example, assuming that the wake-up command corresponds to time t1, and a time before time t1 is time t2, upon receiving the wake-up command, in response to the received wake-up command, the image of the vehicle cabin captured in real time at time t1 can be directly obtained, and the image of the vehicle cabin captured during the time period t1-t2 can be obtained from the preset storage space.
[0043] In order to use the voice assistant function of the vehicle, it is necessary to determine the wake-up sound zone of the vehicle. Therefore, through the method for determining the wake-up sound zone provided in the embodiment of the present invention, if the vehicle receives a wake-up command from the user, it can respond to the wake-up command, obtain the image of the vehicle cabin at the time corresponding to the wake-up command and the image of the vehicle cabin within a preset time period before receiving the wake-up command, and determine the wake-up sound zone of the vehicle based on these images.
[0044] Step 202 : Determine the driver and passenger information in the vehicle and the first sound zone corresponding to each of the driver and passenger based on the image.
[0045] In some embodiments, the above-mentioned driver and passenger information may include driver and passenger quantity information and target state information; wherein the target state information may include: driver and passenger behavior state information and sight state information and touch position.
[0046] In some examples, an image recognition algorithm is used to identify key human feature information such as human skeleton points, facial areas, hand areas, and eye areas in the image to obtain information on the number of drivers and passengers; and the human feature information is combined with a temporal behavior estimation algorithm and the system touch feedback results to obtain the behavior status information of each driver and passenger. For example, the behavior status information may be whether the vehicle screen is touched; the line of sight state information of each driver and passenger is determined based on the facial feature information combined with the line of sight estimation algorithm; and the touch position or gesture information can be determined based on the hand feature information combined with the hand reconstruction model.
[0047] In some embodiments, based on the distribution of different physical spaces inside the vehicle, multiple sound zones included in the vehicle may be determined.
[0048] For example, the distribution of object space in the vehicle includes: the main driver's seat, the front passenger seat, the second row left position, the second row middle position and the second row right position, that is, the multiple sound zones in the vehicle include: the main driver's seat, the front passenger seat, the second row left position, the second row middle position and the second row right position.
[0049] In some embodiments, the first audio zone can be used to indicate the physical space corresponding to each occupant in the vehicle. After acquiring an image of the vehicle cabin, an image recognition algorithm can be used to identify the image and determine the location of each occupant in the vehicle. Each occupant's location in the vehicle is then matched with multiple audio zones in the vehicle, and the audio zone that matches each occupant's location in the vehicle is determined as the first audio zone corresponding to each occupant.
[0050] For example, assuming a vehicle includes a driver and a front passenger, after acquiring an image of the vehicle cabin, the image can be fed into a trained deep learning model to output the distribution of the driver and passenger positions in the vehicle as the driver's seat and the front passenger's seat. Since the driver's position matches the driver's seat among the vehicle's multiple sound zones, the first sound zone corresponding to the driver is the driver's seat, and similarly, the first sound zone corresponding to the front passenger is the front passenger's seat.
[0051] Step 203 : determining a target sound zone according to the wake-up instruction, the driver and passenger information, and the first sound zones corresponding to the driver and passenger, so as to receive a voice control instruction from the driver and passenger corresponding to the target sound zone.
[0052] In the embodiment of the present disclosure, by combining the wake-up command, driver and passenger information, and the first sound zone corresponding to each driver and passenger, the first sound zone corresponding to a certain driver and passenger can be determined as the target sound zone, that is, the target sound zone is the wake-up sound zone, so that after the driver and passenger located in the target sound zone issues a voice control command, the vehicle's voice system can receive the voice control command and perform corresponding operations in response to the voice control command, thereby realizing the human-vehicle interaction function.
[0053] Furthermore, since the target sound zone is the wake-up sound zone, when a driver or passenger in another sound zone issues a voice control command, the voice system cannot respond to the voice control command, thereby preventing voice interference from drivers or passengers in other sound zones.
[0054] The method for determining a wake-up sound zone provided by the embodiments of the present disclosure is such that when a user wakes up the voice assistant function through a wake-up command, the vehicle's driver and passenger information and the first sound zone corresponding to each driver and passenger can be determined by obtaining an image of the vehicle's cabin. Therefore, the target sound zone can be determined based on the wake-up command, the driver and passenger information, and the first sound zone corresponding to each driver and passenger, thereby receiving the voice control command of the driver and passenger corresponding to the target sound zone. Since the target sound zone can be determined in the first sound zone corresponding to the driver and passenger who has performed the wake-up action by combining the wake-up command, the driver and passenger information, and the first sound zone corresponding to each driver and passenger, the wake-up sound zone can be accurately determined when the vehicle's voice assistant function is manually awakened, thereby ensuring that the vehicle executes the voice control command of the user in the wake-up sound zone as expected by the user, thereby improving the voice interaction experience between the user and the vehicle.
[0055] like Figure 3 As shown in the above Figure 2 Based on the illustrated embodiment, step 203 may include the following step 2031 or 2032:
[0056] Step 2031 : In response to the wake-up instruction being a touch operation and the number of drivers and passengers in the driver and passenger information being one, the first sound zone corresponding to the driver and passenger is determined as the target sound zone.
[0057] In some embodiments, when the wake-up command is a command to touch the vehicle screen, if there is only one driver and passenger in the vehicle, that is, when this driver and passenger touches the vehicle screen to wake up the vehicle's voice assistant function, the first sound zone corresponding to this driver and passenger is determined as the target sound zone.
[0058] In one example, assuming that there is only the main driver in the vehicle, when the main driver triggers the vehicle's voice assistant function through a touch operation, since the first sound zone corresponding to the main driver is the main driver's position, the main driver's position can be determined as the target sound zone, so that only the user in the main driver's position can initiate voice control commands to the vehicle's voice system to trigger the execution of corresponding operations.
[0059] In another example, assuming that there is only a front passenger passenger in the vehicle, when the front passenger passenger triggers the vehicle's voice assistant function through a touch operation, since the first sound zone corresponding to the front passenger passenger is the front passenger position, the front passenger position can be determined as the target sound zone, so that only the user in the front passenger position can initiate voice control commands to the vehicle voice system to trigger the execution of corresponding operations.
[0060] Step 2032 : In response to the wake-up instruction being a touch operation and the number of drivers and passengers being multiple, determine a target sound zone based on the target state information in the driver and passenger information.
[0061] In some embodiments, the target state information includes the sight state information or behavior state information of the driver or passenger. For different target state information, different methods can be used to determine the target sound zone based on the target state information of each driver or passenger.
[0062] In other embodiments, if the wake-up command is a touch operation, if the vehicle screen being touched is in a defined sound zone, the sound zone corresponding to the vehicle screen is directly determined as the target sound zone. A defined sound zone means that only passengers in that sound zone can touch the vehicle screen, while passengers in other sound zones cannot touch the vehicle screen. It should be noted that it is possible to pre-determine whether a vehicle screen is in a defined sound zone or an undefined sound zone. If the vehicle screen is in an undefined sound zone, step 2031 or 2032 is executed to determine the target sound zone.
[0063] For example, if the vehicle screen is the front passenger screen, it can only be touched by the front passenger and not by the driver. Another example is the rear seat screen or the screen under the armrest, which cannot be touched by the front passenger. Since these vehicle screens are all in a specific audio zone, the audio zone corresponding to the vehicle screen can be determined as the target audio zone. In other words, the audio zone corresponding to the vehicle screen is activated to receive and recognize voice control commands from the driver or passenger in that audio zone.
[0064] The method for determining a wake-up sound zone provided in an embodiment of the present disclosure, when the wake-up instruction is a touch operation, determines the first sound zone corresponding to a single occupant. Alternatively, if there are multiple occupants, the target sound zone can be determined based on the target state information of the multiple occupants. Because the target sound zone can be accurately determined by combining the number of occupants and the target state information when the wake-up instruction is a touch operation, the vehicle can be guaranteed to receive voice control commands issued by the occupant in the target sound zone, thereby enabling voice interaction between the user and the vehicle when the user wakes up the vehicle's voice function through a touch operation.
[0065] like Figure 4 As shown in the above Figure 3 Based on the illustrated embodiment, step 2032 may include the following steps:
[0066] Step 2032a: Determine the number of target drivers and passengers corresponding to the touch operation based on the operation information in the target state information.
[0067] In some embodiments, the operation information can be used to indicate whether a driver or passenger has touched the vehicle screen. Based on the operation information of each driver or passenger, all drivers or passengers who have touched the vehicle screen are screened out, and the screened out drivers or passengers are determined as target drivers or passengers, and the number of these drivers or passengers is used as the number of target drivers or passengers.
[0068] In some embodiments, the number of the target driver and passenger is at least one. When the number of the target driver and passenger is one, the first sound zone corresponding to the target driver and passenger is determined as the target sound zone; when the number of the target driver and passenger is at least two, the following step 2032b may be performed.
[0069] Step 2032b: In response to the number of target drivers and passengers being at least two, determine a target sound zone according to target state information of the at least two target drivers and passengers and their respective corresponding first sound zones.
[0070] In some embodiments, when there are at least two target occupants, that is, when multiple occupants touch the vehicle screen before the voice assistant function is activated, a first audio zone corresponding to each occupant can be selected as the target audio zone based on the time information, line of sight information, or behavioral information of the touch operations. For details, please refer to the detailed description of the following embodiments, and the present disclosed embodiments will not be repeated here.
[0071] In some embodiments, when the number of target drivers and passengers is at least two, determining the target sound zone based on the target status information of at least two target drivers and passengers and their respective corresponding first sound zones in the above-mentioned step 2032b may specifically include: determining the time sequence in which each target driver and passenger performs the touch operation; and determining the target sound zone in the first sound zone corresponding to each target driver and passenger based on the time sequence in which each target driver and passenger performs the touch operation.
[0072] In some examples, the time information of each target driver and passenger performing a touch operation can be determined first, and then the time information can be sorted according to a preset order to obtain the time sequence of each target driver and passenger performing the touch operation; and then, based on the time sequence of each target driver and passenger performing the touch operation, the first sound zone corresponding to a specific target driver and passenger (for example, the target driver and passenger with the earliest or latest touch operation time) is selected as the target sound zone.
[0073] Exemplarily, assume that there is a driver and a co-driver in the vehicle, and both the driver and the co-driver perform touch operations on the vehicle screen. The first sound zone corresponding to the driver is the driver's position, and the first sound zone corresponding to the co-driver is the co-driver's position. Determine that the time information of the touch operation performed by the driver is t1, and the time information of the touch operation performed by the co-driver is t2. Sort t1 and t2 in ascending order, and the time order of the touch operations performed by the driver and the co-driver is obtained as t2 < t1. Then, according to this time order, select the co-driver who performs the touch operation at time t2, that is, use the co-driver's position corresponding to the co-driver who is closest to the current time when performing the touch operation as the target sound zone.
[0074] In some embodiments, when the number of target occupants is at least two, determining the target sound zone according to the target status information of at least two target occupants and their respective corresponding first sound zones in step 2032b may specifically include: based on the line-of-sight information in the target status information, determining the first matching relationship between the line of sight of each target occupant and the screen wake-up area of the vehicle; and determining the target sound zone based on the first matching relationship.
[0075] In some examples, the line-of-sight information of each of the above target occupants may include the eye position and the line-of-sight direction; the screen wake-up area is the area on the vehicle screen for waking up the vehicle's voice assistant function, and can also be referred to as the voice assistant area. Based on the eye position and the line-of-sight direction of each target occupant, the line-of-sight fixation area of each target occupant on the vehicle screen can be calculated; then, based on whether the line-of-sight fixation area of each target occupant is the screen wake-up area, the first matching relationship between the line of sight of each target occupant and the screen wake-up area of the vehicle is determined, and further the target area can be determined based on the first matching relationship.
[0076] In some examples, the above first matching relationship includes two cases: (1) the line of sight of at least one first occupant falls within the screen wake-up area of the vehicle; (2) the lines of sight of each target occupant do not fall within the screen wake-up area of the vehicle. When the line-of-sight fixation area of at least one first occupant among at least two target occupants is the screen wake-up area, determine that the first matching relationship is that the line of sight of the at least one first occupant falls within the screen wake-up area, or when the line-of-sight fixation areas of each target occupant are not the screen wake-up area of the vehicle, determine that the first matching relationship is that the lines of sight of each target occupant do not fall within the screen wake-up area of the vehicle.
[0077] In some examples, in response to a first matching relationship where the gaze of at least one first occupant falls within the screen wake-up area, a target sound zone is determined based on the first sound zone corresponding to the at least one first occupant. If there is only one at least one first occupant, the first sound zone corresponding to the only first occupant is determined as the target sound zone. Alternatively, if there are multiple at least one first occupants, the first sound zone corresponding to a specific first occupant among the multiple first occupants can be selected as the target sound zone. For example, the first sound zone corresponding to the first occupant who performed a touch operation most recently among the multiple first occupants can be determined as the target sound zone.
[0078] For example, assuming that the main driver and the front passenger in the vehicle both perform touch operations on the vehicle screen, the first sound zone corresponding to the main driver is the main driver position, and the first sound zone corresponding to the front passenger is the front passenger position; based on the main driver's eye position and line of sight direction, the line of sight gaze area 1 on the vehicle screen is determined. Since the line of sight gaze area 1 is the screen wake-up area on the vehicle screen, the main driver's line of sight falls within the screen wake-up area; based on the front passenger's eye position and line of sight direction, the line of sight gaze area 2 on the vehicle screen is determined. Since the line of sight gaze area 2 is not located in the screen wake-up area on the vehicle screen, the front passenger's line of sight does not fall within the screen wake-up area; in response to the first matching relationship that the main driver's line of sight falls within the screen wake-up area, the main driver's position corresponding to the main driver is determined as the target sound zone.
[0079] In some embodiments, when the number of target drivers and passengers is at least two, determining the target sound zone based on the target status information of at least two target drivers and passengers and their respective corresponding first sound zones in the above-mentioned step 2032b may specifically include: determining the second matching relationship between the touch position of each target driver and passenger and the screen wake-up area of the vehicle based on the touch position in the target status information; and determining the target sound zone based on the second matching relationship.
[0080] In some examples, when determining the touch position of each target driver or passenger, an image recognition algorithm can be used to identify the feature information of the human hand area in the image, and the identified human hand feature information can be combined with the human hand reconstruction model to obtain the three-dimensional coordinates of multiple hand key points; the three-dimensional coordinates of each hand key point are converted from the world coordinate system to the screen coordinate system, and then based on the converted coordinates of each hand key point, multiple contact points closest to the screen surface are determined. Then, combined with the pressure data detected by the tactile sensor and the system touch feedback results, the contact point corresponding to the maximum pressure is selected from the multiple contact points as the touch position. For details, please refer to the detailed description in the relevant technology, and the embodiments of the present disclosure will not be repeated here.
[0081] In some examples, based on the touch positions of at least two target occupants, the touch position of each target occupant is compared with the screen wake-up area to determine whether the touch position corresponds to the screen wake-up area, and a second matching relationship is determined based on the determination result. The determination result may be that the touch position corresponds to the screen wake-up area, or that the touch position does not correspond to the screen wake-up area.
[0082] In this way, when the judgment result is that the touch position of at least one target driver and passenger corresponds to the screen wake-up area, the second matching relationship is that the touch position of at least one target driver and passenger matches the screen wake-up area of the vehicle, or, when the judgment result is that the touch position of each target driver and passenger does not correspond to the screen wake-up area, the second matching relationship is that the touch position of each target driver and passenger does not match the screen wake-up area of the vehicle.
[0083] In the embodiment of the present disclosure, the second matching relationship between the touch position of each target driver or passenger and the screen wake-up area of the vehicle based on the touch position in the target state information may be determined in the following two possible implementations:
[0084] A possible implementation
[0085] In some embodiments, the above-mentioned step of determining the second matching relationship between the touch position of each target driver and passenger and the screen wake-up area of the vehicle based on the touch position in the target status information may specifically include: based on the touch position of each target driver and passenger, performing image cropping processing on the image corresponding to each target driver and passenger to obtain a target image area containing the touch position of each target driver and passenger; based on the target image area corresponding to each target driver and passenger, determining the second matching relationship between the touch position of each target driver and passenger and the screen wake-up area of the vehicle.
[0086] In some examples, the image frames corresponding to the target drivers and passengers are multiple image frames containing the actions of the target drivers and passengers touching the vehicle screen. Multiple image frames containing the actions of the target drivers and passengers touching the vehicle screen can be first acquired, and then the multiple image frames are subjected to cutout processing based on the touch positions of the target drivers and passengers, thereby obtaining multiple target image regions containing the touch positions of the target drivers and passengers. The multiple target image regions are then processed using an image classification algorithm to output whether each target image region is a screen wake-up region. In other words, a second matching relationship between the touch positions of the target drivers and passengers and the vehicle's screen wake-up region can be determined based on the output results.
[0087] In some examples, when the output result is that at least one target image area corresponding to at least one target driver and passenger is the screen wake-up area, that is, the touch position of at least one target driver and passenger corresponds to the screen wake-up area, that is, the second matching relationship is that the touch position of at least one target driver and passenger matches the screen wake-up area of the vehicle, that is, the touch position of at least one target driver and passenger is within the screen wake-up area of the vehicle; when the output result is that each target image area corresponding to each target driver and passenger is not the screen wake-up area, that is, the touch position of each target driver and passenger does not correspond to the screen wake-up area, that is, the second matching relationship is that the touch position of each target driver and passenger does not match the screen wake-up area of the vehicle, that is, the touch position of each target driver and passenger is not within the screen wake-up area of the vehicle.
[0088] Another possible implementation
[0089] In other embodiments, the above-mentioned step of determining the second matching relationship between the touch position of each target driver and passenger and the wake-up area of the vehicle screen based on the touch position in the target status information may specifically include: determining the target position of the wake-up area of the vehicle screen on the target screen; and determining the second matching relationship between the touch position of each target driver and passenger and the wake-up area based on the degree of position matching between the touch position of each target driver and passenger and the target position.
[0090] In some examples, since the screen wake-up area is a preset touch area on the vehicle screen, the coordinate information of the target position of the screen wake-up area on the target screen is known, so the coordinate information of the target position of the screen wake-up area on the target screen can be directly obtained from the memory. The coordinate information of the target position of the screen wake-up area on the target screen can be represented by a coordinate range. For example, if the screen wake-up area is a rectangular area, the coordinate information of its target position on the target screen can be represented as [(x1, y1), (x2, y2)].
[0091] In some embodiments, the coordinate information of the touch position of each target driver and passenger is calculated, and then it is determined whether the coordinates corresponding to the touch position of each target driver and passenger are located within the coordinate range corresponding to the screen wake-up area on the target screen; when the coordinates corresponding to the touch position of the target driver and passenger are located within the coordinate range corresponding to the screen wake-up area on the target screen (including the boundary coordinate points), it is determined that the touch position of the target driver and passenger matches the target position, or, when the coordinates corresponding to the touch position of the target driver and passenger are not located within the coordinate range corresponding to the screen wake-up area on the target screen, it is determined that the touch position of the target driver and passenger does not match the target position.
[0092] When judging whether the coordinates corresponding to the touch position of each target driver and passenger are within the coordinate range corresponding to the screen wake-up area on the target screen, it is necessary to calculate the degree of position matching between the touch position of the target driver and passenger and the target position; an error threshold can be set in advance, and then by comparing the error between the touch position of the target driver and passenger and the target position and the relationship between the error and the threshold, it is determined whether the coordinates corresponding to the touch position of each target driver and passenger are within the coordinate range corresponding to the screen wake-up area on the target screen.
[0093] In some examples, when the error between the touch position of the target driver and the target position is less than or equal to the error threshold, it can be determined that the coordinates corresponding to the touch position of the target driver and the target position are within the coordinate range corresponding to the screen wake-up area on the target screen, that is, the touch position of the target driver and the target position corresponds to the screen wake-up area, and at this time the second matching relationship is that the touch position of the target driver and the screen wake-up area matches; when the error between the touch position of the target driver and the target position is greater than the error threshold, it can be determined that the coordinates corresponding to the touch position of the target driver and the target position are not within the coordinate range corresponding to the screen wake-up area on the target screen, that is, the touch position of the target driver and the target position does not correspond to the screen wake-up area, and at this time the second matching relationship is that the touch position of the target driver and the screen wake-up area does not match.
[0094] In some embodiments, in response to the second matching relationship that the touch position of the second driver and passenger among at least two target drivers and passengers is matched with the screen wake-up area, that is, when the touch position of the second driver and passenger is located in the screen wake-up area of the vehicle screen, the first sound zone corresponding to the second driver and passenger is determined as the target sound zone.
[0095] In some embodiments, the number of the second occupants may be one or more. When there is only one second occupant, the first sound zone corresponding to the second occupant may be determined as the target sound zone. When there are multiple second occupants, the first sound zone corresponding to any one of the multiple occupants may be determined as the target sound zone, or the first sound zone corresponding to a specific occupant among the multiple occupants may be determined as the target sound zone.
[0096] For example, when the second occupant includes multiple occupants, a specific occupant can be determined from among the multiple occupants based on the time information of each occupant touching the voice assistant area. For example, the occupant who touches the voice assistant area earliest among the multiple occupants can be determined as the specific occupant, that is, the first sound zone corresponding to the occupant who touches the voice assistant area earliest among the multiple occupants can be determined as the target sound zone. For another example, the occupant who touches the voice assistant area latest among the multiple occupants can be determined as the specific occupant, that is, the first sound zone corresponding to the occupant who touches the voice assistant area latest among the multiple occupants can be determined as the target sound zone.
[0097] In other embodiments, if a target sound zone cannot be determined based on the wake-up command, driver and passenger information, and the first sound zones corresponding to the respective driver and passenger, the sound source location corresponding to the voice wake-up command is determined, and the sound zone corresponding to the sound source location is determined as the target sound zone. The voice wake-up command is used to activate the voice assistant function, and the voice wake-up command is a command received by the vehicle before the wake-up command. Thus, the sound zone corresponding to the sound source location can be determined based on the sound source location corresponding to the voice wake-up command.
[0098] The method for determining the wake-up sound zone provided by the embodiment of the present disclosure, when there are multiple drivers and passengers, can determine the number of target drivers and passengers corresponding to the touch operation based on the operation information in the target state information. Therefore, when the number of target drivers and passengers is at least two, that is, when multiple drivers and passengers touch the vehicle screen before the voice assistant function is activated, the touch operation time sequence, touch position and line of sight information of the multiple drivers and passengers can be combined to accurately determine the target sound zone in the first sound zone corresponding to each of the multiple drivers and passengers, thereby ensuring that when multiple drivers and passengers touch the vehicle screen, the drivers and passengers in the target sound zone will not be disturbed by the voices of drivers and passengers in other sound zones when they interact with the vehicle by voice.
[0099] In some embodiments, after receiving the wake-up command in step 201, the following steps may further include: determining the line of sight of each driver or passenger based on the image; determining a third matching relationship between the line of sight of each driver or passenger and a preset wake-up area on the vehicle screen; and waking up the vehicle's voice assistant function in response to the third matching relationship being that the line of sight of the third driver or passenger falls within the preset wake-up area and in response to a gesture command of the third driver or passenger matching the preset command. The wake-up command includes a gesture command.
[0100] In some examples, by performing feature recognition on the image, the facial key point features of each driver and passenger are obtained, and the line of sight of each driver and passenger is obtained by using a line of sight estimation algorithm based on the facial key point features of each driver and passenger; the gaze area of each driver and passenger on the vehicle screen is determined based on the line of sight and eye position of each driver and passenger, and a third matching relationship is determined based on the matching relationship between the gaze area and the preset wake-up area.
[0101] In some embodiments, when the gaze area of at least one driver and passenger is located in the preset wake-up area, the third matching relationship is that the line of sight of the at least one driver and passenger falls within the preset wake-up area, or, when the gaze area of each driver and passenger is not located in the preset wake-up area, the third matching relationship is that the line of sight of each driver and passenger does not fall within the preset wake-up area.
[0102] In some examples, when only one driver and passenger's line of sight falls within the preset wake-up area, the third driver and passenger is that one driver and passenger; when multiple drivers and passengers' line of sight falls within the preset wake-up area, the third driver and passenger is the driver and passenger among the multiple passengers whose line of sight first falls within the preset wake-up area.
[0103] In some examples, the preset instruction is a preset gesture instruction for triggering the voice assistant wake-up function. For example, the preset instruction can be a pinch gesture instruction, an OK gesture instruction, or any other gesture instruction.
[0104] For example, assume a vehicle includes a driver, passenger A, and passenger B. After acquiring an image of the vehicle cabin, image recognition is performed on the image to determine that the driver's line of sight is within gaze area A of the vehicle screen, passenger A's line of sight is within gaze area B of the vehicle screen, and passenger A's line of sight is within gaze area C of the vehicle screen. Since neither gaze area A nor gaze area B is within the preset wake-up area of the vehicle screen, the line of sight of the driver and passenger A falls outside the preset wake-up area. However, since gaze area C is within the preset wake-up area of the vehicle screen, passenger B's line of sight falls within the preset wake-up area, and passenger B's pinch gesture command matches the preset command. As a result, the vehicle's voice assistant function is awakened by passenger B's line of sight and the specific gesture command.
[0105] In other embodiments, the above-mentioned preset area is the display area corresponding to the voice assistant control; when it is determined that the line of sight of the third driver and passenger falls within the preset wake-up area, the display method of the voice assistant control on the vehicle screen can be adjusted; and then, in response to the gesture command of the third driver and passenger matching the preset command, the voice assistant function of the vehicle is awakened.
[0106] For example, when it is determined through the acquired image that the third driver's line of sight falls within the preset wake-up area, the voice assistant control can be displayed in an enlarged manner, or the display color of the voice assistant control can be updated.
[0107] In some embodiments, when the voice assistant function of the vehicle is awakened by the sight and gesture instructions of the third driver and passenger, the first sound zone corresponding to the third driver and passenger is determined as the target sound zone.
[0108] The method for determining the wake-up sound zone provided by the embodiment of the present disclosure responds to the received wake-up command. Since the line of sight of each driver and passenger can be determined based on the image, and a third matching relationship between the line of sight of each driver and passenger and the preset wake-up area in the vehicle screen is determined, when the third matching relationship is that the line of sight of the third driver and passenger falls within the preset wake-up area and the finger command of the third driver and passenger matches the preset command, the voice assistant function of the vehicle is awakened, so that the user does not need to touch the vehicle screen, and the voice assistant function can be awakened only by the driver and passenger's line of sight and gesture command, thereby achieving the effect of the driver and passenger remotely operating the vehicle screen for voice wake-up.
[0109] Exemplary devices
[0110] Figure 5 This is a schematic diagram of a device for determining a wake-up area according to an exemplary embodiment of the present disclosure. The device can be installed in an electronic device such as a terminal device or a server, or in an object such as a vehicle, to perform the wake-up area determination method according to any of the above embodiments of the present disclosure.
[0111] like Figure 5 As shown, the above-mentioned wake-up sound zone determination device 300 may include:
[0112] The first acquisition module 301 may be configured to acquire, in response to receiving a wake-up command, an image of the vehicle cabin captured by an image sensor; wherein the image includes an image of the vehicle cabin at the time corresponding to the moment when the wake-up command is received and an image of the vehicle cabin within a preset time period before the wake-up command is received;
[0113] A first determining module 302 may be configured to determine information of the vehicle's drivers and passengers and first sound zones corresponding to the drivers and passengers based on the image;
[0114] The second determining module 303 may be configured to determine a target sound zone according to the wake-up instruction, the driver and passenger information, and the first sound zones corresponding to the driver and passenger, so as to receive a voice control instruction from the driver and passenger corresponding to the target sound zone.
[0115] In a possible implementation, the second determining module 303 may be specifically configured to, in response to the wake-up instruction being a touch operation and the number of drivers and passengers in the driver and passenger information being one, determine the first sound zone corresponding to the driver and passenger as the target sound zone; or
[0116] In response to the wake-up instruction being a touch operation and the number of the drivers and passengers being multiple, the target sound zone is determined based on target state information in the driver and passenger information.
[0117] In a possible implementation, the second determining module 303 may be specifically configured to determine the number of target drivers and passengers corresponding to the touch operation based on the operation information in the target state information;
[0118] In response to the number of the target drivers and passengers being at least two, the target sound zone is determined according to target state information of at least two target drivers and passengers and the first sound zones corresponding to the two target drivers and passengers.
[0119] In a possible implementation, the second determining module 303 may be specifically configured to determine a time sequence in which each target driver or passenger performs a touch operation;
[0120] The target sound zone is determined in the first sound zone corresponding to each target driver or passenger according to a time sequence in which each target driver or passenger performs a touch operation.
[0121] In a possible implementation, the second determining module 303 may be specifically configured to determine a first matching relationship between the sight lines of each target driver and passenger and the screen wake-up area of the vehicle based on the sight line information in the target state information;
[0122] The target vocal range is determined based on the first matching relationship.
[0123] In a possible implementation, the second determining module 303 may be specifically configured to determine a second matching relationship between the touch position of each target driver or passenger and the screen wake-up area of the vehicle based on the touch position in the target state information;
[0124] The target vocal range is determined based on the second matching relationship.
[0125] In one possible implementation, the second determining module 303 may be specifically configured to perform image cropping processing on the image corresponding to each target driver or occupant based on the touch position of each target driver or occupant, to obtain a target image region containing the touch position of each target driver or occupant;
[0126] Based on the target image area corresponding to each target driver or passenger, the second matching relationship between the touch position of each target driver or passenger and the screen wake-up area of the vehicle is determined.
[0127] In a possible implementation, the second determining module 303 may be specifically configured to determine a target position of the screen wake-up area of the vehicle on the target screen;
[0128] Based on the degree of position matching between the touch position of each target driver and passenger and the target position, the second matching relationship between the touch position of each target driver and passenger and the screen wake-up area is determined.
[0129] In a possible implementation, the determining device 300 may further include:
[0130] A third determining module may be configured to determine the sight lines of the drivers and passengers based on the image;
[0131] a fourth determining module, which may be used to determine a third matching relationship between the sight lines of each of the driver and passenger and a preset wake-up area on the vehicle screen;
[0132] The voice wake-up module can be used to wake up the voice assistant function of the vehicle in response to the third matching relationship being that the third driver and passenger's line of sight falls within the preset wake-up area, and in response to the third driver and passenger's gesture instruction matching the preset instruction; wherein, the wake-up instruction includes the gesture instruction.
[0133] The beneficial technical effects corresponding to the exemplary embodiment of this device can be found in the corresponding beneficial technical effects of the above exemplary method part, which will not be repeated here.
[0134] Exemplary electronic devices
[0135] Figure 6 This is a structural diagram of an electronic device 11 provided in an embodiment of the present disclosure, including at least one processor 111 and a memory 112.
[0136] The processor 111 may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 11 to perform desired functions.
[0137] The memory 112 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, a flash memory, etc. One or more computer program instructions may be stored on a computer-readable storage medium, and the processor 111 may execute one or more computer program instructions to implement the wake-up sound zone determination method and / or other desired functions of the various embodiments of the present disclosure described above.
[0138] In one example, the electronic device 11 may further include an input device 113 and an output device 114 , and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).
[0139] The input device 113 may include various sensors, including but not limited to: a distance sensor for detecting the distance between a target object and the vehicle; an image sensor for collecting information about the vehicle's surroundings. In some examples, the input device may also include a pressure sensor for detecting seat pressure to determine the presence and location of a passenger; a temperature sensor for monitoring cabin temperature; a humidity sensor for monitoring cabin humidity to assist in regulating the interior environment; an air quality sensor for monitoring interior air quality, such as carbon dioxide and volatile organic compounds (VOCs); a light sensor for detecting light intensity inside and outside the vehicle; an acceleration sensor for detecting changes in vehicle acceleration; a distance sensor for detecting the distance between the vehicle and other objects; a touchscreen sensor for interacting with the vehicle's infotainment system; biometric sensors, such as fingerprint recognition and facial recognition; a heart rate monitor for monitoring the driver's heart rate; a sound sensor for voice recognition and interaction to enable voice control; a seat sensor for monitoring seat occupancy, such as whether the seat is occupied and the passenger's body shape; and wireless communication sensors, such as Bluetooth and Wi-Fi, for connecting to smart devices to enable data transmission and remote control. In addition to the examples given above, the input device may also include more or fewer sensors, which will not be described in detail here.
[0140] The output device 114 can output various information or signals to other hardware or devices, which may include displays, vehicle audio systems, seats, windows, steering wheels, etc., as well as communication networks and remote output devices connected thereto. The displays may include multiple different display screens, such as a driver's screen, a passenger screen, and a rear screen. The vehicle audio systems may include multiple speakers located in different locations within the vehicle cabin, and the different display screens or speakers may operate independently.
[0141] Of course, to simplify, Figure 6 Only some of the components related to the present disclosure in the electronic device 11 are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, the electronic device 11 may further include any other appropriate components according to specific application scenarios.
[0142] Exemplary computer program products and computer-readable storage media
[0143] In addition to the above-mentioned methods and devices, embodiments of the present disclosure may also provide a computer program product, including computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the method for determining the wake-up sound zone of various embodiments of the present disclosure described in the above-mentioned "Exemplary Method" section.
[0144] The computer program product may be written in any combination of one or more programming languages to implement the operations of the disclosed embodiments, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as C or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0145] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the processor executes the steps in the method for determining the wake-up sound zone of various embodiments of the present disclosure described in the above-mentioned "Exemplary Method" section.
[0146] Computer readable storage media can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium is, for example, but not limited to, a system, device or component comprising electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0147] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, the advantages, strengths, and effects mentioned in this disclosure are merely illustrative and not restrictive, and should not be considered as essential to each embodiment of the present disclosure. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, rather than as limitations. These details do not limit the present disclosure to necessarily being implemented using these specific details.
[0148] Those skilled in the art may make various changes and modifications to the present disclosure without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present disclosure and their equivalents, the present disclosure is intended to include these modifications and variations.
Claims
1. A method for determining a wake-up sound zone, comprising: In response to receiving the wake-up command, acquiring an image of the vehicle cabin captured by an image sensor; wherein the image includes an image of the vehicle cabin at a time corresponding to the moment when the wake-up command is received and an image of the vehicle cabin within a preset time period before the wake-up command is received; determining, based on the image, information of the driver and passenger in the vehicle and a first sound zone corresponding to each of the driver and passenger; A target sound zone is determined according to the wake-up instruction, the driver and passenger information, and the first sound zone corresponding to each of the driver and passenger, so as to receive a voice control instruction of the driver and passenger corresponding to the target sound zone.
2. The method according to claim 1, wherein The determining of the target sound zone according to the wake-up instruction, the driver and passenger information, and the first sound zone corresponding to each of the driver and passenger includes: In response to the wake-up instruction being a touch operation and the number of drivers and passengers in the driver and passenger information being one, determining the first sound zone corresponding to the driver and passenger as the target sound zone; or In response to the wake-up instruction being a touch operation and the number of the drivers and passengers being multiple, the target sound zone is determined based on target state information in the driver and passenger information.
3. The method according to claim 2, wherein: The determining the target sound zone based on the target state information in the driver and passenger information includes: Determining the number of target drivers and passengers corresponding to the touch operation based on the operation information in the target state information; In response to the number of the target drivers and passengers being at least two, the target sound zone is determined according to target state information of at least two target drivers and passengers and the first sound zones corresponding to the two target drivers and passengers.
4. The method according to claim 3, wherein: The determining the target sound zone according to the target state information of at least two target drivers and passengers and the first sound zones corresponding to the two targets includes: Determining a time sequence in which each of the target drivers and passengers performs touch operations; The target sound zone is determined in the first sound zone corresponding to each target driver or passenger according to a time sequence in which each target driver or passenger performs a touch operation.
5. The method according to claim 3, wherein The determining the target sound zone according to the target state information of at least two target drivers and passengers and the first sound zones corresponding to the two targets includes: Determining a first matching relationship between the sight lines of each target driver and passenger and the screen wake-up area of the vehicle based on the sight line information in the target state information; The target vocal range is determined based on the first matching relationship.
6. The method according to claim 3, wherein: The determining the target sound zone according to the target state information of at least two target drivers and passengers and the first sound zones corresponding to the two targets includes: Determining, based on the touch position in the target state information, a second matching relationship between the touch position of each target driver or passenger and the screen wake-up area of the vehicle; The target vocal range is determined based on the second matching relationship.
7. The method according to claim 6, wherein: The determining, based on the touch position in the target state information, a second matching relationship between the touch position of each target driver or passenger and the screen wake-up area of the vehicle includes: Based on the touch position of each target driver or passenger, performing image cropping processing on the image corresponding to each target driver or passenger to obtain a target image area containing the touch position of each target driver or passenger; Based on the target image area corresponding to each target driver or passenger, the second matching relationship between the touch position of each target driver or passenger and the screen wake-up area of the vehicle is determined.
8. The method according to claim 6, wherein: The determining, based on the touch position in the target state information, a second matching relationship between the touch position of each target driver or passenger and the screen wake-up area of the vehicle includes: Determine a target position of a screen wake-up area of the vehicle on the target screen; Based on the degree of position matching between the touch position of each target driver and passenger and the target position, the second matching relationship between the touch position of each target driver and passenger and the screen wake-up area is determined.
9. The method according to claim 1, further comprising: determining the sight lines of each of the driver and passenger based on the image; Determining a third matching relationship between the sight lines of each of the driver and passenger and a preset wake-up area on the vehicle screen; In response to the third matching relationship being that the third driver and passenger's line of sight falls within the preset wake-up area, and in response to the third driver and passenger's gesture instruction matching the preset instruction, the voice assistant function of the vehicle is awakened; wherein, the wake-up instruction includes the gesture instruction.
10. A device for determining a wake-up sound zone, comprising: a first acquisition module, configured to acquire, in response to receiving a wake-up command, an image of the vehicle cabin captured by an image sensor; wherein the image includes an image of the vehicle cabin at a time corresponding to the moment the wake-up command is received and an image of the vehicle cabin within a preset time period before the wake-up command is received; a first determining module, configured to determine, based on the image, information of the driver and passenger in the vehicle and a first sound zone corresponding to each of the driver and passenger; The second determining module is used to determine a target sound zone according to the wake-up instruction, the driver and passenger information, and the first sound zone corresponding to each of the driver and passenger, so as to receive a voice control instruction of the driver and passenger corresponding to the target sound zone.
11. A computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the method for determining a wake-up sound zone according to any one of claims 1 to 9.
12. An electronic device, comprising: processor; a memory for storing instructions executable by the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method for determining the wake-up sound zone according to any one of claims 1 to 9.