Device wakeup method, apparatus, device, and storage medium
By acquiring images of the internal and external environments of the target space and combining them with multi-zone arbitration of image and sound acquisition devices, we ensure that the voice wake-up words originate from insiders, thus solving the problem of people outside the target space maliciously waking up the device and achieving a balance between security and cost-effectiveness.
Patent Information
- Application Number
- CN202210564171.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-23
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2042-05-23
AI Technical Summary
Existing technologies cannot effectively prevent people outside the target space from using voice wake-up words to maliciously wake up and control devices in the target space, posing a security risk.
By acquiring images of the internal and external environments of the target space, it is determined whether the interior contains human images and the exterior does not. The target device is woken up only when an internal person issues a voice wake-up word. Existing image acquisition devices and sound acquisition devices are used for multi-zone arbitration to ensure that the wake-up word comes from inside.
This effectively prevents people outside the target space from maliciously waking up or accidentally waking up the target device, improving security and reducing hardware modification costs.
Smart Images

Figure CN114999476B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to the fields of computer vision, speech technology, natural language processing, artificial intelligence, autonomous driving, and automatic parking. Background Art
[0002] With the development of artificial intelligence, many smart devices (such as in-vehicle devices, smart speakers, computers, and mobile terminals) are equipped with voice interaction functions. After waking up the smart device with a voice wake-up word, you can interact with the smart device to achieve relevant control of the smart device. Summary of the Invention
[0003] The present disclosure provides a device wake-up method, apparatus, device, and storage medium.
[0004] According to one aspect of the present disclosure, a device wake-up method is provided, comprising:
[0005] In response to the received voice wake-up word, acquiring an internal environment image and an external environment image of the target space;
[0006] In response to determining that the internal environment image includes a person image and determining that the external environment image does not include a person image, waking up a target device inside the target space according to a voice wake-up word.
[0007] According to another aspect of the present disclosure, a device waking up an apparatus is provided, comprising:
[0008] A first acquisition module is configured to acquire an internal environment image and an external environment image of a target space in response to a received voice wake-up word;
[0009] The first wake-up module is configured to wake up a target device in the target space according to a voice wake-up word in response to determining that the internal environment image contains a person image and determining that the external environment image does not contain a person image.
[0010] According to another aspect of the present disclosure, there is provided an electronic device, comprising:
[0011] at least one processor; and
[0012] a memory communicatively connected to at least one processor; wherein,
[0013] The memory stores instructions that can be executed by at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method in any embodiment of the present disclosure.
[0014] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method in any embodiment of the present disclosure.
[0015] According to another aspect of the present disclosure, there is provided a computer program product comprising a computer program which, when executed by a processor, implements the method in any of the embodiments of the present disclosure.
[0016] According to the scheme of the present disclosure, it can effectively prevent the personnel outside the target space from maliciously waking up the target device in the target space by using the voice wake-up word and further controlling the target device.
[0017] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0018] The accompanying drawings are used to better understand the present scheme and do not limit the present disclosure. Among them:
[0019] Figure 1 is a schematic diagram of a device wake-up method according to an embodiment of the present disclosure;
[0020] Figure 2 is an application schematic diagram of a device wake-up method according to an embodiment of the present disclosure;
[0021] Figure 3 is an application schematic diagram of a device wake-up method according to an embodiment of the present disclosure;
[0022] Figure 4 is a schematic diagram of a device wake-up method according to another embodiment of the present disclosure;
[0023] Figure 5 is a schematic diagram of a device wake-up method according to another embodiment of the present disclosure;
[0024] Figure 6 is a schematic diagram of a device wake-up method according to another embodiment of the present disclosure;
[0025] Figure 7 is a schematic diagram of a device wake-up method according to another embodiment of the present disclosure;
[0026] Figure 8 is a schematic diagram of a device wake-up device according to an embodiment of the present disclosure;
[0027] Figure 9 is a block diagram of an electronic device for implementing a device wake-up method according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0028] Exemplary embodiments of the present disclosure are described herein with reference to the accompanying drawings, which are meant to be exemplary in nature, and include various details intended to facilitate understanding of the present disclosure. Accordingly, it should be understood that various changes and modifications to the embodiments described herein can be made by those skilled in the art without departing from the scope and spirit of the present disclosure. Also, for the purpose of clarity and a concise description, descriptions of well-known functions and constructions are omitted from the following description.
[0029] Embodiments of the present disclosure provide a device wake-up method, which can include the following steps: Figure 1 As shown in the flowchart of the device wake-up method of the present embodiment, the method can include the following steps:
[0030] S100: In response to the received voice wake-up word, obtaining an internal environment image and an external environment image of the target space.
[0031] S101: In response to determining that the internal environment image contains a person image and determining that the external environment image does not contain a person image, waking up the target device inside the target space according to the voice wake-up word.
[0032] It should be noted that the target space can be understood as a space environment formed by the internal structure of a certain specific object. For example, the target space can be understood as a vehicle space, a house space, etc.
[0033] The target device can include a smart device with voice interaction function, such as a vehicle-mounted device, a smart speaker, a computer, a mobile terminal, etc. As long as it can realize the control of the smart device on itself or other connected devices through voice interaction.
[0034] The internal environment image can be understood as an image showing the internal environment of the target space. The external environment image can be understood as an image showing the external environment of the location where the target space is located.
[0035] According to the embodiments of the present disclosure, the internal environment image and the external environment image of the target space can be used to know the sound source position corresponding to the voice wake-up word, and ensure that the target device is woken up only when the sound is emitted by the person inside the target space, effectively preventing the person outside the target space from maliciously waking up the target device inside the target space and further controlling the target device by using the voice wake-up word.
[0036] In one example, the device wake-up method of the present embodiment can be applied to a vehicle-mounted device of a vehicle. Wherein the target space can include a vehicle space. The target device can include a vehicle-mounted device, such as a car machine. The internal environment image can include an in-vehicle environment image. The external environment image can include an out-of-vehicle environment image.
[0037] S100 can specifically include: in response to the received voice wake-up word, acquiring an in-vehicle environment image within the vehicle space by using the first image acquisition device and acquiring an out-of-vehicle environment image outside the vehicle space by using the second image acquisition device.
[0038] S101 can specifically include: in response to determining that the in-vehicle environment image contains a person image and the out-of-vehicle environment image does not contain a person image, awakening the vehicle-mounted device of the vehicle according to the voice wake-up word.
[0039] It should be noted that the first image acquisition device can be arranged inside or outside the vehicle. The first image acquisition device can be a camera or a journey recorder, etc., as long as it can realize image acquisition, which is not limited here. The image acquisition end of the first image acquisition device is arranged towards the position of the in-vehicle seat to realize the shooting of the in-vehicle environment. The acquired in-vehicle environment image can be the entire environment image of the in-vehicle trim (including all seats), the acquired in-vehicle environment image can also be the environment image of the front seat position of the in-vehicle trim, and the acquired in-vehicle environment image can also be the environment image of the front main driver seat position of the in-vehicle trim.
[0040] The second image acquisition device can be arranged inside or outside the vehicle. The second image acquisition device can be a camera or a journey recorder, etc., as long as it can realize image acquisition, which is not limited here. The image acquisition end of the second image acquisition device is arranged towards the position of the vehicle periphery to realize the shooting of the out-of-vehicle environment. The acquired out-of-vehicle environment image can be a 360° range environment image around the vehicle. The acquired out-of-vehicle environment image can also be an environment image of a certain position area around the vehicle, for example, an environment image within a certain distance range outside the main driver door, an environment image within a certain distance range outside the main driver door and outside the co-driver door, an environment image within a certain distance range outside the main driver door and outside the rear seat door on the same side, etc., and the specific acquired out-of-vehicle environment image can be selected and adjusted as needed.
[0041] The vehicle-mounted device can be understood as any device that can perform voice interaction and realize the function of voice assistant on the vehicle. For example, the vehicle-mounted device can be a vehicle machine or a vehicle control system that can perform voice interaction. Through voice interaction, the vehicle machine or the vehicle control system can realize the control of some functions of the vehicle. For example, opening the window, opening the air conditioner, opening the automatic driving mode, playing music, etc.
[0042] The in-vehicle environment image containing a person image can be understood as that there is a person at any seat position of the entire vehicle trim shown in the in-vehicle environment image. It can also be understood that there is a person at a certain specific seat position of the vehicle trim shown in the in-vehicle environment image, for example, there is a person at the main driver seat position.
[0043] It can be understood that there is no person within a certain distance range around the vehicle. It can also be understood that there is no person within a certain distance range at a certain specific direction around the vehicle. For example, the specific direction can be a certain distance range outside the main driver door. In this case, even if the collected vehicle external environment image is a 360° environment image around the vehicle, and there is a person appearing in the environment image close to the rear of the vehicle, it can also be considered that the vehicle external environment image does not contain a person image. The main driver door is a specific direction only for illustration, and can also be a certain distance range outside the main driver door and the passenger door, or a certain distance range outside the main driver door and the rear door on the same side, etc., which is not limited here.
[0044] According to the scheme of the present disclosure, it can effectively prevent the person outside the vehicle from maliciously waking up the vehicle-mounted device by using the voice wake-up word and further controlling the vehicle, or the person outside the vehicle from mistakenly waking up the vehicle-mounted device by using the voice wake-up word. At the same time, the implementation of the embodiment of the present disclosure does not need to additionally modify the hardware of the vehicle, and the existing image acquisition device on the vehicle can be reused to implement the method of the present embodiment, thereby reducing the cost of implementing the present method.
[0045] In one example, the device wake-up method of the present embodiment can be applied to a control device of a house. Wherein, the target space can include a space of the house. The target device can include a control device with voice interaction function and control function. The internal environment image can include an environment image inside the house. The external environment image can include an environment image outside the house.
[0046] S100 can specifically include: in response to the received voice wake-up word, acquiring an environment image inside the space of the house by using the first image acquisition device and acquiring an environment image outside the space of the house by using the second image acquisition device.
[0047] S101 can specifically include: in response to determining that the environment image inside the space of the house contains a person image and the environment image outside the space of the house does not contain a person image, waking up the control device according to the voice wake-up word.
[0048] It should be noted that the first image acquisition device and the second image acquisition device can be arranged inside or outside the house. The first image acquisition device and the second image acquisition device can be a camera, a cat eye with camera function, etc., as long as image acquisition can be realized, which is not limited here.
[0049] A control device can be understood as any device capable of voice interaction and further control functions, such as a computer, mobile terminal, cabinet, air conditioner, etc. If the control device is an air conditioner, voice interaction can be used to turn on the air conditioner to cool the room. If the control device is a computer, voice interaction can be used to turn on any other device that the computer can control.
[0050] In one example, the device wake-up method provided by the embodiment of the present disclosure can be applied to Figure 2 In the scenario framework shown. Figure 2 In the figure, 10 represents the vehicle control system, 20 represents the server, and 30 represents the distributed server cluster. The device wake-up method disclosed herein can be executed by the vehicle control system 10, the server 20, or the distributed server cluster 30. Alternatively, the device wake-up method can be executed by both the vehicle control system 10 and the server 20, or by both the vehicle control system 10 and the distributed server cluster 30.
[0051] In one embodiment, the device wake-up method provided by the embodiment of the present disclosure includes steps S100 to S101 above, wherein step S100 of the method may include:
[0052] In response to the received voice wake-up word, an internal environment image of the target space and an external environment image of a preset position of the target space are acquired.
[0053] It should be noted that the preset position may include any position in the external area around the target space.
[0054] According to the solution disclosed herein, by collecting environmental images of a preset area outside the target space, it can be specifically determined whether a person at a certain location maliciously wakes up the target device in an attempt to control the target device.
[0055] In one example, the device wakeup method according to an embodiment of the present disclosure can be applied to an in-vehicle device. The target space can include the vehicle space. The target device can include an in-vehicle device, such as a vehicle computer. The interior environment image can include an image of the in-vehicle environment. The exterior environment image can include an image of the exterior environment.
[0056] In response to the received voice wake-up word, obtaining an internal environment image of the target space and an external environment image of a preset position of the target space may specifically include: in response to the received voice wake-up word, obtaining an internal environment image of the vehicle captured by a first image acquisition device and an external environment image of a preset position captured by a second image acquisition device, wherein the preset position at least includes a preset area range A outside the main driver's door (such as Figure 3 shown).
[0057] When the collected out-of-vehicle environment image at the preset area range A has no human image (i.e., no person 1 appears) and the in-vehicle environment image has a human image, the step of: in response to determining that the in-vehicle environment image contains a human image and the out-of-vehicle environment image does not contain a human image, waking up the vehicle-mounted device of the vehicle according to the voice wake-up word, is not continued to be executed. At this time, it is indicated that the voice wake-up word is not spoken by the in-vehicle person.
[0058] When the collected out-of-vehicle environment image at the preset area range A has a human image (i.e., person 1 appears) and the in-vehicle environment image has no human image, the step of: in response to determining that the in-vehicle environment image contains a human image and the out-of-vehicle environment image does not contain a human image, waking up the vehicle-mounted device of the vehicle according to the voice wake-up word, is not continued to be executed. At this time, it is indicated that the voice wake-up word is not spoken by the in-vehicle person.
[0059] The preset position can include one or more of the preset area range outside the co-driver door, the preset area range in front of the vehicle head, the preset area range outside the rear door, and the preset area range behind the parking space in addition to the preset area range A outside the main driver door.
[0060] According to the scheme of the present disclosure, by collecting the out-of-vehicle environment image of the preset area range outside the main driver door, it can be determined whether the out-of-vehicle person maliciously wakes up the vehicle-mounted device from the main driver position and attempts to control the vehicle by using the vehicle-mounted device.
[0061] In one example, the out-of-vehicle environment image can include an environment image within a 360° range around the vehicle. As shown in Figure 3 When any one of the person 1, the person 2, the person 3, or the person 4 around the vehicle appears in the collected out-of-vehicle environment image, it can be considered that the out-of-vehicle environment image has a human image. When none of the person 1, the person 2, the person 3, or the person 4 appears in the collected out-of-vehicle environment image, it is considered that the out-of-vehicle environment image has no human image.
[0062] In one example, the device wake-up method of the embodiments of the present disclosure can be applied to a control device of a house. The target space can include a space of the house. The target device can include a control device with voice interaction function and control function. The internal environment image can include an environment image inside the house. The external environment image can include an environment image outside the house.
[0063] In response to the received voice wake-up word, the internal environment image of the target space and the external environment image of the preset position of the target space can be obtained, which can specifically include: in response to the received voice wake-up word, obtaining an environment image inside the house collected by a first image collection device and an environment image of a preset position outside the house collected by a second image collection device. The preset position at least includes the outside of the door and / or the outside of the window.
[0064] When the preset position outside the house has a person image in the environment image, and the house has no person image in the environment image, the step of waking up the control device according to the voice wake-up word in response to the determination that the environment image in the space of the house contains a person image and the environment image outside the space of the house does not contain a person image is not continued. At this time, it is indicated that the voice wake-up word is not spoken by the person inside the house.
[0065] When the preset position outside the house has a person image in the environment image, and the house has no person image in the environment image, the step of waking up the control device according to the voice wake-up word in response to the determination that the environment image in the space of the house contains a person image and the environment image outside the space of the house does not contain a person image is not continued. At this time, it is indicated that the voice wake-up word is not spoken by the person inside the house.
[0066] In an embodiment, as shown in Figure 4 The device wake-up method provided by the embodiment of the present disclosure includes the steps S100-S101, and can further include:
[0067] S400: In response to the determination that the internal environment image and the external environment image both contain person images, a face image of the person inside the target space is acquired.
[0068] S401: According to the face image, a mouth shape change of the person inside the target space is determined.
[0069] S402: In response to the mouth shape change of the person meeting a preset shape change requirement, a target device is woken up according to a voice wake-up word.
[0070] It should be noted that the preset shape change can be understood as the mouth being in a state of continuously opening and closing, or the mouth being in an open state in one frame of image.
[0071] The person image in the external environment image can be understood as the appearance of the person at any position within the 360° range around the target space. It can also be understood as the appearance of the person at a specific position within the 360° range around the target space.
[0072] According to the scheme of the present disclosure, in the case that the internal environment image and the external environment image both contain person images, by determining the mouth shape change of the person inside the target space, it can be further determined whether the voice wake-up word is spoken by the person inside the target space, preventing the person outside the target space from waking up the target device and further controlling the target device through the voice wake-up word in the case that the person inside the target space does not speak the voice wake-up word.
[0073] In one example, the device wake-up method of the embodiments of the present disclosure can be applied to a vehicle-mounted device of a vehicle. Among them, the target space can include an internal space of the vehicle. The target device can include the vehicle-mounted device. The internal environment image can include an in-vehicle environment image. The external environment image can include an out-of-vehicle environment image.
[0074] S400 can specifically include: in response to determining that the in-vehicle environment image and the out-of-vehicle environment image both contain a person image, acquiring a face image of an in-vehicle person in the vehicle space by using a third image acquisition device.
[0075] S401 can specifically include: determining a mouth shape change of the in-vehicle person according to the face image.
[0076] S402 can specifically include: in response to the mouth shape change of the in-vehicle person meeting a preset shape change requirement, waking up the vehicle-mounted device according to the voice wake-up word.
[0077] It should be noted that the third image acquisition device can be the same image acquisition device as the first image acquisition device, or a different image acquisition device. The face image of the in-vehicle person acquired by the third image acquisition device can be one frame of image, or a plurality of continuous frames of image.
[0078] The preset shape change can be understood as the mouth being in a state of continuously opening and closing, or can be understood as the mouth being in an open state in one frame of image.
[0079] The in-vehicle environment image containing a person image can be understood as a person appearing at any position in the 360° out-of-vehicle environment image around the vehicle, for example, Figure 3 The personnel 1, personnel 2, personnel 3, and personnel 4 shown. It can also be understood that only the appearance of personnel 1 is considered as the out-of-vehicle environment image containing a person image, and the appearance of personnel 2, personnel 3, and personnel 4 is not considered as the basis for determining whether the out-of-vehicle environment image contains a person image. It can also be understood that only the appearance of personnel 1 and / or personnel 2 is considered as the out-of-vehicle environment image containing a person image, and the appearance of personnel 3 and personnel 4 is not considered as the basis for determining whether the out-of-vehicle environment image contains a person image.
[0080] According to the scheme of the present disclosure, in the case that the in-vehicle environment image and the out-of-vehicle environment image both contain a person image, by determining the mouth shape change of the in-vehicle person, it can be further determined whether the voice wake-up word is spoken by the in-vehicle person, and the situation that the in-vehicle person does not speak the voice wake-up word is prevented, the out-of-vehicle person wakes up the vehicle-mounted device by the voice wake-up word and further controls the vehicle, effectively avoiding the threat to the personal and property safety of the in-vehicle person.
[0081] In one example, the device wake-up method of the embodiments of the present disclosure can be applied to a control device of a house. The target space can include a space of the house. The target device can include a control device with voice interaction function and control function. The internal environment image can include an environment image inside the house. The external environment image can include an environment image outside the house.
[0082] S400 can specifically include: in response to determining that the environment image inside the house and the environment image outside the house both contain a person image, acquiring a face image of the person inside the house by using a third image acquisition device.
[0083] S401 can specifically include: determining a mouth shape change of the person inside the house according to the face image.
[0084] S402 can specifically include: in response to the mouth shape change of the person inside the house meeting a preset shape change requirement, waking up the control device according to the voice wake-up word.
[0085] It should be noted that the third image acquisition device can be the same image acquisition device as the first image acquisition device, or different image acquisition devices. The face image of the person inside the house acquired by the third image acquisition device can be one frame of image, or a plurality of continuous frames of image.
[0086] The preset shape change can be understood as the mouth being in a state of continuously opening and closing, or the mouth being in an open state in one frame of image.
[0087] According to the scheme of the present disclosure, in the case that the environment image inside the house and the environment image outside the house both contain a person image, by determining the mouth shape change of the person inside the house, it can be further determined whether the voice wake-up word is spoken by the person inside the house, preventing the person outside the house from waking up the control device by the voice wake-up word and further controlling the control device in the case that the person inside the house does not speak the voice wake-up word, effectively avoiding threatening the personal and property safety of the person inside the house.
[0088] In one embodiment, as shown in Figure 5 The device wake-up method provided by the embodiments of the present disclosure includes the above steps S100 to S101, and can further include:
[0089] S500: in response to determining that the internal environment image and the external environment image both contain a person image, determining a sound acquisition device for collecting the voice wake-up word.
[0090] S501: in response to the sound acquisition device inside the target space collecting the voice wake-up word and the sound acquisition device outside the target space not collecting the voice wake-up word, waking up the target device according to the voice wake-up word.
[0091] It should be noted that the sound collection device includes at least two, one of which is arranged outside the target space, and the other is arranged inside the target space. The sound collection device can adopt any structure in the prior art, which is not limited here, for example, the sound collection device can adopt a microphone.
[0092] When the sound collection device inside the target space collects the voice wake-up word, and the sound collection device outside the target space does not collect the voice wake-up word, it indicates that the voice wake-up word is spoken by the person inside the target space, and the target device can be woken up in response to the voice wake-up word.
[0093] When the sound collection device inside the target space does not collect the voice wake-up word, and the sound collection device outside the target space collects the voice wake-up word, it indicates that the voice wake-up word is spoken by the person outside the target space, and the target device is not woken up in response to the voice wake-up word. At this time, it indicates that the person outside the target space uses the voice wake-up word to maliciously wake up the target device and intends to further control.
[0094] According to the scheme of the present disclosure, the sound collection devices inside and outside the target space are used for multi-sound area arbitration, which can further accurately determine whether the voice wake-up word is spoken by the person inside the target space, and prevent the person outside the target space from waking up the target device through the voice wake-up word and further controlling when the person inside the target space does not speak the voice wake-up word.
[0095] In one example, the device wake-up method of the embodiments of the present disclosure can be applied to the vehicle-mounted device of a vehicle. Among them, the target space can include the vehicle space. The target device can include a vehicle-mounted device, such as a car machine. The internal environment image can include an in-vehicle environment image. The external environment image can include an out-of-vehicle environment image.
[0096] S500 can specifically include: in response to determining that the in-vehicle environment image and the out-of-vehicle environment image both contain a person image, determining a sound collection device that collects the voice wake-up word.
[0097] S501 can specifically include: in response to the in-vehicle sound collection device collecting the voice wake-up word and the out-of-vehicle sound collection device not collecting the voice wake-up word, waking up the vehicle-mounted device according to the voice wake-up word.
[0098] It should be noted that the sound collection device includes at least two, one of which is arranged outside the target space, and the other is arranged inside the target space. The sound collection device can adopt any structure in the prior art, which is not limited here, for example, the sound collection device can adopt a microphone.
[0099] When the in-vehicle sound collection device collects the voice wake-up word and the out-of-vehicle sound collection device does not collect the voice wake-up word, it is determined that the voice wake-up word is spoken by the person in the vehicle, and the vehicle-mounted device can be woken up in response to the voice wake-up word.
[0100] When the in-vehicle sound collection device does not collect the voice wake-up word and the out-of-vehicle sound collection device collects the voice wake-up word, it is determined that the voice wake-up word is spoken by the person outside the vehicle, and the vehicle-mounted device is not woken up in response to the voice wake-up word. At this time, it is indicated that the person outside the vehicle wants to maliciously wake up the vehicle-mounted device and further control the vehicle by using the voice wake-up word.
[0101] The out-of-vehicle environment image contains a person image, which can be understood as that the person appears at any position in the 360° out-of-vehicle environment image around the vehicle, for example, Figure 3 The person 1, the person 2, the person 3, and the person 4 shown. It can also be understood that only when the person 1 appears, it is considered that the out-of-vehicle environment image contains a person image, and whether the person 2, the person 3, and the person 4 appear or not is not a basis for determining whether the out-of-vehicle environment image contains a person image. It can also be understood that only when the person 1 and / or the person 2 appear, it is considered that the out-of-vehicle environment image contains a person image, and whether the person 3 and the person 4 appear or not is not a basis for determining whether the out-of-vehicle environment image contains a person image.
[0102] According to the scheme of the present disclosure, the sound collection devices inside and outside the vehicle are used for multi-sound area arbitration, which can further accurately determine whether the voice wake-up word is spoken by the person inside the vehicle, and prevent the person outside the vehicle from waking up the vehicle-mounted device and further controlling the vehicle by using the voice wake-up word when the person inside the vehicle does not speak the voice wake-up word, thereby effectively avoiding threats to the personal and property safety of the person inside the vehicle. At the same time, the implementation of the embodiment of the present disclosure does not need to additionally modify the hardware of the vehicle, and the existing sound collection devices on the vehicle can be reused to implement the method of the present embodiment, thereby reducing the cost of implementing the present method.
[0103] In one example, the device wake-up method of the embodiment of the present disclosure can be applied to a control device of a house. The target space can include a space of the house. The target device can include a control device with a voice interaction function and a control function. The internal environment image can include an environment image inside the house. The external environment image can include an environment image outside the house.
[0104] S500 can specifically include: in response to determining that the environment image inside the house and the environment image outside the house both contain person images, determining a sound collection device that collects the voice wake-up word.
[0105] S501 can specifically include: in response to the in-vehicle sound collection device collecting the voice wake-up word and the in-vehicle and out-of-vehicle sound collection devices not collecting the voice wake-up word, waking up the control device according to the voice wake-up word.
[0106] It should be noted that the sound collection device is at least two, one is set outside the house, the other is set inside the house. The sound collection device can adopt any structure in the prior art, which is not limited here, for example, the sound collection device can adopt a microphone.
[0107] When the sound collection device inside the house collects the voice wake-up word and the sound collection device outside the house does not collect the voice wake-up word, it means that the voice wake-up word is spoken by the person inside the house, and the control device can be woken up in response to the voice wake-up word.
[0108] When the sound collection device inside the house does not collect the voice wake-up word and the sound collection device outside the house collects the voice wake-up word, it means that the voice wake-up word is spoken by the person outside the house, and the target device will not be woken up in response to the voice wake-up word. At this time, it means that the person outside the house wants to maliciously wake up the control device using the voice wake-up word.
[0109] According to the scheme of the present disclosure, the sound collection devices inside and outside the house are used for multi-sound area arbitration, which can further accurately determine whether the voice wake-up word is spoken by the person inside the house, and prevent the person outside the house from waking up the control device and further controlling it through the voice wake-up word when the person inside the house does not speak the voice wake-up word, effectively avoiding threats to the personal and property safety of the person inside the house.
[0110] In one embodiment, as shown in Figure 6 The device wake-up method provided by the embodiment of the present disclosure includes steps S100-S101, and can further include:
[0111] S600: In response to determining that the internal environment image and the external environment image both contain a person image, and determining that the sound collection device inside the target space and the sound collection device outside the target space both collect the voice wake-up word, the volume of the voice wake-up word collected by the sound collection device inside the target space and the sound collection device outside the target space is determined respectively.
[0112] S601: In response to determining that the volume of the voice wake-up word collected by the sound collection device inside the target space is large, the target device is woken up according to the voice wake-up word.
[0113] According to the scheme of the present disclosure, the sound collection devices inside and outside the target space are used for multi-sound area arbitration, and the volume can be used to further accurately determine whether the voice wake-up word is spoken by the person inside the target space, and prevent the person outside the target space from waking up the target device and further controlling it through the voice wake-up word when the person inside the target space does not speak the voice wake-up word, effectively avoiding threats to the personal and property safety of the person inside the target space.
[0114] In one example, the device wake-up method of the embodiments of the present disclosure can be applied to a vehicle-mounted device of a vehicle. Among them, the target space can include a vehicle space. The target device can include a vehicle-mounted device, such as a car machine. The internal environment image can include an in-vehicle environment image. The external environment image can include an out-of-vehicle environment image.
[0115] S600 can specifically include: in response to determining that the in-vehicle environment image and the out-of-vehicle environment image both contain a person image, and determining that the in-vehicle sound collection device and the out-of-vehicle sound collection device both collect a voice wake-up word, respectively determining the volume of the voice wake-up word collected by the in-vehicle sound collection device and the out-of-vehicle sound collection device.
[0116] S601 can specifically include: in response to determining that the volume of the voice wake-up word collected by the in-vehicle sound collection device is larger, responding to the voice wake-up word to wake up the vehicle-mounted device.
[0117] It should be noted that when the vehicle is parked and the driver opens the window for rest, the situation that the voice wake-up word maliciously spoken by the person outside the vehicle is collected by the in-vehicle sound collection device and the out-of-vehicle sound collection device at the same time can occur.
[0118] The fact that the out-of-vehicle environment image contains a person image can be understood as that there is a person appearing at any position in the 360° out-of-vehicle environment image around the vehicle, for example, Figure 3 person 1, person 2, person 3, and person 4. It can also be understood that only when person 1 appears is the out-of-vehicle environment image considered to contain a person image, and whether person 2, person 3, or person 4 appears is not considered as a basis for determining whether the out-of-vehicle environment image contains a person image. It can also be understood that only when person 1 and / or person 2 appears is the out-of-vehicle environment image considered to contain a person image, and whether person 3 or person 4 appears is not considered as a basis for determining whether the out-of-vehicle environment image contains a person image.
[0119] According to the scheme of the present disclosure, the in-vehicle and out-of-vehicle sound collection devices are used for multi-sound area arbitration, and the volume size can be used to further accurately determine whether the voice wake-up word is spoken by the in-vehicle person, so as to prevent the situation that the in-vehicle person does not speak the voice wake-up word, and the out-of-vehicle person wakes up the vehicle-mounted device through the voice wake-up word and further controls the vehicle, effectively avoiding the threat to the personal and property safety of the in-vehicle person. At the same time, the implementation of the embodiments of the present disclosure does not need to additionally modify the hardware of the vehicle, and the existing sound collection device on the vehicle can be reused to implement the method of the present embodiment, thereby reducing the cost of implementing the present method.
[0120] In one example, the device wake-up method of the embodiments of the present disclosure can be applied to a control device of a house. The target space can include a space of the house. The target device can include a control device with voice interaction function and control function. The internal environment image can include an environment image inside the house. The external environment image can include an environment image outside the house.
[0121] S600 can specifically include: in response to determining that the environment image inside the house and the environment image outside the house both contain a person image, and determining that the sound collection device inside the house and the sound collection device outside the house both collect a voice wake-up word, respectively determining the volume of the voice wake-up word collected by the sound collection device inside the house and the sound collection device outside the house.
[0122] S601 can specifically include: in response to determining that the volume of the voice wake-up word collected by the sound collection device inside the house is larger, responding to the voice wake-up word to wake up the control device.
[0123] In one embodiment, as shown in FIG. 1, Figure 7 The device wake-up method provided by the embodiments of the present disclosure includes the above steps S100-S101, and can further include:
[0124] S700: in response to determining that the internal environment image and the external environment image both contain a person image, and determining that the sound collection device inside the target space and the sound collection device outside the target space both collect a voice wake-up word, determining the sound wave receiving angle of the voice wake-up word collected by the sound collection device inside the target space.
[0125] S701: in response to the sound wave receiving angle satisfying a threshold angle, waking up the target device according to the voice wake-up word.
[0126] It should be noted that the threshold angle can be confirmed and adjusted according to the setting position of the sound collection device inside the target space. The external environment image containing a person image can be understood as that there is a person appearing at any position in the 360° environment image around the external environment image. It can also be understood that only a specific position in the 360° environment image around the external environment image contains a person.
[0127] According to the scheme of the present disclosure, the sound collection devices inside and outside the target space are used for multi-sound area arbitration, and the sound wave receiving angle can be used to further accurately determine whether the voice wake-up word is spoken by the person inside the target space, so as to prevent the person outside the target space from waking up the target device through the voice wake-up word and further controlling the target device in the case that the person inside the target space does not speak the voice wake-up word, effectively avoiding threats to the personal and property safety of the person inside the target space.
[0128] In one example, the device wake-up method of the embodiments of the present disclosure can be applied to a vehicle-mounted device of a vehicle. Among them, the target space can include an internal space of the vehicle. The target device can include the vehicle-mounted device. The internal environment image can include an in-vehicle environment image. The external environment image can include an out-of-vehicle environment image.
[0129] S700 can specifically include: in response to determining that the in-vehicle environment image and the out-of-vehicle environment image both contain a person image, and determining that the in-vehicle sound collection device and the out-of-vehicle sound collection device both collect a voice wake-up word, determining the sound wave receiving angle of the voice wake-up word collected by the in-vehicle sound collection device.
[0130] S701 can specifically include: in response to the sound wave receiving angle satisfying a threshold angle, waking up the vehicle-mounted device according to the voice wake-up word.
[0131] It should be noted that the threshold angle can be confirmed and adjusted according to the setting position of the in-vehicle sound collection device. For example, when the in-vehicle sound collection device is set at a vehicle interior position facing the main driver's seat, the sound wave angle of the sound wave received by the in-vehicle sound collection device when the driver of the main driver's seat speaks the wake-up word is approximately an acute angle or 0°. When the angle of the sound wave received by the in-vehicle sound collection device is an obtuse angle, it means that it may be a malicious voice wake-up word spoken by a person in the external environment of the vehicle.
[0132] The fact that the out-of-vehicle environment image contains a person image can be understood as that a person appears at any position in the 360° out-of-vehicle environment image around the vehicle, for example, Figure 3 person 1, person 2, person 3, and person 4. It can also be understood that only when person 1 appears is the out-of-vehicle environment image considered to contain a person image, and whether person 2, person 3, or person 4 appears is not a basis for determining whether the out-of-vehicle environment image contains a person image. It can also be understood that only when person 1 and / or person 2 appear is the out-of-vehicle environment image considered to contain a person image, and whether person 3 or person 4 appears is not a basis for determining whether the out-of-vehicle environment image contains a person image.
[0133] According to the scheme of the present disclosure, the in-vehicle and out-of-vehicle sound collection devices are used for multi-sound area arbitration, and the sound wave receiving angle can be used to further accurately determine whether the voice wake-up word is spoken by the in-vehicle person, so as to prevent the out-of-vehicle person from waking up the vehicle-mounted device through the voice wake-up word and further controlling the vehicle without the in-vehicle person speaking the voice wake-up word, thereby effectively avoiding threats to the personal and property safety of the in-vehicle person. At the same time, the implementation of the embodiments of the present disclosure does not need to additionally modify the hardware of the vehicle, and the existing sound collection device on the vehicle can be reused to implement the method of the present embodiment, thereby reducing the cost of implementing the present method.
[0134] In one example, the device wakeup method according to an embodiment of the present disclosure can be applied to a control device in a house. The target space can include a space in the house. The target device can include a control device with voice interaction and control functions. The internal environment image can include an image of the environment inside the house. The external environment image can include an image of the environment outside the house.
[0135] S700 may specifically include: in response to determining that both the environment image inside the house and the environment image outside the house contain human images, and determining that both the sound collection device inside the house and the sound collection device outside the house have collected the voice wake-up word, determining the sound wave reception angle of the voice wake-up word collected by the sound collection device inside the house.
[0136] S701 may specifically include: in response to the sound wave receiving angle meeting the threshold angle, waking up the control device according to the voice wake-up word.
[0137] The embodiment of the present disclosure provides a device waking up a device, such as Figure 8 As shown in FIG, it is a structural block diagram of the device wake-up apparatus of this embodiment, and the apparatus may include:
[0138] The first acquisition module 810 is configured to acquire an internal environment image and an external environment image of a target space in response to a received voice wake-up word.
[0139] The first wake-up module 820 is configured to wake up a target device in the target space according to a voice wake-up word in response to determining that the internal environment image includes a person image and determining that the external environment image does not include a person image.
[0140] In one embodiment, the first acquisition module 810 is configured to acquire an internal environment image of the target space and an external environment image of a preset position of the target space in response to a received voice wake-up word.
[0141] In one embodiment, the device waking up the device further includes:
[0142] The second acquisition module is configured to acquire a facial image of a person inside the target space in response to determining that both the internal environment image and the external environment image contain a person image.
[0143] The first determination module is used to determine the mouth shape changes of people in the target space based on the facial image.
[0144] The second wake-up module is used to wake up the target device according to the voice wake-up word in response to the change of the person's mouth shape meeting the preset shape change requirements.
[0145] In one embodiment, the device waking up the device further includes:
[0146] The second determining module is configured to determine the sound collection device that collects the voice wake-up word in response to determining that the internal environment image and the external environment image both contain the image of the person.
[0147] The third wake-up module is configured to wake up the target device according to the voice wake-up word in response to the sound collection device inside the target space collecting the voice wake-up word and the sound collection device outside the target space not collecting the voice wake-up word.
[0148] In an embodiment, the device wake-up apparatus further comprises:
[0149] The third determining module is configured to determine the volume of the voice wake-up word collected by the sound collection device inside the target space and the sound collection device outside the target space respectively in response to determining that the internal environment image and the external environment image both contain the image of the person and determining that the sound collection device inside the target space and the sound collection device outside the target space both collect the voice wake-up word.
[0150] The fourth wake-up module is configured to wake up the target device according to the voice wake-up word in response to determining that the volume of the voice wake-up word collected by the sound collection device inside the target space is large.
[0151] In an embodiment, the device wake-up apparatus further comprises:
[0152] The fourth determining module is configured to determine the sound wave receiving angle of the voice wake-up word collected by the sound collection device inside the target space in response to determining that the internal environment image and the external environment image both contain the image of the person and determining that the sound collection device inside the target space and the sound collection device outside the target space both collect the voice wake-up word.
[0153] The fifth wake-up module is configured to wake up the target device according to the voice wake-up word in response to the sound wave receiving angle satisfying a threshold angle.
[0154] In the technical solution of the present disclosure, the acquisition, storage and application of user personal information comply with relevant laws and regulations and do not violate public order and good customs.
[0155] The device wake-up apparatus of the embodiments of the present disclosure can perform the device wake-up method of any of the above embodiments. The device wake-up apparatus of the embodiments of the present disclosure can be applied to the scenarios to which the device wake-up method of any of the above embodiments is applied.
[0156] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.
[0157] Figure 9A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0158] like Figure 9 As shown, the device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the device 900 can also be stored in the RAM 903. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0159] Various components in device 900 are connected to I / O interface 905, including an input unit 906, such as a keyboard and mouse; an output unit 907, such as various types of displays and speakers; a storage unit 908, such as a magnetic disk and optical disk; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0160] The computing unit 901 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs various methods and processes described above, such as the device wake-up method. For example, in some embodiments, the device wake-up method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded onto the RAM 903 and executed by the computing unit 901, one or more steps of the device wake-up method described above can be performed. Alternatively, in other embodiments, the computing unit 901 can be configured to perform the device wake-up method by any other suitable means, such as by means of firmware.
[0161] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0162] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or the block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.
[0163] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0164] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0165] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0166] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0167] It should be understood that the various forms of flow shown above can be used to reorder, add, or delete steps. For example, the steps recited in the present disclosure can be performed in parallel, in series, or in a different order, without limitation herein, so long as the desired results of the technology disclosed in the present disclosure are achieved.
[0168] The specific implementation described above does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A device wake-up method, comprising: In response to the received voice wake-up word, acquiring an internal environment image and an external environment image of the target space; In response to determining that the internal environment image includes a person image and determining that the external environment image does not include a person image, waking up a target device in the target space according to the voice wake-up word; The method further comprises: In response to determining that both the internal environment image and the external environment image contain human images, obtaining a facial image of the person inside the target space; determining a mouth morphology change of the person inside the target space based on the facial image; in response to the mouth morphology change of the person meeting a preset morphology change requirement, waking up the target device according to the voice wake-up word; or, In response to determining that both the internal environment image and the external environment image contain human images, determining a sound collection device for collecting the voice wake-up word; in response to the sound collection device inside the target space collecting the voice wake-up word and the sound collection device outside the target space not collecting the voice wake-up word, waking up the target device according to the voice wake-up word; or In response to determining that both the internal environment image and the external environment image contain human images, and determining that both the sound collection device inside the target space and the sound collection device outside the target space have collected the voice wake-up word, respectively determining the volume of the voice wake-up word collected by the sound collection device inside the target space and the sound collection device outside the target space; in response to determining that the volume of the voice wake-up word collected by the sound collection device inside the target space is greater, waking up the target device according to the voice wake-up word; or, In response to determining that both the internal environment image and the external environment image contain human images, and determining that both the sound collection device inside the target space and the sound collection device outside the target space have collected the voice wake-up word, a sound wave reception angle of the voice wake-up word collected by the sound collection device inside the target space is determined; in response to the sound wave reception angle meeting a threshold angle, the target device is woken up according to the voice wake-up word.
2. The method according to claim 1, wherein The step of acquiring an internal environment image and an external environment image of a target space in response to a received voice wake-up word includes: In response to the received voice wake-up word, an internal environment image of a target space and an external environment image of a preset position of the target space are acquired.
3. A device wake-up apparatus, comprising: A first acquisition module is configured to acquire an internal environment image and an external environment image of a target space in response to a received voice wake-up word; a first wake-up module, configured to wake up a target device in the target space according to the voice wake-up word in response to determining that the internal environment image includes a person image and determining that the external environment image does not include a person image; The device waking up device further includes: a second acquisition module for acquiring a facial image of a person in the target space in response to determining that both the internal environment image and the external environment image contain a person image; a first determination module for determining a mouth morphology change of the person in the target space based on the facial image; a second wake-up module for waking up the target device according to the voice wake-up word in response to the mouth morphology change of the person meeting a preset morphology change requirement; or a second determining module for determining a sound collection device for collecting the voice wake-up word in response to determining that both the internal environment image and the external environment image contain human images; a third waking-up module for waking up the target device according to the voice wake-up word in response to the sound collection device inside the target space collecting the voice wake-up word and the sound collection device outside the target space not collecting the voice wake-up word; or a third determining module for, in response to determining that both the internal environment image and the external environment image contain human images and determining that both the sound collection device inside the target space and the sound collection device outside the target space have collected the voice wake-up word, determining the volume of the voice wake-up word collected by the sound collection device inside the target space and the sound collection device outside the target space respectively; a fourth waking-up module for, in response to determining that the volume of the voice wake-up word collected by the sound collection device inside the target space is larger, waking up the target device according to the voice wake-up word; or The fourth determination module is used to determine the sound wave reception angle of the voice wake-up word collected by the sound collection device inside the target space in response to determining that both the internal environment image and the external environment image contain human images, and determining that both the sound collection device inside the target space and the sound collection device outside the target space have collected the voice wake-up word; the fifth wake-up module is used to wake up the target device according to the voice wake-up word in response to the sound wave reception angle meeting the threshold angle.
4. The device according to claim 3, wherein The first acquisition module is configured to acquire an internal environment image of a target space and an external environment image of a preset position of the target space in response to a received voice wake-up word.
5. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 2.
6. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 2.
7. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 2.
Citation Information
Patent Citations
Vehicle voice starting device
CN106627491A