Method for waking up device, and system and chip

Through the cascading neural network, the target object in the moving area image is identified in image processing, which solves the problem that existing devices are prone to being accidentally awakened, and achieves lower power consumption and higher wake-up accuracy.

WO2025124369A1PCT designated stage expired Publication Date: 2025-06-19HUAWEI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/138048
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-12
Filing Date
2024-12-10
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing wake-up devices are prone to be accidentally awakened due to non-target objects or ambient temperature changes, resulting in increased power consumption.

Method used

Using a cascading first neural network and a second neural network, the image processing is used to determine whether there is a target object in the moving area image, and a trigger signal wake-up device is generated. The first neural network makes preliminary judgments, and the second neural network performs secondary confirmation to reduce false wake-up.

Benefits of technology

It effectively reduces the frequency of false wake-up, reduces system power consumption, and improves wake-up speed and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024138048_19062025_PF_FP_ABST
    Figure CN2024138048_19062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides a method for waking up a device, and a system and a chip. The method may comprise: processing a collected first image, determining a motion area image in the first image, and by means of a primary neural network, determining whether a target object is present in the motion area image. When it is determined that the target object is present in the motion area image, by means of a secondary neural network cascaded with the primary neural network, secondary confirmation is performed on a result outputted by the primary neural network. Hence, determination on the target object is performed by means of the primary neural network and the secondary neural network, thereby reducing the false alarm frequency, reducing the system power consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Method, system and chip for waking up device

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on December 12, 2023, with application number 202311702984.3, and invention name “Method, system and chip for waking up device”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] This application relates to the field of artificial intelligence, and in particular to methods, systems, and chips for waking up devices. Background Art

[0003] As technology advances, the way we interact with devices is changing. We now use voice, gestures, and visual interactions to interact with devices. These interactions increase the amount of data computation required on devices, requiring them to use processors with greater computing power and higher power consumption. To reduce the power consumption of these processors, smaller "wake-up" units are used to detect input from multiple sensor interfaces. Once a target object is detected, the "wake-up" unit wakes up the processor.

[0004] A common "wake-up" unit can be a pyroelectric infrared sensor (PIR). See Figure 1A, which shows a schematic diagram of a PIR-based device wake-up scenario. As shown in Figure 1A, the PIR sensor is a passive sensor based on infrared radiation. It can detect infrared radiation generated by an object and use it as a trigger signal to wake up the detection device. The working principle of the low-power wake-up strategy based on PIR is as follows: when an object that can generate infrared radiation enters the detection area of ​​the PIR sensor, the infrared radiation is enhanced by the Fresnel lens and focused on the PIR sensor. The PIR sensor can then sense this radiation and output a trigger signal. When the device is in sleep mode, it uses the detected trigger signal to determine whether an object has entered and wakes up the detection device to perform the corresponding operation.

[0005] Since the PIR sensor can wake up the device by sensing infrared radiation, it is not robust to non-target objects or changes in ambient temperature, which can easily cause the device to be woken up by mistake and increase device power consumption. Summary of the Invention

[0006] The present application provides a method, system, and chip for waking up a device, which can reduce the frequency of false wake-ups and lower system power consumption.

[0007] In a first aspect, the present application provides a method for waking up a device, which may include:

[0008] acquiring a first image;

[0009] determining a motion region image in the first image according to the first image;

[0010] Inputting the motion region image into a first neural network to obtain a first output result;

[0011] When the first output result is used to indicate that there is a target object in the motion area image, the second neural network is used to determine that the motion area image contains the target object, and a trigger signal is generated, wherein the trigger signal is used to wake up the device, and the first neural network is cascaded to the second neural network.

[0012] Among them, the first image is any frame image in the video. Through the present application, the motion area image in the image frame can be determined first, and then when the target object is determined to exist in the motion area image based on the cascaded first neural network and the second neural network, a trigger signal will be generated to wake up the device to perform corresponding work, such as waking up the processor in the device to identify the target object in the image and other processing. It can be seen that the first neural network first makes a first judgment on whether there is a moving target object, which can filter out dynamic and static false detection problems in the scene. The second neural network confirms the result of the first neural network, and only wakes up the device when it is confirmed that there is a moving target object. Otherwise, the device is in a dormant state. The secondary confirmation of the second neural network can reduce the frequency of false alarms and reduce system power consumption. In addition, the first neural network and the second neural network are cascaded, and the input of the two-stage neural network only requires one frame of data, which can reduce the wake-up delay and increase the wake-up speed.

[0013] In a possible implementation of the first aspect, determining, by the second neural network, that the moving region image contains a target object includes:

[0014] Inputting feature data of the middle layer of the first neural network into the second neural network to determine whether the motion area image contains the target object, wherein the accuracy of the second neural network is greater than the accuracy of the first neural network.

[0015] It can be seen that the input of the second neural network comes from the middle layer of the first neural network, and the resolution of the feature data of the middle layer is smaller than the resolution of the original image, which can reduce the input cache of the second neural network and alleviate the cache pressure.

[0016] In a possible implementation of the first aspect, determining a motion region image in the first image according to the first image includes:

[0017] A motion region image in the first image is determined based on the first image and an average image, wherein the first image includes an image acquired at a first moment, and the average image is determined based on images acquired before the first moment.

[0018] It can be seen that when the present application perceives the changing target or area (i.e., the moving area image) of the image frame in the scene, the reference frame is not the previous frame data, but an "average background frame" (i.e., the average image) is introduced. It is understandable that if the previous frame image is used as the reference frame, it will be more sensitive to the noise and illumination changes between consecutive frames, and will continue to cause false alarms in scenes with frequent scene noise and illumination changes, thereby increasing system power consumption. The average image used as the reference frame can weaken the noise and illumination changes, thereby reducing false alarms.

[0019] In a possible implementation of the first aspect, determining the motion region image in the first image according to the first image and the average image includes:

[0020] Acquire the average image, and determine an average image corresponding to the first image based on the first image and the average image;

[0021] The motion region image in the first image is determined by performing inter-frame difference according to the first image and an average image corresponding to the first image.

[0022] It can be seen that the calculation of the current average image (i.e., the average image of the first image) depends only on the current image (i.e., the first image) and the original average image. The original average image can weaken the noise and illumination changes in the first image, reduce the motion caused by noise and illumination changes, and thus reduce the number of false alarms.

[0023] In a possible implementation of the first aspect, the average image is an image obtained by averaging images acquired before the first moment according to preset parameters.

[0024] The preset parameters can be adaptively adjusted according to the actual scene requirements. For example, for low-contrast scenes such as business, the sensitivity of motion detection can be improved by lowering the preset parameters.

[0025] In a possible implementation of the first aspect, after determining, by the second neural network, that the motion region image contains a target object and before generating the trigger signal, the method further includes:

[0026] It is determined that the moving direction of the target object in the moving area image is consistent with a preset direction, where the preset direction includes a moving direction set by a user.

[0027] It can be seen that the present application can perform targeted wake-up based on the characteristic motion state of the target object (i.e., the preset direction). For example, in the doorbell application scenario, assuming that the target object moving forward is the key focus object (such as the owner opening the door), and the target object moving backward (such as the owner leaving) or left and right (such as neighbors) is the secondary focus object, the system can focus on awakening the target object in a specific motion direction in the scene (such as forward motion), and filter out the awakening of target objects moving in other directions (such as backward motion, left and right motion). This can not only achieve accurate wake-up, but also reduce system power consumption.

[0028] In a possible implementation of the first aspect, determining that the moving direction of the target object in the moving area image is consistent with a preset direction includes:

[0029] Acquire key points of the target object in the motion area image;

[0030] determining a moving direction of the target object according to the key points;

[0031] It is determined that the movement direction is consistent with the movement direction.

[0032] It can be seen that the key points can reflect the motion state of the target object. The method of determining the motion direction of the target object based on the key points is easy to implement. Therefore, the wake-up mechanism based on the motion direction of the target object is easy to apply in actual scenarios.

[0033] In a possible implementation of the first aspect, determining the movement direction of the target object according to the key point includes:

[0034] The key point includes at least one of a first key point, a second key point, and a third key point, wherein the first key point is a point located on the head of the target object, the second key point is a point located on the left shoulder of the target object, and the third key point is a point located on the right shoulder of the target object;

[0035] Obtaining a first distance and a second distance, wherein the first distance is a perpendicular distance from the first key point to a first connecting line, the first connecting line is a connecting line between the second key point and the third key point, and the second distance is a distance between the second key point and the third key point;

[0036] The moving direction of the target object is determined according to the ratio of the first distance to the second distance.

[0037] It can be seen that the key point in this application may be skeletal data. By analyzing the changes in multiple skeletal data of the target object to comprehensively determine the movement direction of the target object, the accuracy of judging the movement direction can be improved, thereby reducing the number of false wake-ups and reducing system power consumption.

[0038] In a possible implementation of the first aspect, determining the movement direction of the target object according to the ratio of the first distance to the second distance includes:

[0039] If a ratio of the first distance to the second distance is less than a first threshold, it is determined that the moving direction of the target object is forward movement or backward movement.

[0040] It can be seen that the first distance can indicate the left-right movement trend of the target object, and the second distance can indicate the forward-backward movement trend of the target object. When the first distance is less than the second distance, it means that the forward-backward movement trend of the target object is less than the left-right movement trend, and the target object is more likely to move forward and backward. The first threshold is a value determined by this application based on a large amount of data to determine the direction of movement of the target object. Therefore, the forward or backward movement of the target object determined based on the first threshold is credible. The wake-up method based on forward movement or backward movement can reduce the number of false wake-ups and reduce system power consumption.

[0041] In a possible implementation of the first aspect, determining whether the moving direction of the target object is forward movement or backward movement includes:

[0042] Acquire a motion region image of N frames of images acquired before the first moment, where N is a positive integer greater than or equal to 1;

[0043] The moving direction of the target object is determined to be forward or backward according to the second distance and an average of the second distances corresponding to the key points of the target object in the moving area images of the N frames of images.

[0044] It can be seen that the present application estimates the forward and backward motion by the difference between the current frame (i.e., the first image) and the previous N frames (i.e., those captured before the first moment), which can improve the credibility. The wake-up method based on the forward and backward motion can reduce the number of false wake-ups and reduce system power consumption.

[0045] In a possible implementation of the first aspect, determining whether the moving direction of the target object is forward or backward based on the second distance and an average of the second distances corresponding to the moving region images of the N frames of images includes:

[0046] The absolute value of the difference between the second distance and the mean is greater than a second threshold, and the moving direction of the target object is forward; or

[0047] An absolute value of a difference between the second distance and the mean is less than or equal to a second threshold, and the moving direction of the target object is backward.

[0048] It can be seen that the second threshold is a value determined by this application based on a large amount of data to judge the forward or backward movement of the target object. Therefore, the forward or backward movement of the target object determined is credible. The wake-up method based on forward or backward movement can reduce the number of false wake-ups and reduce system power consumption.

[0049] In a possible implementation of the first aspect, determining the movement direction of the target object according to the ratio of the first distance to the second distance includes:

[0050] If a ratio of the first distance to the second distance is greater than or equal to a first threshold, it is determined that the moving direction of the target object is leftward or rightward.

[0051] It can be seen that the first distance can indicate the left-right movement trend of the target object, and the second distance can indicate the forward-backward movement trend of the target object. When the first distance is greater than the second distance, it means that the forward-backward movement trend of the target object is greater than the left-right movement trend, and the target object is more likely to move left-right. The first threshold is a value determined by this application based on a large amount of data to determine the direction of movement of the target object. Therefore, the leftward or rightward movement of the target object determined based on the first threshold is credible. The wake-up method based on leftward or rightward movement can reduce the number of false wake-ups and reduce system power consumption.

[0052] In a possible implementation of the first aspect, determining the movement direction of the target object according to the key point includes:

[0053] The key point includes the center point of the motion area image;

[0054] According to the fact that a change of the center point of the motion region image on the X-axis is greater than a change on the Y-axis, it is determined that the motion direction of the target object is leftward motion or rightward motion.

[0055] In a possible implementation of the first aspect, determining the movement direction of the target object according to the key point includes:

[0056] The key point includes the center point of the motion area image;

[0057] According to the fact that a change of the center point of the motion region image on the X-axis is less than or equal to a change on the Y-axis, it is determined that the motion direction of the target object is forward motion or backward motion.

[0058] It can be seen that the present application can determine the left-right movement or the forward-backward movement of the target object based on the change of the center point on the motion area image. The implementation is simple and feasible, and the execution speed can be improved.

[0059] In a second aspect, the present application provides a computing device, which includes a communication module and a processing module, wherein:

[0060] The communication module is used to acquire a first image;

[0061] The processing module is configured to determine a motion region image in the first image based on the first image;

[0062] The processing module is further configured to input the motion region image into a first neural network to obtain a first output result;

[0063] The processing module is further configured to, when the first output result indicates that a target object exists in the motion area image, determine through a second neural network that the motion area image contains a target object, and generate a trigger signal, wherein the trigger signal is used to wake up the device, and the first neural network is cascaded to the second neural network.

[0064] In a possible implementation manner of the second aspect, the processing module is specifically configured to:

[0065] Inputting feature data of the middle layer of the first neural network into the second neural network to determine whether the motion area image contains the target object, wherein the accuracy of the second neural network is greater than the accuracy of the first neural network.

[0066] In a possible implementation manner of the second aspect, the processing module is specifically configured to:

[0067] A motion region image in the first image is determined based on the first image and an average image, wherein the first image includes an image acquired at a first moment, and the average image is determined based on images acquired before the first moment.

[0068] In a possible implementation manner of the second aspect, the processing module is specifically configured to:

[0069] Acquire the average image, and determine an average image corresponding to the first image based on the first image and the average image;

[0070] The motion region image in the first image is determined by performing inter-frame difference according to the first image and an average image corresponding to the first image.

[0071] In a possible implementation manner of the second aspect, the processing module is further configured to:

[0072] It is determined that the moving direction of the target object in the moving area image is consistent with a preset direction, where the preset direction includes a moving direction set by a user.

[0073] In a possible implementation manner of the second aspect, the processing module is specifically configured to:

[0074] Acquire key points of the target object in the motion area image;

[0075] determining a moving direction of the target object according to the key points;

[0076] It is determined that the movement direction is consistent with the movement direction.

[0077] In a possible implementation manner of the second aspect, the processing module is specifically configured to:

[0078] The key point includes at least one of a first key point, a second key point, and a third key point, wherein the first key point is a point located on the head of the target object, the second key point is a point located on the left shoulder of the target object, and the third key point is a point located on the right shoulder of the target object;

[0079] Obtaining a first distance and a second distance, wherein the first distance is a perpendicular distance from the first key point to a first connecting line, the first connecting line is a connecting line between the second key point and the third key point, and the second distance is a distance between the second key point and the third key point;

[0080] The moving direction of the target object is determined according to the ratio of the first distance to the second distance.

[0081] In a possible implementation of the second aspect, the processing module is configured to:

[0082] If a ratio of the first distance to the second distance is less than a first threshold, it is determined that the moving direction of the target object is forward movement or backward movement.

[0083] In a possible implementation manner of the second aspect, the processing module is specifically configured to:

[0084] Acquire a motion region image of N frames of images acquired before the first moment, where N is a positive integer greater than or equal to 1;

[0085] The moving direction of the target object is determined to be forward or backward according to the second distance and an average of the second distances corresponding to the key points of the target object in the moving area images of the N frames of images.

[0086] In a possible implementation manner of the second aspect, the processing module is specifically configured to:

[0087] The absolute value of the difference between the second distance and the mean is greater than a second threshold, and the moving direction of the target object is forward; or

[0088] An absolute value of a difference between the second distance and the mean is less than or equal to a second threshold, and the moving direction of the target object is backward.

[0089] In a possible implementation manner of the second aspect, the processing module is specifically configured to:

[0090] If a ratio of the first distance to the second distance is greater than or equal to a first threshold, it is determined that the moving direction of the target object is leftward or rightward.

[0091] In a possible implementation manner of the second aspect, the processing module is specifically configured to:

[0092] The key point includes the center point of the motion area image;

[0093] According to the fact that a change of the center point of the motion region image on the X-axis is greater than a change on the Y-axis, it is determined that the motion direction of the target object is leftward motion or rightward motion.

[0094] In a possible implementation manner of the second aspect, the processing module is specifically configured to:

[0095] The key point includes the center point of the motion area image;

[0096] According to the fact that a change of the center point of the motion region image on the X-axis is less than or equal to a change on the Y-axis, it is determined that the motion direction of the target object is forward motion or backward motion.

[0097] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, and the processor is used to execute instructions stored in a memory, or to run a logic circuit, so that the communication device implements the method described in any one of the first aspect or the method described in any one of the second aspect.

[0098] In a possible implementation, the communication device further includes a communication interface, where the communication interface is used to receive and / or send data, and / or the communication interface is used to provide input and / or output for the processor.

[0099] In one possible implementation, the communication device further includes a memory for storing at least one of an instruction, a configuration file of a logic circuit, and data. Optionally, the processor and the memory may be integrated into one device.

[0100] The above embodiments are described using a processor (or general-purpose processor) that executes a method by calling a computer instruction. In specific implementations, the processor may also be a dedicated processor, in which case the computer instructions are pre-loaded into the processor. Alternatively, the processor may include both a dedicated processor and a general-purpose processor.

[0101] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores instructions. When the instructions are executed on at least one processor, the method described in any one of the first aspect or the method described in any one of the second aspect is implemented.

[0102] In a fifth aspect, the present application provides a computer program product, which includes computer instructions. When the instructions are executed on at least one processor, the method described in any one of the first aspect or the method described in any one of the second aspect is implemented.

[0103] Optionally, the computer program product may be a software installation package or an image package. When the aforementioned method is required, the computer program product may be downloaded and executed on a computing device.

[0104] In the sixth aspect, the present application provides a chip system, which includes at least one processor, a memory and an interface circuit, wherein the memory, the interface circuit and the at least one processor are interconnected through lines, and a computer program is stored in the at least one memory; when the computer program is executed by the processor, the method described in any one of the first aspect or the method described in any one of the second aspect is implemented.

[0105] In a seventh aspect, the present application provides a communication system, which includes the communication device described in the fourth aspect and the communication device described in the fifth aspect.

[0106] The beneficial effects of the technical solutions provided in the second to seventh aspects of this application can refer to the beneficial effects of the technical solution in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0107] FIG1A is a schematic diagram showing a scenario of waking up a device based on PIR;

[0108] FIG1B is a schematic diagram of the architecture of low-power wake-up using PIR combined with motion detection;

[0109] Figure 1C shows a schematic diagram of the architecture of motion detection combined with CNN low-power wake-up;

[0110] FIG2 is a schematic diagram showing the structure of a wake-up system provided in an embodiment of the present application;

[0111] FIG3 is a flow chart of a method for waking up a device according to an embodiment of the present application;

[0112] FIG4A is a schematic diagram of a process of generating a motion region image in a first image according to an embodiment of the present application;

[0113] FIG4B is a schematic diagram of a process for determining an average image according to an embodiment of the present application;

[0114] FIG4C is a schematic diagram of key points of a target object provided in an embodiment of the present application;

[0115] FIG5 is a schematic diagram of a wake-up architecture provided in an embodiment of the present application;

[0116] FIG6 is a schematic diagram of another wake-up architecture provided in an embodiment of the present application;

[0117] FIG7 is a schematic structural diagram of a computing device 70 provided in an embodiment of the present application;

[0118] FIG8 is a schematic structural diagram of an electronic device 80 provided in an embodiment of the present application. DETAILED DESCRIPTION

[0119] The following describes the embodiments of the present application in conjunction with the accompanying drawings. In the description of the embodiments of the present application, unless otherwise specified, " / " represents or, for example, A / B can represent A or B; "and / or" in the text is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, "multiple" refers to two or more than two.

[0120] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of the technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of this application, unless otherwise specified, "plurality" means two or more.

[0121] The following is a detailed analysis of the technical problems that need to be solved by the embodiments of the present application and the corresponding application scenarios.

[0122] Identifying target objects or scenes is a fundamental research topic in the field of deep learning, and many deep learning optimization techniques are based on image recognition. Currently, low-power target wake-up technology based on deep learning has been successfully applied in many fields such as security monitoring, doorbell access control, and autonomous driving. The biggest problem facing low-power wake-up technology based on deep learning is the huge demand for computing and storage resources. Many network miniaturization technologies such as network pruning, network sparseness, and low-bit quantization have emerged, making it possible for deep learning-based applications to be implemented. However, in the field of low-power end-side devices, especially in the field of low-power chip design at the milliwatt and microwatt levels, existing neural network optimization technologies still cannot meet the computing resources and perception accuracy requirements in actual application scenarios.

[0123] 1B , which shows a schematic diagram of a low-power wakeup architecture for PIR combined with motion detection. The low-power wakeup architecture 100 shown in FIG1B includes at least one of an image detection module 1001 , a primary wakeup module 1002 , a secondary wakeup module 1003 , and a processor 1004 .

[0124] Image detection module 1001 obtains a region of interest (ROI) from the captured image (e.g., the i-th frame image or the i+1-th frame image, where i is a positive integer) and inputs the ROI into primary wake-up module 1002. The ROI is the area to be processed, outlined in the processed image using a box, circle, ellipse, or irregular polygon. The ROI is typically a pre-defined area based on the range to be detected, such as an entrance, exit, or corridor.

[0125] The primary wake-up module 1002, acting as a primary wake-up unit, detects the region of interest (ROI) based on at least one of a low-frame-rate smart motion detection (SMD) and a PIR sensor to detect whether there is a moving object or a thermally moving object in the ROI. When the primary wake-up module 1002 detects at least one of a moving object or a thermally moving object, an interrupt is triggered, prompting the secondary wake-up module 1003 to perform further detection.

[0126] Among them, SMD can also be called motion detection, which detects the movement of the target object by comparing the pixel differences between adjacent frames through the frame difference method.

[0127] The secondary wake-up module 1003, acting as a secondary wake-up unit, further determines whether a target object (e.g., a humanoid object) is present in the captured image based on the high frame rate SMD. Upon detecting a target object, the secondary wake-up module 1003 wakes up the processor 1004 for appropriate processing. Alternatively, if the secondary wake-up module 1003 fails to detect a target object within a preset time, it returns to the primary wake-up mode.

[0128] It can be seen that the secondary wake-up module 1003 can filter the results from the primary wake-up module 1002, thereby reducing the frequency of false wake-ups and saving power consumption.

[0129] 1C , which shows a schematic diagram of a low-power wakeup architecture for motion detection combined with CNN. The low-power wakeup architecture 101 shown in FIG1C includes at least one of a primary wakeup module 1011 , a secondary wakeup module 1012 , and a processor 1013 .

[0130] The first-level wake-up module 1011 is used to perform motion detection on the captured video images (for example, the i-th frame image, the i+1-th frame image). When a motion area is detected (for example, an area where there is a difference between the previous and next frames), the first-level wake-up is triggered, and the second-level wake-up module 1012 performs further detection.

[0131] Secondary wakeup module 1012 is used to further detect the motion region and generate a motion region event, which is then fed into the neural network. The neural network then further determines whether a target object (e.g., a humanoid object) is present within the motion region event. If a target object is detected, secondary wakeup module 1012 wakes up processor 1004 to perform appropriate processing.

[0132] It can be seen that the secondary wake-up module 1003 can filter the results from the primary wake-up module 1002 to improve the end-to-end wake-up accuracy.

[0133] In summary, the architectures shown in Figures 1B and 1C have the following problems:

[0134] 1. Uncertainty in ROI selection. The method shown in Figure 1B typically manually sets or uses a specific algorithm to detect the ROI region where the target object is located in the full scene image. However, in some doorbell access control scenarios, the target object (such as a human figure) is often close to the acquisition device (such as a camera). The manually or algorithmically set ROI region can easily cause the target object to be truncated, thereby affecting the target perception accuracy.

[0135] 2. The frequency of false awakenings in motion detection is high. The method shown in Figures 1B and 1C uses the frame difference method to sense whether there is a pixel difference between the previous and next frames in the captured image, and uses this as a first-level awakening. However, the frame difference method can only sense the area where there is a pixel jump in the image, and cannot confirm whether the pixel jump is caused by the movement of the target object. Therefore, the misjudgment area determined by the pixel difference between the previous and next frames caused by the movement of non-target objects or changes in illumination may bring computing and storage pressure to the subsequent detection process, and frequent first-level false alarms or triggering of second-level awakenings will increase system power consumption.

[0136] Understandably, the frame difference method is sensitive to noise and subtle changes in illumination between consecutive frames. As for noise, since the frame difference method detects targets based on pixel differences, even the presence of noise in the scene can be mistaken for motion of the target object. This can lead to persistent false alarms, increasing the system's power consumption and processing burden. Furthermore, the frame difference method is sensitive to subtle changes in illumination. When the illumination in the scene changes slightly, the grayscale value of the pixel also changes slightly, which can be mistaken for motion of the target object. Similarly, this can lead to persistent false alarms, increasing the system's power consumption and processing burden.

[0137] 3. Slow wakeup. Using SMD for wakeup often requires at least two frames of images. In scenes where the target object is at the edge or in rapid motion, this can lead to missed wakeups due to the target object leaving the scene. Furthermore, SMD can easily misjudge the motion of non-target objects, resulting in a large number of misjudgment areas and increased storage overhead.

[0138] In view of this, the present application provides a method, system, and chip for waking up a device. The method may include: processing a captured first image, determining a motion area image in the first image, and determining whether a target object exists in the motion area image through a first-level neural network. When the target object is confirmed to exist in the motion area image, a second neural network cascaded with the first-level neural network is used to perform a second confirmation of the output of the first-level neural network. It can be seen that judging the target object through the first-level neural network and the second-level neural network can reduce the frequency of false alarms and reduce system power consumption.

[0139] Below, the parts involved in the neural network are explained for the purpose of understanding by those skilled in the art.

[0140] (1) Deep Neural Networks (DNN) is a broad concept that includes, in a sense, convolutional neural networks (CNN), recurrent neural networks (RNN), and generative adversarial networks (GAN). DNN refers to a neural network that contains multiple hidden layers. The neural network provided in the embodiments of the present application may include a convolutional neural network.

[0141] (2) Convolutional neural network (CNN) is a multi-layer neural network. Each layer consists of multiple two-dimensional planes, and each plane is composed of multiple independent neurons. Multiple neurons in each plane share weights. Weight sharing can reduce the number of parameters in the neural network. Currently, in convolutional neural networks, the processor usually performs convolution operations by converting the convolution of input signal features and weights into a matrix multiplication operation between the signal matrix and the weights.

[0142] (3) The convolution kernel can be initialized in the form of a matrix of random size, and the convolution kernel can obtain reasonable weights through learning during the training process of the convolutional neural network. In addition, the direct benefit of shared weights is to reduce the number of connections between the layers of the convolutional neural network, while reducing the risk of overfitting.

[0143] (4) A filter is a series connection of multiple convolution kernels, each of which is assigned to a specific channel of the input. When the number of channels is 1, the filter is the convolution kernel. When the number of channels is greater than 1, the filter refers to the series connection of multiple convolution kernels. For example, if an image is stored as a tensor in RGB format, the input includes three channels, namely the R matrix, the G matrix, and the B matrix (red, green, and blue, corresponding to three images of the same size). The matrix of each channel is convolved with a corresponding convolution kernel, and all the convolution kernels corresponding to all channels form a filter. Each filter is used to extract different feature data. For example, consider an image with four ARGB channels (transparency and red, green, and blue, corresponding to four images of equal size). Assuming a convolution kernel size of 100*100, 16 convolution kernels w1 to w16 are used. Kernels w1 to w4 form the first filter, w5 to w8 form the second, w9 to w12 form the third, and w13 to w6 form the fourth. Different filters are used to extract different features of the input image. Convolution of the ARGB image with the first filter—that is, convolution of w1 to w4 with the four images corresponding to the four channels—yields the first image. The first pixel in the upper left corner of this image is the weighted sum of the pixels within the 100*100 area in the upper left corner of the four input images, and so on. Similarly, including the other filters, the output of this layer corresponds to four "images." Each image pair responds to a different feature in the original image.

[0144] (5) The convolutional neural network can use the back propagation (BP) algorithm to correct the size of the parameters in the initial neural network model during the training process, so that the reconstruction error loss of the neural network model becomes smaller and smaller. Specifically, the forward transmission of the input signal to the output will generate error loss, and the parameters in the initial neural network model are updated by back propagating the error loss information, so that the error loss converges. The back propagation algorithm is a back propagation movement dominated by error loss, which aims to obtain the optimal parameters of the super-resolution model, such as the weights and attention vectors in the embodiments of the present application.

[0145] (6) Convolution is the extraction of feature data on the original input. In short, the extraction of feature data is the extraction of features in a small area on the original input. Expressed in mathematical terms, convolution is the operation of the convolution kernel and the input matrix of the convolution layer. Usually, the input matrix is ​​the matrix extracted from the image matrix according to the stride of the convolution kernel during convolution. The convolution kernel is a small window that records the weights. The convolution kernel slides on the image matrix according to the stride. Each time the convolution kernel slides, it corresponds to a sub-matrix of the image matrix. The weight in the convolution kernel is multiplied and added by the value contained in the sub-matrix, and assigned to an element corresponding to the current output feature map (output matrix) of the convolution kernel. Among them, convolution is not limited to convolution of the original input, but also includes the use of convolution on the output result after convolution. The embodiment of the present application does not limit this. For example, the first convolution extracts low-level feature data, the second convolution extracts mid-level feature data, the third convolution extracts high-level features, and so on. Features can be continuously extracted and compressed. The higher-level features finally extracted can be understood as a further concentration of the original features, making the final features more reliable. The last layer of features can be used to process various tasks, such as classification and regression.

[0146] Next, the application scenarios of the embodiments of the present application are introduced.

[0147] FIG2 is a schematic diagram of the structure of the wake-up system provided in an embodiment of the present application. The wake-up system 20 includes an acquisition device 201, a data processing device 202, a working device 203, and a storage device 204.

[0148] The acquisition device 201 may be a camera for acquiring multiple images of the surrounding environment. The camera may be a still camera or a video camera (i.e., a video camera), a visible light camera, or an infrared camera, and may be any camera for acquiring images, and the present embodiment does not limit this.

[0149] The data processing device 202 is used to process the image to be identified acquired by the acquisition device 201 to identify the motion area therein, and then determine whether there is a target object in the motion area.

[0150] In one implementation, data processing device 202 includes a primary wake-up module 2021 and a secondary wake-up module 2022. Primary wake-up module 2021, based on intelligent motion detection (SMD), can determine the motion region image in the image to be identified. It then uses a primary CNN to determine whether a target object exists in the motion region image. If the target object is confirmed to be present in the motion region image, secondary wake-up module 2022 is awakened. Secondary wake-up module 2022 verifies the output of primary wake-up module 2021 based on a secondary CNN cascaded with the primary CNN. If the secondary CNN output indicates the presence of a target object in the motion region image, it awakens working device 203. As can be seen, using a secondary wake-up architecture for judgment can reduce the frequency of false positives and lower end-to-end system power consumption.

[0151] In one implementation, when the secondary CNN output indicates the presence of a target object in the moving area image, the secondary wake-up module 2022 is further configured to determine the direction of motion of the target object in the moving area image and, when the target object's motion direction is consistent with a preset direction, generate a trigger signal to wake up the working device 203 to perform the corresponding operation. The preset direction includes a user-set direction of interest, including but not limited to forward motion, backward motion, leftward motion, and leftward motion.

[0152] The target object may be set by the user through a user device (not shown in FIG2 ), and the target object may be an object that the user wants to detect, such as a human-shaped target, a vehicle, an animal, and so on.

[0153] The data processing device 202 can be of various types, such as a cloud server, a network server, an application server, or a management server, and can also be a device or server with data processing capabilities, such as a chip, a software module, or an integrated circuit. The data processing device 202 receives a detection request from a user device (not shown in FIG. 2 ) via an interactive interface, and then performs detection processing of target objects in a moving area using methods such as machine learning, deep learning, search, reasoning, and decision-making through a memory that stores data and a data processing link. The memory in the data processing device is a general term that includes local storage and a database that stores historical data. The database can be on the data processing device or on another network server.

[0154] The working device 203 is used to wake up the processor to perform corresponding work when the data processing device 202 determines that there is a moving target object, such as displaying the moving target object identified by the data processing device 202, and / or generating alarm information, which is used to indicate that a moving target object has been detected.

[0155] The storage device 204 is used to store the motion region and the target object in the motion region in each frame of image identified by the data processing device 202 .

[0156] For example, the data acquisition device 201 and the data processing device 202 may be integrated into a single device, for example, the data acquisition device 201 and the data processing device 202 may be integrated into the same security monitoring device or the same vehicle. The data acquisition device 201 and the data processing device 202 may also be provided separately, for example, the data acquisition device 201 and the data processing device 202 may be provided separately as a camera and a server.

[0157] Exemplarily, the acquisition device 201 and the data processing device 202 can be directly connected in communication. For example, when the acquisition device 201 and the data processing device 202 are integrated into the same device (, the acquisition device 201 and the data processing device 202 can be directly connected through corresponding connecting devices. The acquisition device 201 and the data processing device 202 can be indirectly connected in communication. For example, when the acquisition device 201 and the data processing device 202 are separately provided, the acquisition device 201 and the data processing device 202 can be indirectly connected in communication through wireless communication or other means.

[0158] The wake-up system shown in Figure 2 can be applied to a variety of scenarios. These scenarios are described below using a camera as an example of acquisition device 201. In this scenario, the camera is typically positioned in a fixed position to ensure that the non-moving areas of each frame of the captured video remain essentially the same, without noticeable changes. Three of these scenarios are described below, but this application is not limited to these three scenarios.

[0159] Scenario 1: The wake-up system shown in Figure 2 is used in indoor or outdoor security monitoring scenarios. To protect personal and property safety in places like homes, schools, and construction sites, surveillance equipment is installed in at least one location, such as a hallway, doorway, or room. The equipment then captures and displays surveillance footage. The "surveillance footage" refers to the image displayed on the monitor after the surveillance equipment captures the scene.

[0160] In this scenario, the monitoring device can identify the moving area image in each frame of the video by waking up the system, and use this moving area image to search for the target object (such as a human or animal). If a human or animal is found in the moving area, the processor can be woken up to perform corresponding tasks, such as identifying the target object and displaying an alarm message.

[0161] Scenario 2: The wake-up system shown in Figure 2 is used in a traffic monitoring scenario. At some highway intersections, gates, or crossroads, road monitoring equipment is usually installed to monitor and adjust vehicle traffic flow. The equipment is used to collect and display monitoring images.

[0162] In this scenario, the monitoring device wakes up the system to identify the moving area images in each frame of the video, and uses this moving area image to find the target object (such as a vehicle). If a large number of vehicles are detected in a certain direction, the processor can be woken up to perform corresponding tasks, such as prompting relevant personnel to control and adjust the traffic lights.

[0163] Scenario 3: The wake-up system shown in Figure 2 is applied to photography. Taking the wake-up system application in a mobile phone as an example, in one scenario, when a user uses a mobile phone to shoot, to improve the shooting effect, the mobile phone can analyze the moving area image in the shooting picture based on intelligent motion detection. When a target is identified in the moving area image, the processor is woken up to perform corresponding operations, such as marking and displaying the identified target object (such as a moving puppy).

[0164] Please refer to FIG3 , which is a flowchart of a method for waking up a device provided in an embodiment of the present application. The method can be applied to the system shown in FIG2 , and the method includes but is not limited to the following steps:

[0165] Step S301: Acquire a first image.

[0166] Specifically, the electronic device can obtain each frame of image in the video captured by the acquisition device (such as a camera) of the target scene (such as a security monitoring scene, a traffic monitoring scene, a shooting scene, an intelligent driving scene, etc.). The first image is a frame of image in the video, such as an image captured at the first moment.

[0167] It should be noted that the electronic device can be a device with communication and computing capabilities. In different scenarios, the electronic device can be different devices, such as smart cameras, surveillance doorbells, smart door locks, vehicles, etc.

[0168] Step S302: determining a motion region image in the first image according to the first image.

[0169] It is understood that consecutive frames in a captured video exhibit continuity. If there are no moving objects in the target scene, the changes between consecutive frames are subtle. However, if there are moving objects, there will be noticeable changes between consecutive frames. Because the objects in the target scene are in motion, the positions of the objects' images vary between frames. Therefore, the motion region in the first image represents the region of noticeable change between the first image and the other images caused by the moving objects.

[0170] Please refer to Figure 4A, which is a flow chart of a motion area image in a first image provided by an embodiment of the present application. As shown in Figure 4A, the electronic device performs Gaussian blur processing on the first image to obtain a processed first image. The resolution of the first image can be 80×64. The electronic device subtracts the processed first image from the average image to obtain a differential map (activation) of the candidate motion area between frames, and takes the average value of the first image and the average image to update the average image. Then, the electronic device binarizes the differential map (activation) to obtain a binary map (Mask), and performs morphological processing on the binary map (Mask), such as dilation (Dilate) and corrosion (Erode), to obtain a morphologically processed binary map (Mask), so as to obtain a complete and accurate Mask of the motion target area. Next, the electronic device performs grid processing on the morphologically processed binary map (Mask) to obtain a grid binary map (Grid Mask), and the resolution of the grid binary map is 16×8. Finally, the electronic device performs connected component analysis (CCA) on the grid binary image to obtain bounding boxes, wherein the image occupied by the bounding boxes in the first image is the motion area image.

[0171] In one possible implementation, the electronic device determines the motion region image in the first image using an inter-frame difference method. Unlike conventional inter-frame difference methods, except for the first frame in the video, the reference frame used in the embodiment of the present application does not directly extract the previous frame image of the current frame, but instead introduces the probability of an "average background frame" (i.e., an average image). For example, the average image is an image obtained by averaging images captured before the first moment.

[0172] Please refer to Figure 4B, which is a flowchart of determining an average image provided by an embodiment of the present application. As shown in Figure 4A, the i-th (i=0) frame image is the first frame image in the video captured after the acquisition device is powered on. Because there are no other images in the video before the i-th (i=0) frame image, the i-th (i=0) frame image can be directly considered as the average image. For the i+1-th frame image after the i-th (i=0) frame image, it can be averaged with the average image first, and the average value of the i+1-th frame image and the average image is taken to update the average image. Therefore, except for the first frame after power-on, the average value can be calculated to update the average image when each subsequent frame image is input. Furthermore, the calculation of the average image of the current image depends only on the current image and the original average image.

[0173] In one implementation, the electronic device obtains an average image, determines an average image corresponding to the first image based on the first image and the average image, and then determines a motion region image in the first image using an inter-frame difference method based on the first image and the average image corresponding to the first image. Specifically, a difference operation is performed on the first image and the average image of the first image, and the corresponding pixels of the first image and the average image of the first image are subtracted to determine the absolute value of the grayscale difference. When the absolute value of the grayscale difference corresponding to a region image exceeds a certain threshold, the region image is determined to be a motion region image.

[0174] As can be seen from Figure 4B, for the first image captured at the first moment, before the first image is input, the average image is the average value of the images captured before the first moment, and the average image of the first image is related to the first image and the average image before the first moment.

[0175] Exemplarily, avg_frm(i+1)=α*avg_frm(i)+(1-α)*cur_frm(i+1), where i is a positive integer, avg_frm(i+1) is the average image of the current image, avg_frm(i) is the original average image, cur_frm is the current image, α is a preset parameter, and α is used to represent the update parameter of the average image. In one implementation, the α value can be determined according to the usage scenario. For example, in low-contrast scenes such as at night, the sensitivity of motion detection can be improved by lowering the α value. The traditional inter-frame difference method is more sensitive to noise and weak illumination changes between consecutive frames, and will continue to generate false alarms in scenes with frequent scene noise and illumination changes, thereby increasing system power consumption. In an embodiment of the present application, false alarms can be reduced by constructing an average image to weaken noise and illumination changes.

[0176] Step S303: input the motion area image into the first neural network to obtain a first output result.

[0177] Specifically, the first neural network can be a convolutional neural network, and is trained based on sample data. When the electronic device recognizes an object in the moving area image using the first neural network, it can directly calculate based on the trained model parameters to obtain a first output result. It is understandable that there may be one or more objects in the moving area image, the target object is one of the one or more objects, and the sample data is sample data containing the target object. Therefore, the electronic device can recognize the target object in the moving area image based on the first neural network trained based on the sample data.

[0178] Step S304 : When the first output result indicates that the target object exists in the motion region image, the second neural network is used to determine that the motion region image contains the target object, and a trigger signal is generated.

[0179] It can be seen that through step S302 and step S303, the moving target object in the target scene can be detected, thereby filtering out the dynamic (such as fluttering curtains, leaves, etc.) and static (such as statues, posters, etc.) false detection problems in the target scene. In order to avoid false awakening caused by false detection and missed detection, when the first output result is used to indicate that there is a target object in the moving area image, the present application confirms the first output result of the first neural network output through the second neural network, and when the output result of the second neural network is also used to indicate that there is a target object in the moving area image, a trigger signal is generated to wake up the processor for corresponding processing. When the output result of the second neural network is used to indicate that there is no target object in the moving area image, no trigger signal is generated and the processor will not be awakened for corresponding processing.

[0180] Among them, the accuracy of the first neural network is lower than that of the second neural network. Therefore, the first neural network can be used for preliminary judgment, and the second neural network can be used for secondary judgment, thereby improving the accuracy of judgment, reducing the number of false wake-ups, and reducing system power consumption.

[0181] In one possible implementation, the first neural network is cascaded with the second neural network, and the electronic device inputs the intermediate layer feature data of the first neural network into the second neural network, and uses the second neural network to identify whether there is a target object in the motion area image. Therefore, in the present application, wake-up can be achieved by inputting a single-frame image (that is, the first image). Compared with the wake-up that relies on at least two frames of images and the wake-up that relies on at least four frames of images, the wake-up that relies on a single-frame image in the present application can reduce the wake-up delay and increase the wake-up speed.

[0182] As can be seen, the input of the second neural network does not rely on the original first image, but instead uses the feature data of the intermediate layer of the first neural network as input. For example, the feature data output to the second neural network can come from the intermediate layer of the first neural network that has been downsampled eightfold. This can reduce the input buffer of the second neural network and alleviate buffer pressure.

[0183] In one possible implementation, in order to achieve accurate wake-up and reduce system power consumption, wake-up can be achieved in the direction of motion that the user is interested in. Therefore, after the electronic device determines that the target object is included in the motion area image through the second neural network, it determines the motion direction of the target object in the motion area image, and generates a trigger signal when the motion direction belongs to the motion direction that the user is interested in, to wake up the processor in the electronic device to perform corresponding work. For example. If the motion direction that the user is interested in is forward motion, and the motion direction of the target object in the motion area image is forward motion, it is consistent with the motion direction that the user is interested in, then a trigger signal is generated to wake up the processor; if the motion direction of the target object in the motion area image is at least one of left motion, right motion or forward motion, and is inconsistent with the motion direction that the user is interested in, then the processor will not be woken up, and the system can continue to remain in sleep mode.

[0184] In a possible implementation, the electronic device obtains key points of the target object in the moving area image, and then determines the moving direction of the target object according to the movement trend of the key points.

[0185] In one implementation, the key point includes the center point of the moving area image, and the electronic device estimates the target object's direction of motion based on the changing trends of the center point on the X-axis and Y-axis. It will be appreciated that the computing device stores the moving area image before the first moment, and therefore, based on the changes in the center point of the moving area image before the first moment on the X-axis and Y-axis, the changing trends of the center point of the moving area image at the first moment on the X-axis and Y-axis can be estimated.

[0186] For example, the change of the center point of the motion area image on the X-axis is greater than the change on the Y-axis, indicating that in the time domain, the movement of the center point on the X-axis is significantly reduced or increased, and the movement on the Y-axis is a small fluctuation. The electronic device can determine whether the movement direction of the target object is to the left or to the right.

[0187] For example, the change of the center point of the motion area image on the X-axis is less than or equal to the change on the Y-axis, indicating that in the time domain, the movement of the center point on the X-axis is a small fluctuation, and the movement on the Y-axis is a significant decrease or increase. The electronic device can determine whether the movement direction of the target object is forward or backward.

[0188] In one implementation, the key point includes at least one of a first key point, a second key point, and a third key point. The key point may be an imaging coordinate point of a skeletal point, where the imaging coordinate point is the coordinate of the skeletal point in the motion region image. Referring to FIG. 4C , FIG. 4C is a schematic diagram of key points of a target object provided in an embodiment of the present application. As shown in FIG. 4C , the first key point is located on the target object's head, the second key point is located on the target object's left shoulder, and the third key point is located on the target object's right shoulder.

[0189] In one possible implementation, the electronic device may predict the direction of motion of the target object based on the key points of the target object in the motion region image, specifically using at least one of the first key point, the second key point, and the third key point to comprehensively determine the direction of motion of the target object. Further, the electronic device obtains a first distance L1 and a second distance L2, wherein, as shown in FIG4C , the first distance L1 is the vertical distance from the first key point to the first connecting line, the first connecting line is the connecting line between the second key point and the third key point, and the second distance L2 is the distance between the second key point and the third key point. The electronic device may then determine the direction of motion of the target object based on the ratio of the first distance L1 to the second distance L2 (r=L1 / L2). For example, the electronic device may determine whether the target object is facing the acquisition device (such as a camera) or facing the acquisition device sideways based on the ratio of the first distance L1 to the second distance L2 (r=L1 / L2), and determine the direction of motion of the target object based on this.

[0190] In one implementation, when the ratio of the first distance to the second distance is less than a first threshold (for example, the first threshold = 0.8), it can be said that the target object is facing the camera, and the electronic device can determine whether the target object is moving forward or backward.

[0191] Exemplarily, for an image captured before the first moment, if it is determined that a motion area image exists in the image, the electronic device may store it. Therefore, the electronic device may obtain the motion area image of N frames of images captured before the first moment, calculate the second distance between the second key point and the third key point of the target object in the motion area image of each frame of the N frames, and calculate the average of the above N second distances, where N is a positive integer greater than or equal to 1. Then, the electronic device determines whether the target object's motion direction is forward or backward based on the second distance (the distance between the second key point and the second key point of the target object in the motion area image of the first image captured at the first moment) and the average of the second distances corresponding to the key points of the target object in the motion area image of the N frames.

[0192] For example, the electronic device may estimate the forward and backward motion of the target object based on the difference between the second distance and the above average. When the difference is greater than zero and the absolute value of the difference is greater than a second threshold (for example, the second threshold = second distance * 0.05), the electronic device determines that the target object is moving forward. When the absolute value of the difference is less than or equal to the second threshold, the electronic device determines that the target object is moving backward.

[0193] In one implementation, when the ratio of the first distance to the second distance is greater than or equal to a first threshold (e.g., first threshold = 0.8), it can be indicated that the target object is facing the camera, and the electronic device can determine whether the target object is moving to the left or right. For example, the electronic device can estimate the left and right movement of the target object based on the change of the center point of the moving area image on the X-axis and Y-axis. If the change of the center point on the X-axis is greater than the change on the Y-axis, the target object's movement direction is left or right.

[0194] Please refer to Figure 5, which is a schematic diagram of a wake-up architecture provided in an embodiment of the present application, which is applied to the system shown in Figure 2. As shown in Figure 5, the wake-up architecture 50 includes at least one of a primary wake-up module 501, a secondary wake-up module 502, and a processor 503. The primary wake-up module 501 includes a motion detection module 5011 and a first neural network 5012, and the secondary wake-up module 502 includes a second neural network 5021. The first neural network 5012 and the second neural network 5021 are cascaded.

[0195] As can be seen from FIG5 , the electronic device can sense the changing targets or regions between consecutive frames in the target scene through the motion detection module 5011, thereby obtaining images of the moving region in the consecutive frames. The electronic device specifically obtains the moving region images using the inter-frame difference method based on the average frame provided in this application. A detailed description can be found in FIG4A and FIG4B , and will not be repeated here.

[0196] As shown in FIG5 , the motion region image input to the first neural network 5012 includes the region obtained by clipping the bounding box in FIG4A on the original resolution image (for example, the resolution is 64×64) and the mask image of the motion region image (for example, the resolution is 64×64). In one implementation, the motion region image input to the first neural network 5012 is a mask image superimposed on the above region. It can be seen that in order to introduce the position information of the target object to improve the classification accuracy, the present application adds a mask image channel corresponding to the motion region image to the input of the first neural network 5012.

[0197] In the first neural network 5012 shown in FIG5 , the moving region image is analyzed to obtain a first output result. If the first output result indicates that the target object is not present in the moving region image, it indicates that the first neural network 5012 has detected no target object motion in the current image frame. Therefore, the secondary wake-up module 502 does not perform further processing on the image frame, and therefore does not wake up the processor 503 to perform corresponding operations.

[0198] As shown in Figure 5, when the first output indicates the presence of a target object in the moving area image, the secondary wake-up module 502 is activated. By cascading two neural networks, both neural networks only require a single frame of data as input. The second neural network's input does not rely on the original image frame data, but instead uses the feature data from the intermediate layers of the first neural network as input. In one implementation, the feature data extracted from the first neural network's eightfold downsampling of the intermediate layers can reduce the input buffer of the secondary neural network to 40% of the original resolution (for example, 160×64).

[0199] Therefore, the second neural network 5021 analyzes the input feature data and generates a second output result. When the second output result indicates that the target object is not present in the motion region image, it indicates that the second neural network 5021 has detected no motion of the target object in the current image frame and, therefore, does not wake up the processor 503 to perform corresponding operations. It is understood that the second neural network 5021 has a higher precision than the first neural network 5012, and therefore the recognition accuracy of the second neural network 5021 is higher than that of the first neural network 5012, and the result output by the second neural network 5021 is credible.

[0200] As shown in FIG5 , when the second output result is used to indicate that a target object exists in the motion region image, the secondary awakening module 502 generates a trigger signal to awaken the processor 503 to perform corresponding work.

[0201] Please refer to Figure 6, which is a schematic diagram of another wake-up architecture provided in an embodiment of the present application, which is applied to the system shown in Figure 2. As shown in Figure 6, wake-up architecture 60 includes at least one of a primary wake-up module 601, a secondary wake-up module 602, and a processor 603. Primary wake-up module 601 includes a motion detection module 6011 and a first neural network 6012, and secondary wake-up module 602 includes a second neural network 6021 and a motion direction module 6022. The first neural network 6012 and the second neural network 6021 are cascaded.

[0202] For the description of the primary awakening module 601 in FIG6 and the description of the second neural network 6021 of the secondary awakening module 602, reference can be made to the relevant description of the primary awakening module 501 and the second neural network 5021 of the secondary awakening module 502 in FIG5, which will not be repeated here.

[0203] Motion direction module 6022 shown in FIG6 is used to identify the motion direction of a moving target object. Specifically, it is used to determine the key points of the target object through a key point detection network and further used to determine the target object's motion method based on the key points. For a description of "identifying the motion direction of a moving target object," please refer to step S304 in FIG3 and will not be repeated here.

[0204] When the motion direction of the target object determined by the motion direction module 6022 is inconsistent with the motion direction of the user, it means that the motion direction is not the direction of interest to the user, and there is no need to wake up the processor 603 to perform corresponding work.

[0205] As can be seen from Figure 6, when the movement direction of the target object determined by the movement direction module 6022 is consistent with the movement direction of the user, it means that the movement direction is the direction of interest to the user, and the secondary wake-up module 602 generates a trigger signal to wake up the processor 603 to perform corresponding work.

[0206] It is understandable that the introduction of the motion direction module 6022 in the wake-up architecture 60 can accurately identify the motion direction of the target object in the scene, thereby guiding the system to wake up in the direction of interest to the user and filtering out wake-ups in directions that the user does not care about, allowing the user to flexibly configure the wake-up state of the target object and also saving system power consumption to a large extent. For example, for a common smart doorbell system, when someone passes by, it will be awakened to perform the corresponding work, and it is impossible to identify whether the state of the moving target object is moving left and right or forward and backward. If the user only needs to pay attention to the awakening of a human-shaped target moving forward, then left and right movement and backward movement are irrelevant directions, and the system can continue to remain in sleep mode, theoretically saving 76% of power consumption.

[0207] The method of the embodiment of the present application is described above, and the device of the embodiment of the present application is provided below.

[0208] Please refer to Figure 7, which is a schematic diagram of the structure of a computing device 70 provided in an embodiment of the present application. The computing device 70 may include a communication module 701 and a processing module 702. The details of each module are as follows:

[0209] The communication module 701 can implement corresponding communication functions, and the processing module 702 is used to process data. The communication module 701 can also be called a communication interface or a transceiver unit. Optionally, the processing module 702 can be implemented by at least one processor or processor-related circuit.

[0210] Optionally, the communication module 701 may further include a storage module, which may be used to store instructions and / or data. The processing module 702 may read the instructions and / or data in the storage module to implement the aforementioned method embodiment.

[0211] Optionally, the communication module 701 may include a transmitting unit and a receiving unit. The transmitting unit is configured to perform the transmitting operation in the above method embodiment. The receiving unit is configured to perform the receiving operation in the above method embodiment. Optionally, the communication module 701 may be implemented by a transceiver or transceiver-related circuits.

[0212] It should be noted that the computing device 70 may include a sending unit but not a receiving unit. Alternatively, the computing device 70 may include a receiving unit but not a sending unit. This may depend on whether the above solution executed by the computing device 70 includes both sending and receiving actions.

[0213] Optionally, the computing device 70 may be used to execute the actions performed by the electronic device in the above method embodiment. The computing device 70 may be an electronic device or a component that can be configured in an electronic device (e.g., a processor, a chip, or a chip system). For example, the computing device 70 is used to execute the following scheme:

[0214] The communication module is used to acquire a first image;

[0215] The processing module is configured to determine a motion region image in the first image based on the first image;

[0216] The processing module is further configured to input the motion region image into a first neural network to obtain a first output result;

[0217] The processing module is further configured to, when the first output result indicates that a target object exists in the motion area image, determine through a second neural network that the motion area image contains a target object, and generate a trigger signal, wherein the trigger signal is used to wake up the device, and the first neural network is cascaded to the second neural network.

[0218] In a possible implementation, the processing module is specifically configured to:

[0219] Inputting feature data of the middle layer of the first neural network into the second neural network to determine whether the motion area image contains the target object, wherein the accuracy of the second neural network is greater than the accuracy of the first neural network.

[0220] In a possible implementation, the processing module is specifically configured to:

[0221] A motion region image in the first image is determined based on the first image and an average image, wherein the first image includes an image acquired at a first moment, and the average image is determined based on images acquired before the first moment.

[0222] In a possible implementation, the processing module is specifically configured to:

[0223] Acquire the average image, and determine an average image corresponding to the first image based on the first image and the average image;

[0224] The motion region image in the first image is determined by performing inter-frame difference according to the first image and an average image corresponding to the first image.

[0225] In a possible implementation, the processing module is further configured to:

[0226] It is determined that the moving direction of the target object in the moving area image is consistent with a preset direction, where the preset direction includes a moving direction set by a user.

[0227] In a possible implementation, the processing module is specifically configured to:

[0228] Acquire key points of the target object in the motion area image;

[0229] determining a moving direction of the target object according to the key points;

[0230] It is determined that the movement direction is consistent with the movement direction.

[0231] In a possible implementation, the processing module is specifically configured to:

[0232] The key point includes at least one of a first key point, a second key point, and a third key point, wherein the first key point is a point located on the head of the target object, the second key point is a point located on the left shoulder of the target object, and the third key point is a point located on the right shoulder of the target object;

[0233] Obtaining a first distance and a second distance, wherein the first distance is a perpendicular distance from the first key point to a first connecting line, the first connecting line is a connecting line between the second key point and the third key point, and the second distance is a distance between the second key point and the third key point;

[0234] The moving direction of the target object is determined according to the ratio of the first distance to the second distance.

[0235] In a possible implementation, the processing module is configured to:

[0236] If a ratio of the first distance to the second distance is less than a first threshold, it is determined that the moving direction of the target object is forward movement or backward movement.

[0237] In a possible implementation, the processing module is specifically configured to:

[0238] Acquire a motion region image of N frames of images acquired before the first moment, where N is a positive integer greater than or equal to 1;

[0239] The moving direction of the target object is determined to be forward or backward according to the second distance and an average of the second distances corresponding to the key points of the target object in the moving area images of the N frames of images.

[0240] In a possible implementation, the processing module is specifically configured to:

[0241] The absolute value of the difference between the second distance and the mean is greater than a second threshold, and the moving direction of the target object is forward; or

[0242] An absolute value of a difference between the second distance and the mean is less than or equal to a second threshold, and the moving direction of the target object is backward.

[0243] In a possible implementation, the processing module is specifically configured to:

[0244] If a ratio of the first distance to the second distance is greater than or equal to a first threshold, it is determined that the moving direction of the target object is leftward or rightward.

[0245] In a possible implementation, the processing module is specifically configured to:

[0246] The key point includes the center point of the motion area image;

[0247] According to the fact that a change of the center point of the motion region image on the X-axis is greater than a change on the Y-axis, it is determined that the motion direction of the target object is leftward motion or rightward motion.

[0248] In a possible implementation, the processing module is specifically configured to:

[0249] The key point includes the center point of the motion area image;

[0250] According to the fact that a change of the center point of the motion region image on the X-axis is less than or equal to a change on the Y-axis, it is determined that the motion direction of the target object is forward motion or backward motion.

[0251] FIG8 is a schematic diagram of the structure of an electronic device 80 provided in an embodiment of the present application. The electronic device 80 is a device with computing capabilities. The device here can be a physical device, such as a controller, processor, server (such as a rack server), host, etc., or a virtual device, such as a virtual machine, container, etc.

[0252] As shown in Figure 8, the electronic device 80 includes: a processor 802 and a memory 801, and optionally includes a bus 804 and a communication interface 803. The processor 802 and the memory 801 communicate with each other via the bus 804. It should be understood that this application does not limit the number of processors and memories in the electronic device 80. Among them, the processor 802 and the memory 801, as well as the optional bus 804 and the communication interface, can be integrated into a system on a chip (SoC). SoC is a chip that integrates multiple functional modules. It integrates multiple functions on a single chip, which can achieve the characteristics of high integration, high performance and low power consumption.

[0253] The memory 801 is used to provide storage space, which can optionally store application data, user data, operating systems, and computer programs. The memory 801 may include volatile memory, such as random access memory (RAM). The memory 801 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0254] The processor 802 is a module for performing calculations and may include any one or more of a controller (e.g., a storage controller), a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), a digital signal processor (DSP), a coprocessor (to assist the central processor in completing corresponding processing and applications), an application-specific integrated circuit (ASIC), a microcontroller unit (MCU), and the like.

[0255] The communication interface 803 is used to provide information input or output for the at least one processor. And / or, the communication interface 803 can be used to receive data sent externally and / or send data to the outside. The communication interface 803 can be a wired link interface such as an Ethernet cable, or a wireless link (Wi-Fi, Bluetooth, general wireless transmission and other wireless communication technologies, etc.) interface. Optionally, the communication interface 803 can also include a transmitter (such as a radio frequency transmitter, antenna, etc.) coupled to the interface, or a receiver, etc.

[0256] Bus 804 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, for example. Buses can be categorized as address buses, data buses, control buses, and the like. For ease of illustration, FIG8 shows only one bus line, but this does not imply a single bus or type of bus. Bus 804 may include a path for transmitting information between various components of electronic device 80 (e.g., memory 801, processor 802, and communication interface 803).

[0257] In the embodiment of the present application, the memory 801 stores executable instructions, and the processor 802 executes the executable instructions to implement the aforementioned method of waking up the device, such as the method of waking up the device in the embodiments of Figures 3, 5, or 6. That is, the memory 801 stores instructions for executing the method of waking up the device.

[0258] An embodiment of the present application also provides a chip device, which includes at least one processor, and the at least one processor is used to call a computer program or instruction stored in a memory so that the processor executes the method of waking up the device in the embodiments such as Figure 3, Figure 5 or Figure 6 above.

[0259] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program or instruction. When the computer program or instruction runs on a processor, it enables the method of waking up the device in the embodiments such as Figures 3, 5 or 6 above.

[0260] An embodiment of the present application also provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are run on a processor, the method for waking up the device in the embodiments of Figures 3, 5 or 6 above is executed.

[0261] It is understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.

[0262] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, a register, a hard disk, a mobile hard disk, a CD-ROM or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.

[0263] In the above embodiments, all or part of the embodiments may be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are performed in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device, or other programmable device. The computer program or instructions may be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program or instructions may be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; an optical medium, such as a digital video disk; or a semiconductor medium, such as a solid-state drive. The computer-readable storage medium may be a volatile or nonvolatile storage medium, or may include both volatile and nonvolatile types of storage media.

[0264] In the various embodiments of the present application, unless otherwise specified or there is a logical conflict, the terms and / or descriptions between different embodiments are consistent and can be referenced by each other. The technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationships.

[0265] In the description of this application, words such as "first", "second", "S301", or "S302" are only used to distinguish the description and facilitate the context. Different sequence numbers themselves do not have specific technical meanings and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying the order of execution of operations. The execution order of each process should be determined by its function and internal logic.

[0266] In this application, the term "and / or" simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. A and B can be singular or plural. Additionally, the character " / " in this document indicates that the related objects are in an "or" relationship.

[0267] In this application, "transmission" may include the following three situations: sending of data, receiving of data, or sending of data and receiving of data. In this application, "data" may include business data and / or signaling data.

[0268] In this application, the terms "comprise" or "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process / method comprising a series of steps, or a system / product / apparatus comprising a series of units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes / methods / products / apparatus.

[0269] In the description of this application, unless otherwise specified, the number of nouns refers to "singular or plural," that is, "one or more." "At least one" means one or more. "Including at least one of the following: A, B, C" means that it may include A, or include B, or include C, or include A and B, or include A and C, or include B and C, or include A, B, and C. A, B, and C can be single or plural.

Claims

1. A method for waking up a device, characterized in that: The method comprises: Acquire a first image, wherein the first image is a frame of image in a video; determining a motion region image in the first image according to the first image; Inputting the motion region image into a first neural network to obtain a first output result; In the case where the first output result is used to indicate that there is a target object in the motion area image, the motion area image is determined to contain the target object through a second neural network, and a trigger signal is generated, wherein the trigger signal is used to wake up the device, and the first neural network is cascaded to the second neural network.

2. The method according to claim 1, characterized in that Determining that the moving area image contains a target object by using a second neural network includes: Inputting feature data of the middle layer of the first neural network into the second neural network to determine whether the motion area image contains the target object, wherein the accuracy of the second neural network is greater than the accuracy of the first neural network.

3. The method according to claim 1 or 2, characterized in that: The determining, according to the first image, a motion region image in the first image comprises: A motion region image in the first image is determined according to the first image and an average image, wherein the first image includes an image acquired at a first moment, and the average image is determined according to images acquired before the first moment.

4. The method according to claim 3, characterized in that The determining the motion region image in the first image according to the first image and the average image comprises: Acquire the average image, and determine the average image corresponding to the first image according to the first image and the average image; According to the first image and an average image corresponding to the first image, a motion region image in the first image is determined by frame difference.

5. The method according to any one of claims 1 to 4, characterized in that: After determining that the target object is included in the moving area image by the second neural network and before generating the trigger signal, the method further includes: It is determined that the moving direction of the target object in the moving area image is consistent with a preset direction, where the preset direction includes a moving direction set by a user.

6. The method according to claim 5, characterized in that The determining that the moving direction of the target object in the moving area image is consistent with a preset direction includes: Acquire key points of the target object in the motion region image; Determine the moving direction of the target object according to the key points; Determine that the movement direction is consistent with the movement direction.

7. The method according to claim 6, characterized in that Determining the moving direction of the target object according to the key point includes: The key point includes at least one of a first key point, a second key point and a third key point, wherein the first key point is a point located at the head of the target object, the second key point is a point located at the left shoulder of the target object, and the third key point is a point located at the right shoulder of the target object; Acquire a first distance and a second distance, wherein the first distance is a vertical distance from the first key point to a first connecting line, the first connecting line is a connecting line between the second key point and the third key point, and the second distance is a distance between the second key point and the third key point; The moving direction of the target object is determined according to the ratio of the first distance to the second distance.

8. The method according to claim 7, characterized in that The determining the moving direction of the target object according to the ratio of the first distance to the second distance includes: If a ratio of the first distance to the second distance is less than a first threshold, it is determined that the moving direction of the target object is forward movement or backward movement.

9. The method according to claim 8, characterized in that The determining whether the moving direction of the target object is forward movement or backward movement includes: Acquire a motion region image of N frames of images acquired before the first moment, where N is a positive integer greater than or equal to 1; According to the second distance and the average of the second distances corresponding to the key points of the target object in the moving area images of the N frames of images, it is determined that the moving direction of the target object is forward movement or backward movement.

10. The method according to claim 9, characterized in that The step of determining, based on the second distance and an average of the second distances corresponding to the moving region images of the N frames of images, that the moving direction of the target object is forward movement or backward movement comprises: The absolute value of the difference between the second distance and the mean is greater than a second threshold, and the moving direction of the target object is forward movement; or, The absolute value of the difference between the second distance and the mean is less than or equal to a second threshold, and the moving direction of the target object is backward.

11. The method according to claim 7, characterized in that The determining the moving direction of the target object according to the ratio of the first distance to the second distance includes: If a ratio of the first distance to the second distance is greater than or equal to a first threshold, it is determined that the moving direction of the target object is moving to the left or to the right.

12. The method according to claim 6, characterized in that Determining the moving direction of the target object according to the key point includes: The key point includes the center point of the motion area image; According to the fact that the change of the center point of the motion region image on the X-axis is greater than the change on the Y-axis, it is determined that the motion direction of the target object is leftward motion or rightward motion.

13. The method according to claim 6, characterized in that Determining the moving direction of the target object according to the key point includes: The key point includes the center point of the motion area image; According to the change of the center point of the motion region image on the X-axis being less than or equal to the change on the Y-axis, it is determined that the motion direction of the target object is forward motion or backward motion.

14. A computing device, characterized in that: The device comprises a communication module and a processing module, wherein: The communication module is used to acquire a first image; The processing module is used to determine a motion region image in the first image according to the first image; The processing module is further used to input the motion area image into a first neural network to obtain a first output result; The processing module is further used to determine that the motion area image contains the target object through a second neural network when the first output result is used to indicate that there is a target object in the motion area image, and generate a trigger signal, wherein the trigger signal is used to wake up the device, and the first neural network is cascaded to the second neural network.

15. The device according to claim 14, characterized in that The processing module is specifically used for: Inputting feature data of the middle layer of the first neural network into the second neural network to determine whether the motion area image contains the target object, wherein the accuracy of the second neural network is greater than the accuracy of the first neural network.

16. The device according to claim 14 or 15, characterized in that The processing module is specifically used for: A motion region image in the first image is determined according to the first image and an average image, wherein the first image includes an image acquired at a first moment, and the average image is determined according to images acquired before the first moment.

17. The device according to claim 16, characterized in that The processing module is specifically used for: Acquire the average image, and determine the average image corresponding to the first image according to the first image and the average image; According to the first image and an average image corresponding to the first image, a motion region image in the first image is determined by frame difference.

18. The device according to any one of claims 14 to 17, characterized in that: The processing module is further used for: It is determined that the moving direction of the target object in the moving area image is consistent with a preset direction, where the preset direction includes a moving direction set by a user.

19. The device according to claim 18, characterized in that The processing module is specifically used for: Acquire key points of the target object in the motion region image; Determine the moving direction of the target object according to the key points; Determine that the movement direction is consistent with the movement direction.

20. The device according to claim 19, characterized in that The processing module is specifically used for: The key point includes at least one of a first key point, a second key point and a third key point, wherein the first key point is a point located at the head of the target object, the second key point is a point located at the left shoulder of the target object, and the third key point is a point located at the right shoulder of the target object; Acquire a first distance and a second distance, wherein the first distance is a vertical distance from the first key point to a first connecting line, the first connecting line is a connecting line between the second key point and the third key point, and the second distance is a distance between the second key point and the third key point; The moving direction of the target object is determined according to the ratio of the first distance to the second distance.

21. The device according to claim 20, characterized in that The processing module is used for: If a ratio of the first distance to the second distance is less than a first threshold, it is determined that the moving direction of the target object is forward movement or backward movement.

22. The device according to claim 21, characterized in that The processing module is specifically used for: Acquire a motion region image of N frames of images acquired before the first moment, where N is a positive integer greater than or equal to 1; According to the second distance and the average of the second distances corresponding to the key points of the target object in the moving area images of the N frames of images, it is determined that the moving direction of the target object is forward movement or backward movement.

23. The device according to claim 22, characterized in that The processing module is specifically used for: The absolute value of the difference between the second distance and the mean is greater than a second threshold, and the moving direction of the target object is forward movement; or, The absolute value of the difference between the second distance and the mean is less than or equal to a second threshold, and the moving direction of the target object is backward.

24. The device according to claim 20, characterized in that The processing module is specifically used for: If a ratio of the first distance to the second distance is greater than or equal to a first threshold, it is determined that the moving direction of the target object is moving to the left or to the right.

25. The device according to claim 19, characterized in that The processing module is specifically used for: The key point includes the center point of the motion area image; According to the fact that the change of the center point of the motion region image on the X-axis is greater than the change on the Y-axis, it is determined that the motion direction of the target object is leftward motion or rightward motion.

26. The device according to claim 19, characterized in that The processing module is specifically used for: The key point includes the center point of the motion area image; According to the change of the center point of the motion region image on the X-axis being less than or equal to the change on the Y-axis, it is determined that the motion direction of the target object is forward motion or backward motion.

27. An electronic device, characterized in that: The electronic device includes at least one processor and at least one memory, wherein the at least one memory stores computer instructions; the at least one processor is used to call the computer instructions to implement the method described in any one of claims 1-13.

28. A computer-readable storage medium, characterized in that: The computer-readable storage medium comprises computer program instructions, and when the computer program instructions are executed by a processor, the method of any one of claims 1 to 13 is implemented.

29. A computer program product, characterized in that The computer program product includes a computer program or instructions, and when the computer program or instructions are executed, the method according to any one of claims 1 to 13 is performed.

30. A chip system, characterized in that: The chip system includes at least one processor, a memory and an interface circuit. The memory, the interface circuit and the at least one processor are interconnected through lines. Computer program instructions are stored in the at least one memory. When the computer program instructions are executed by the at least one processor, the method described in any one of claims 1 to 13 is implemented.

Citation Information

Patent Citations

  • Method, system and chip for waking up equipment

    CN117891516A

  • Robot awakening method and device, terminal equipment and storage medium

    CN109955257A

  • Neural Network Systems for Vehicles

    US20080144944A1

  • Method and apparatus for waking up screen

    WO2020181523A1