Method and apparatus for waking up electronic device, sensor, device, and medium

WO2026113140A1PCT designated stage Publication Date: 2026-06-04SONY SEMICON SOLUTIONS CORP +1

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SONY SEMICON SOLUTIONS CORP
Filing Date
2025-01-26
Publication Date
2026-06-04

Smart Images

  • Figure CN2025075229_04062026_PF_FP_ABST
    Figure CN2025075229_04062026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method and apparatus for waking up an electronic device, a sensor, a device, and a medium. In the method, when an electronic device is in a sleep state, image data is captured by means of a sensor, and the electronic device is awakened on the basis of a wake-up signal outputted by the sensor on the basis of the captured image data. On the basis of the above-described solution, by using the sensor to capture and process image data to automatically wake up the electronic device, the convenience and intelligence of device wake-up can be improved, and the overall power consumption of the device can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, apparatus, sensors, devices, and media for waking up electronic devices. Technical Field

[0001] This disclosure relates to the field of information processing, and more specifically, to methods, apparatus, sensors, devices, and media for waking up electronic devices in the field of information processing. Background Technology

[0002] With the development of technology, electronic devices are gradually becoming integrated into people's lives and have increasingly diverse functions. For example, smart glasses, as wearable devices, are receiving more and more attention. They can include integrated batteries, controllers, and cameras, as well as components such as microphones, speakers, and displays, to help meet user needs such as information retrieval, communication, and entertainment.

[0003] Currently, smart glasses that are in sleep mode to maintain low power consumption are often woken up by clicking a button or inputting voice. However, clicking a button requires the user to raise their hand to operate it. Frequent hand raising can be uncomfortable and even inconvenient for users. While voice input eliminates the need to raise the hand, it is unsuitable in quiet environments. Furthermore, the accuracy of voice recognition is also a concern, especially in noisy environments where accuracy drops significantly, making it difficult to wake up the device correctly. In addition, both button clicking and voice input methods can lead to the device not being properly woken up when needed due to user oversight, resulting in missed operations. Similar problems exist for other electronic devices that are woken up by clicking buttons or inputting voice.

[0004] Therefore, there is a need to provide a new way to wake up electronic devices, which can wake up electronic devices more conveniently and intelligently. Summary of the Invention

[0005] One aspect of this disclosure relates to a method for waking up an electronic device. The method may include: capturing image data via a sensor while the electronic device is in a sleep state; and waking up the electronic device based on a wake-up signal output by the sensor based on the captured image data.

[0006] Another aspect of this disclosure relates to a sensor. The sensor may include: a photoelectric conversion unit configured to capture image data by converting incident light signals into electrical signals; and an on-chip processor configured to output a wake-up signal based on the captured image data to wake up an electronic device.

[0007] Another aspect of this disclosure relates to an electronic device. The electronic device may include: the sensor described above; and a device configured to receive image data output from the sensor.

[0008] Another aspect of this disclosure relates to an apparatus for waking up an electronic device. The apparatus may include: a processor; and a memory, including computer program instructions, wherein the memory and computer program instructions are configured to cause the apparatus to perform operations via the processor. The operations may include: processing image data captured by sensors using a machine learning-based model or image processing operations via sensors included in the electronic device; and waking up the electronic device based on the processing result.

[0009] Another aspect of this disclosure relates to a computer-readable storage medium storing one or more computer program instructions. According to embodiments of this disclosure, the one or more computer program instructions can cause the processing device to perform the methods described above when executed by a processing device.

[0010] The above overview is provided to summarize some exemplary embodiments to provide a basic understanding of the aspects of the subject matter described herein. Therefore, the features described above are merely examples and should not be construed as narrowing the scope or spirit of the subject matter described herein in any way. Other features, aspects, and advantages of the subject matter described herein will become apparent from the following detailed description, taken in conjunction with the accompanying drawings. Attached Figure Description

[0011] A better understanding of this disclosure can be obtained by considering the following detailed description of the embodiments in conjunction with the accompanying drawings. The same or similar reference numerals are used in the drawings to denote the same or similar parts. The drawings, together with the following detailed description, are incorporated in and form a part of this specification to illustrate embodiments of the disclosure and explain the principles and advantages of the disclosure. Wherein:

[0012] Figure 1 is a schematic diagram of smart glasses, an example of an electronic device in the related technology.

[0013] Figure 2A is a schematic diagram of the processing sequence corresponding to the button triggering method of waking up an electronic device using a button click in related technologies.

[0014] Figure 2B is a schematic diagram of the processing sequence corresponding to the voice triggering method for waking up electronic devices using voice input in related technologies.

[0015] Figure 3 is a flowchart of a method for waking up an electronic device according to an embodiment of the present disclosure.

[0016] Figures 4A and 4B are examples of targets that can be detected, for example, by a DNN (deep neural network) model, according to embodiments of the present disclosure.

[0017] Figure 5A is a schematic diagram of a DNN model according to an embodiment of the present disclosure.

[0018] Figure 5B is a flowchart of an example method for detecting gestures using image processing according to an embodiment of the present disclosure.

[0019] Figure 6A is a schematic diagram of the processing sequence corresponding to the DNN triggering method for waking up an electronic device using a DNN according to an embodiment of the present disclosure.

[0020] Figure 6B is a schematic diagram of the processing sequence corresponding to the image processing triggering method for waking up an electronic device using a non-DNN according to an embodiment of the present disclosure.

[0021] Figure 7A is a schematic diagram of sensor mode switching corresponding to the DNN triggering method according to an embodiment of the present disclosure.

[0022] Figure 7B is a schematic diagram of sensor mode switching corresponding to the image processing triggering method according to an embodiment of the present disclosure.

[0023] Figure 8 is a schematic diagram illustrating an example of an application scenario of an electronic device according to an embodiment of the present disclosure.

[0024] Figure 9 is a block diagram illustrating an example of a schematic configuration of an electronic device according to an embodiment of the present disclosure.

[0025] Figure 10 is a block diagram of another example of a schematic configuration of an electronic device according to an embodiment of the present disclosure.

[0026] Figure 11 is a block diagram of yet another example of a schematic configuration of an electronic device according to an embodiment of the present disclosure.

[0027] While the embodiments described in this disclosure may be readily modified and alternatively implemented, specific embodiments thereof are shown by way of example in the accompanying drawings and are described in detail herein. However, it should be understood that the drawings and the detailed description thereof are not intended to limit the embodiments to the specific forms disclosed, but rather are intended to cover all modifications, equivalents, and alternatives that fall within the spirit and scope of the claims. Detailed Implementation

[0028] The following description illustrates representative applications of the devices and methods described herein. These examples are provided merely to provide context and aid in understanding the described embodiments. Therefore, it will be apparent to those skilled in the art that the embodiments described below can be practiced without some or all of the specific details provided. In other instances, well-known process steps have not been described in detail to avoid unnecessarily obscuring the described embodiments. Other applications are also possible, and the scope of this disclosure is not limited to these examples.

[0029] Referring first to Figure 1, a schematic diagram of smart glasses, an example of an electronic device in the related art, is shown in Figure 1.

[0030] Smart glasses can integrate electronic components such as processors, memory, cameras, microphones, and displays. When a user wears smart glasses, they remain in standby mode to conserve power. After the user wakes the smart glasses, the processor starts up from sleep mode and can then provide various services, including virtual reality (VR) or augmented reality (AR), based on information from the microphone, camera, etc., thereby expanding the user's capabilities and sensory experience. As shown on the left side of Figure 1, a camera can be mounted at the end of the temple of the smart glasses, as indicated by the circled area. The right side of Figure 1 shows the camera as seen from the front of the smart glasses.

[0031] Currently, users can wake up electronic devices such as smart glasses using button triggering and voice triggering. In button triggering, the user raises their hand to click the wake-up button on the electronic device, thus activating it from standby mode to normal operation. However, frequently raising the hand to wake up the device is cumbersome and unpleasant for the user. Figure 2A illustrates a schematic diagram of the processing sequence 200-A corresponding to the button triggering method for waking up an electronic device using button clicks in related technologies.

[0032] In S210-A, the button used to wake up the electronic device (e.g., a power button, a designated area, etc.) receives a button click from the user and transmits the trigger signal generated by the button click to the System-on-Chip (SoC), which acts as the processor of the electronic device. In S220-A, the SoC responds to the trigger signal and transmits it via I... 2 The C (integrated circuit bus) sends command signals to the image sensor corresponding to the camera to enable the image sensor to function properly. In S230-A, the image sensor responds to the command signal to capture an image or video. As shown in Figure 2A, the normal operation of the image sensor is triggered by a button click, and once triggered, the image sensor will enter normal operating mode.

[0033] In voice-triggered mode, users activate electronic devices from standby to normal operation by speaking (e.g., uttering specific commands). However, voice input is cumbersome and inconvenient for users when they need to frequently speak to wake up electronic devices, or when they need to wake up electronic devices while remaining quiet. Figure 2B shows a schematic diagram of the processing sequence 200-B corresponding to the voice-triggered mode of waking up electronic devices using voice input.

[0034] In S210-B, the microphone in the electronic device in standby mode remains on, and when voice input is received, the input voice is analyzed (e.g., the voice is converted to text and analyzed). In response to determining that the voice is a specific instruction for waking up the electronic device, a trigger signal is transmitted to the SoC in the electronic device. In S220-B, the SoC responds to the trigger signal via I... 2 C sends a command signal to the image sensor corresponding to the camera to enable the image sensor to function properly. In S230-B, the image sensor responds to the command signal to capture an image or video. As shown in Figure 2B, the normal operation of the image sensor is triggered by voice input, and once triggered, the image sensor will enter normal operating mode.

[0035] Besides being cumbersome and inconvenient, button-triggered and voice-triggered methods may also lead to users neglecting to properly activate the electronic device when needed, resulting in missed operations. To alleviate or resolve at least some of these problems, embodiments of this disclosure provide a method for waking up an electronic device using sensor processing. This method not only avoids the problems caused by button-triggered and voice-triggered methods but also automatically wakes up the electronic device based on the sensor's processing of the image data it captures. This improves the convenience and intelligence of device wake-up and reduces the overall power consumption of the device.

[0036] [Corrected according to Rule 91 04.03.2025] Figure 3 is a flowchart of a method 300 for waking up an electronic device according to an embodiment of the present disclosure. Although the following description uses smart glasses as an example of an electronic device, the present disclosure is not limited thereto. Those skilled in the art will readily realize that the electronic device can be other information processing devices with image sensors, such as AR devices with cameras (e.g., head-mounted or wrist-worn AR devices), smart cameras, or devices with smart cameras (e.g., smart appliances, smart toys, etc.).

[0037] In S310, image data is captured via a sensor when the electronic device is in sleep mode.

[0038] Since an image sensor (such as the camera shown in Figure 1) can be installed on an electronic device, the device can capture image data via the image sensor. The image data captured by the sensor may include image data containing a specific target. This specific target may be pre-defined to wake the device when it is in sleep mode. Such a target can be pre-set in the electronic device's processing program or can be set or changed by the user as needed. When the captured image data contains this specific target, it can indicate that the electronic device needs to be woken up to enter normal operating mode.

[0039] In one or more examples, the specific target could be a user gesture. For instance, a user could make a specific gesture when they need to wake up an electronic device, allowing the sensor to capture it and then output a wake-up signal to wake up the device by recognizing or analyzing the gesture. Since users can make gestures without lifting their hands to touch a switch or speaking a wake-up command, and only need to keep the gesture within the sensor's field of view to wake up the device, this greatly simplifies user operation.

[0040] In one or more examples, the specific target can be a predetermined object that requires the electronic device to take action. For example, certain items can be pre-defined as predetermined objects, such as keys, mobile phones, computers, blackboards, beds, study desks, or gaming devices. When the sensor determines that the captured image data contains the predetermined object, the sensor can output a wake-up signal to wake up the electronic device to take the relevant action. Thus, even if the user is unaware that the electronic device needs to be woken up to record something, the electronic device can be automatically woken up to perform the relevant operation.

[0041] As large models develop, their memory capabilities are becoming increasingly powerful. When sensors detect a predetermined object, electronic devices can be activated to record information related to a specific scene or environment, for example, by capturing images. This allows for querying and / or processing of information recorded by the electronic device and the large model when the user forgets something, thereby expanding the functionality and capabilities of the electronic device.

[0042] For example, when a sensor captures image data containing a blackboard, it can process this data to wake up an electronic device to instruct the sensor to record video and / or instruct the microphone to record audio, thus recording content in possible teaching or meeting scenarios for review or retrieval. As another example, when a sensor captures image data containing both a key and a door, it can process this data to wake up an electronic device to instruct the sensor to record video and / or take a picture, allowing users to check whether the door was locked when unsure. Those skilled in the art can conceive of many similar processes, such as automatically taking pictures upon detecting famous buildings, automatically recording video upon detecting collisions caused by traffic accidents, and automatically issuing an alarm when an obstacle is detected and the user is getting closer. Since the electronic device can be automatically woken up to record relevant information based on predetermined objects, users can obtain relevant information by asking the electronic device when they forget or cannot remember something. This greatly expands the user's memory capacity, thereby improving their learning, work, and life abilities.

[0043] Because the sensor continues to capture image data even when the electronic device is in sleep mode, it can be required to keep the sensor always on. However, while the sensor remains on, it can consume less power (or receive less power) when the electronic device is in sleep mode, thus allowing the sensor to operate in a low-power always-on state.

[0044] The following is a detailed description of the sensor. Sensors can be directly integrated into electronic devices. Compared to augmented reality (AR) environments where special devices need to be worn on the user's arm or fingers to detect gestures or movements and transmit the results to head-mounted devices such as smart glasses, directly capturing information such as gestures or movements through sensors within electronic devices avoids the need for additional equipment and prevents the distraction caused by such equipment, thus reducing the user's burden.

[0045] An electronic device may contain only one sensor related to image capture (also called an image sensor). This sensor can have two modes: a first mode (also called wake-up mode) and a second mode (also called viewing mode or streaming mode). The first mode can be further specified as an AI mode that uses a machine learning-based model to generate the wake-up signal, and an image processing mode that uses a non-machine learning-based model to generate the wake-up signal. When the machine learning-based model is a DNN, the AI ​​mode can also be called the DNN mode. The image processing mode, because it does not use a machine learning-based model, can process image data using algorithms that do not involve neurons, such as conventional image processing algorithms. Regardless of the wake-up signal generation method, when the sensor is in the first mode, it determines whether to wake up the electronic device based on the acquired image data. Once the electronic device is woken up, the sensor can automatically enter the second mode.

[0046] Given that the sensor can have a first mode and a second mode, the camera containing the sensor can have two corresponding modes: a always-on wake-up mode and an image data transfer stream mode. When the electronic device is in sleep mode, the sensor is in the first mode to determine whether to output a wake-up signal to wake up the electronic device based on the captured image data. When the electronic device is awakened and in normal working condition, the sensor switches from the first mode to the second mode for regular image capture, etc. Since the same sensor can have two modes and switch between them, hardware resources and costs can be saved, hardware utilization efficiency can be improved, and the need for additional equipment can be avoided. Of course, those skilled in the art will understand that the electronic device can also have two, three, or more image sensors, where at least one image sensor has both of the above modes.

[0047] According to embodiments of this disclosure, the sensor may include a photoelectric conversion unit and an on-chip processor. The mode in which the sensor operates also corresponds to the mode in which the photoelectric conversion unit and the on-chip processor operate. The photoelectric conversion unit may be configured to capture image data by converting incident light signals into electrical signals, and the on-chip processor may be configured to output a wake-up signal based on the captured image data. The photoelectric conversion unit may be a conventional photoelectric conversion unit, such as a CIS (CMOS image sensor), which corresponds to a conventional sensor for capturing and outputting image data, but does not contain an on-chip processor like the sensor according to embodiments of this disclosure. The photoelectric conversion unit can perform photoelectric conversion on the light signals incident upon it to generate pixel signals, thereby obtaining image data composed of pixel signals. An on-chip processor means that the processor is located on the chip where the sensor is located, that is, the on-chip processor and the photoelectric conversion unit are located on the same chip. In this way, when the photoelectric conversion unit outputs a signal to the on-chip processor, the signal transmission can be performed only within the sensor chip, thereby reducing interference, delay and other problems in the signal transmission process to improve signal processing efficiency, and since only the sensor chip operates to determine whether to wake up the device when the device is in sleep mode, the overall power consumption of the device can be reduced. Because such sensor chips can be installed in different electronic devices, it is easy to upgrade existing electronic devices and reduce upgrade costs.

[0048] In the first mode, the photoelectric conversion unit outputs the captured image data to the on-chip processor, so that the on-chip processor determines whether to output a wake-up signal based on the captured image data. For example, the on-chip processor can use an image processing algorithm to process the image data to determine whether to output a wake-up signal. Any existing or future image processing algorithm can be used, as long as it can generate a wake-up signal when a specific target (e.g., a gesture) for waking up the electronic device is detected. FIG. 5B shows a flowchart of an example of a method 500 for detecting gestures using image processing according to an embodiment of the present disclosure. Although this method involves gesture detection algorithms in the related art, the related art does not use such detected gestures to wake up the electronic device, nor does it place the processor for detecting gestures on the chip where the sensor is located to generate a wake-up signal only through signal interaction within the sensor chip, thus failing to achieve the advantages of improved signal processing efficiency and reduced power consumption as in the embodiments of the present disclosure. Those skilled in the art will understand that, in addition to the image processing method of FIG. 5B, other image processing methods can be used to identify other gestures or predetermined objects, as long as a wake-up signal can be output when a specific target is detected. The method of FIG. 5B performed by the on-chip processor is briefly described below.

[0049] In S510, image acquisition is performed via the photoelectric conversion unit in the sensor. Since the images acquired by the photoelectric conversion unit are all the same size, no resizing operation is required. In S520, the acquired image is converted into a grayscale image. This conversion reduces the influence of lighting and color on the image. In S530, an edge detection algorithm is used to detect the gesture contour based on the grayscale image. For example, the Canny operator can be used to extract the gesture contour. In S540, edge enhancement is performed. For example, noise in the image can be removed and the contour enhanced by erosion followed by dilation, and the extracted gesture contour is scaled to the same size for comparison. In S550, the features of the gesture are extracted. For example, the area of ​​the gesture contour, the perimeter of the contour, and the features of the convex hull (the area enclosed by the gesture contour) can be calculated. In S560, feature comparison and classification are performed to determine what gesture is contained in the image. For example, suppose the method in Figure 5B is used to identify rock, scissors, and paper. If the detected gesture has a large convex hull area and no obvious indentation points, the gesture is judged to be "rock". If the detected gesture has two prominent fingertips in its convex hull (e.g., two indentations), the gesture is identified as "scissors." If the detected gesture has a relatively flat outline and complex edges, and the convex hull features multiple fingertips, the gesture is identified as "cloth." In S570, when a gesture matching the gesture used to wake up the device is detected, a wake-up signal instructing the electronic device to be woken up is output to a device outside the sensor (e.g., SoC, application processor).

[0050] In addition to using image processing algorithms to generate wake-up signals, the on-chip processor can also use machine learning-based models to generate wake-up signals. In some embodiments, the on-chip processor may use image processing algorithms to generate wake-up signals, in others it may use machine learning-based models, and in still others it may use both methods simultaneously to further improve the accuracy of wake-up signal generation.

[0051] When the on-chip processor uses a machine learning-based model to generate a wake-up signal, the machine learning-based model can be obtained through pre-training. This model can be trained to output a trigger signal related to waking up an electronic device based on captured image data input to its input. For example, when the trigger signal is high, it acts as a wake-up signal indicating that the electronic device (e.g., the application processor within the electronic device) needs to be woken up; otherwise, it does not. The model can be trained using any suitable supervised, unsupervised, semi-supervised, or reinforcement learning techniques. For example, various image forms that require waking up the electronic device (e.g., specific gestures, predetermined objects, etc.) can be pre-labeled for training the model. Alternatively, the model can be rewarded or penalized for the image analysis results to automatically learn which situations require waking up the electronic device. In embodiments of this disclosure, the machine learning-based model can be a DNN, and DNNs and their operations will be described in detail below.

[0052] In the second mode, the photoelectric conversion unit outputs the captured image data to a device external to the sensor. Specifically, in the first mode, the photoelectric conversion unit outputs the captured image data to the on-chip processor to determine whether to wake up the electronic device. At this time, the sensor does not need to transmit image data to its external device. After the electronic device is woken up to switch to the second mode, the photoelectric conversion unit can stop outputting the captured image data to the on-chip processor and instead output it to a device external to the sensor. At this time, the sensor communicates with the external device, outputting image data at normal resolution so that the external device can perform corresponding operations based on the image data. In the second mode, since the device has already been woken up and no wake-up operation is required, the on-chip processor can be stopped to further save power. Those skilled in the art will understand that the on-chip processor can also be used in the second mode. According to embodiments of this disclosure, the shutdown of the on-chip processor can be selective. For example, after entering the second mode, when the photoelectric conversion unit outputs image data to the external device for a predetermined time, power supply to the on-chip processor can be stopped or the output of the on-chip processor can be ignored. For example, after entering the second mode, when the woken electronic device does not need the on-chip processor to work, it can stop supplying power to the on-chip processor or not respond to the output of the on-chip processor. When the woken electronic device needs the on-chip processor to work, it can supply power to it or request it to perform the corresponding operation.

[0053] The power consumed (or received) by the sensor in the first mode can be less than that consumed in the second mode, which helps reduce power consumption. Although the lower power consumption of the sensor in the first mode means that the photoelectric conversion unit may also consume less power and may not be able to acquire image data with the same resolution or quality as conventional image capture, such image data is sufficient to provide the on-chip processor with relevant information to determine whether to wake up the electronic device.

[0054] Because the on-chip processor resides on the same chip as the sensor—that is, the on-chip processor and the photoelectric conversion unit (e.g., CIS) are on the same chip—the image data acquired by the sensor in the first mode does not need to be output to external memory or a processor (e.g., SoC). Instead, it is output to the on-chip processor for image processing or analysis used for device wake-up. This simplifies the hardware involved in device wake-up, ensuring that only the sensor chip remains operational during the wake-up process. This facilitates device management and reduces power consumption of other hardware components.

[0055] After the on-chip processor outputs a wake-up signal to wake up the electronic device based on the image data captured by the photoelectric conversion unit in the first mode, the electronic device is awakened, and the sensor is automatically switched to the second mode for regular image capture. In the second mode, the sensor can output images or videos with a higher resolution than the image data in the first mode to external memory or the processor. Since the on-chip processor implementing the image processing algorithm or machine learning-based model is included in the sensor, it is possible to avoid adding additional hardware and increasing the device's size.

[0056] When the DNN is implemented in an on-chip processor, the sensor according to embodiments of this disclosure may include, in addition to a photoelectric conversion unit like a conventional sensor, a DNN module related to waking up the electronic device. This DNN module can be implemented in the on-chip processor as a software program or a hardware module implementing the software program. The photoelectric conversion unit can be configured to convert captured optical signals into electrical signals, thereby obtaining captured image data. The DNN module can be configured to generate a processing result based on the electrical signal output from the photoelectric conversion unit when the electronic device is in a sleep state. This processing result may include a trigger signal related to waking up the electronic device. When the trigger signal meets predetermined conditions, such as a high level or a value greater than a specific threshold, the trigger signal is used as the wake-up signal.

[0057] As an example of a machine learning-based model, a DNN can be implemented using various structures, such as one or more convolutional layers, one or more fully connected layers, etc., as long as it can output a trigger signal related to waking up an electronic device based on the input image data. According to embodiments of this disclosure, a DNN for object detection can be used to determine whether to set the output to a form for waking up an electronic device based on whether a predetermined target is detected. For example, a classifier containing fully connected layers can be used to implement the DNN to detect whether a target is present in the input image. As another example, a DNN can be implemented based on a pruned MobileNet SSD (Single Shot MultiBox Detector). That is, a DNN is constructed by pruning an existing MobileNet SSD model to detect a target at a certain location in the input image. A pruned DNN can have a simpler network structure, less computation and model parameters, and faster processing speed, which is beneficial for use in resource-constrained devices such as smart glasses for object detection. Those skilled in the art will understand that, in addition to implementing DNN in the manner described above, other models can also be used to implement DNN, such as the YOLO (You only look once) series of models.

[0058] Figures 4A and 4B illustrate examples of targets that the DNN model according to embodiments of the present disclosure can detect. When a target is detected, a trigger signal for waking up the electronic device can be output as a wake-up signal. In Figure 4A, as shown in the box, the on-chip DNN sensor located on the sensor can detect gestures in various environments. As shown in Figure 4B, the on-chip DNN sensor can perform human detection to identify human bodies in the sensor's field of view. Of course, the on-chip DNN sensor can also detect other objects (not shown), such as beds or blackboards.

[0059] Figure 5A illustrates a schematic diagram of a DNN model according to an embodiment of the present disclosure, implemented based on Pruned MobileNet SSD. In addition to showing the DNN structure, Figure 5A also shows some parameters. Those skilled in the art will understand that the DNN structure and parameters in Figure 5A are merely examples to aid in a better understanding of the present disclosure and do not constitute a limitation thereof. Those skilled in the art, with the aid of the teachings of the present disclosure and in conjunction with existing machine learning techniques, can implement other DNNs that output trigger signals related to waking up electronic devices based on image data, using other model architectures and / or parameters. Furthermore, while Figure 5A describes a DNN using the detection of user gestures as an example, those skilled in the art will understand that the DNN can also detect other targets such as actions, animals, humans, blackboards, keys, etc.

[0060] As shown in Figure 5A, the image data captured by the sensor is input as a 120*160*1 image to the DNN, where 120*160 represents the size of the input image and 1 represents the number of channels. After passing through the downsampling module and convolutional layer, an intermediate result of 15*20*48 is obtained, where 15*20 represents the feature size and 48 represents the number of channels. Subsequent parameters have similar meanings and will not be elaborated further. Next, the 15*20*48 intermediate result is input to the next-level downsampling module and convolutional layer, resulting in an 8*10*80 intermediate result. Then, the 8*10*80 intermediate result is input to the next-level downsampling module, which downsamples it to obtain a 4*5*80 intermediate result. This result is then input to the next-level downsampling module, which downsamples it to obtain a 2*3*80 intermediate result, and then inputs it to the next-level downsampling module, which downsamples it to obtain a 1*1*64 intermediate result. Each of the aforementioned intermediate results is then transformed into an intermediate result of the same size through convolution and reshape operations, and concatenated to form a feature matrix. Next, an anchoring operation is performed based on the formed feature matrix to generate a large number of detection boxes on the input image, thereby helping to locate the target in the image. The anchoring operation can have 1242 parameters to add 1242 anchor boxes to the input image. Then, non-maximum suppression (NMS) is used to filter the anchor boxes and perform an offset operation to determine the anchor box with the highest confidence that contains the target, thus obtaining the final detection result for the target, such as a gesture in the input image. The anchoring and NMS operations are existing operations that, through their processing, can accurately detect targets based on features extracted from the input image.

[0061] When the DNN detects a user gesture or other predetermined target, it sets the trigger signal to a level used to wake up the electronic device (e.g., a high level) as a wake-up signal. Otherwise, if the DNN does not detect a gesture or other target that needs to wake up the electronic device, it does not change the trigger signal (e.g., it remains at a low level), and no wake-up signal is output from the sensor.

[0062] In S320, the electronic device is woken up based on a wake-up signal output by the sensor based on the captured image data.

[0063] Because sensors have information processing capabilities (e.g., including on-chip processors), wake-up signals can be output based on captured image data using image processing algorithms and / or machine learning-based models, as described above. Wake-up signals can include signals capable of waking up electronic devices (e.g., enabling the CPU (Central Processing Unit), MCU (Microcontroller Unit), DSP (Digital Signal Processor), host processor, etc., of the electronic device to begin operation). For example, wake-up signals can be interrupt signals, start signals, power-on signals, specific digitally encoded signals, etc. The electronic device can be woken up in response to this signal. For example, the chip operates when the start pin of a processor chip located outside the sensor receives a start level. For example, the SoC or application processor (AP) of the electronic device receives a wake-up signal and is woken up when the trigger signal output by the DNN meets predetermined conditions (e.g., goes high), thereby bringing the electronic device into normal operation.

[0064] According to embodiments of this disclosure, image processing algorithms that do not involve neurons can output wake-up signals based on image data, and can also output feature data related to the image data. This feature data can enable the electronic device to provide corresponding operations to the user after being woken up. For example, the image processing algorithm can analyze the image data to determine the scene associated with the image (e.g., classroom, conference room, traffic accident, etc.) and take corresponding operations (e.g., recording audio / video, alarm, etc.).

[0065] Similar to image processing algorithms, machine learning-based models can output feature data related to image data in addition to wake-up signals. The wake-up signal and feature data can be jointly included in the output of the machine learning model, providing different information to a processor outside the electronic device's sensors. The feature data output can also be obtained through pre-training, for example, by pre-specifying different outputs of the model or combinations of multiple outputs to indicate image-related scenes, language translation needs, audio / video recording needs, etc.

[0066] For example, after detecting a target such as a user's gesture or a predetermined object, a DNN can issue a trigger signal as a wake-up signal and output feature data related to the image data. During training, the DNN is not only trained to output trigger signals related to waking up the electronic device based on the input image, but also trained to output relevant feature data based on the input image. For example, when a blackboard with text is detected in an image, feature data related to the blackboard with text is output, thereby instructing the electronic device to record video and / or audio after being woken up; as another example, when non-native language text is detected in an image, feature data related to the non-native language text is output, thereby instructing the electronic device to perform language conversion to the native language after being woken up; as yet another example, when a traffic jam scene is detected in an image, feature data related to the traffic jam scene is output, thereby instructing the electronic device to replan its route; and so on.

[0067] Feature data can indicate the operations that an electronic device should perform under normal operating conditions. Feature data obtained using image processing algorithms or machine learning-based models can help electronic devices perform more appropriate operations, thereby improving the device's intelligence. For example, an electronic device can use sensors to take pictures based on feature data output by a DNN, rather than necessarily using sensors to take pictures upon wake-up. Another example is that an electronic device can provide users with information for decision-making, such as providing path guidance for users including the blind, providing translation support to convert non-native languages ​​into native languages, and providing reading support to extract viewpoints and proofs from obscure materials. Yet another example is that an electronic device can provide information to the user terminal to launch relevant applications, such as enabling smart glasses to communicate with a mobile phone to launch relevant applications on the phone. Still another example is that an electronic device can communicate with a remote server (such as a cloud server, roadside device, etc.) to enable AI services, such as enabling an AI assistant. The above operations are merely examples, and they can appear individually or in combination in different embodiments.

[0068] According to the above technical solution, by utilizing the sensor itself, which remains active while the electronic device is in sleep mode, to capture and process image data, the electronic device can be woken up based on a wake-up signal output by the sensor based on the captured image data. This eliminates the need for user operations such as clicking or speaking, enhancing the user experience and providing a more convenient way to wake up electronic devices. Furthermore, since the electronic device is automatically woken up by the sensor's own processing of the captured image data, it can not only automatically wake up the device when it is needed but the user is unaware of the need for wake-up, thus improving the intelligence of device wake-up, but also confines the wake-up operation to within the sensor, avoiding interference and delays in inter-chip communication caused by excessive hardware involvement in the wake-up operation, and simplifying the operation of other hardware components while saving power consumption.

[0069] The flow of the method for waking up an electronic device has been described above. Below, with reference to FIG6A, a schematic diagram of the processing sequence 600-A corresponding to the DNN triggering method for waking up an electronic device according to an embodiment of this disclosure is described. Although the method for waking up an electronic device using a DNN is described herein, those skilled in the art will realize that other machine learning-based models can also be used to wake up the electronic device. Furthermore, although FIG6A shows the photoelectric conversion unit as a CIS, those skilled in the art will understand that other forms of photoelectric conversion units for capturing and outputting image data are also applicable.

[0070] In the S610-A, when the electronic device is in standby or sleep mode, the CIS remains on in DNN mode (first mode), a specific example of AI mode, to capture image data. In this mode, the CIS consumes less power than it would for regular image capture, capturing lower-resolution image data while maintaining low power consumption. This makes the sensor consume less power in first mode than in second mode. The CIS sends the captured image data to the DNN included in the on-chip processor to determine whether to wake up the electronic device. The DNN can be a software module or a hardware module embedded in the sensor (e.g., a sensor chip). Since such a sensor containing a DNN can function as part of a camera, it eliminates the need for additional sensors or other devices to generate wake-up signals (e.g., interrupt signals to an external SoC), thus saving hardware overhead and cost.

[0071] In the S620-A, the DNN in the sensor's on-chip processor processes the received input image to detect targets it contains. For example, it can detect predetermined objects or gestures by using bounding boxes. Alternatively, it can detect the presence of predetermined objects or gestures using a classifier. Upon detecting a target for waking up the electronic device, the DNN sends a trigger signal to an external processor (e.g., a SoC) as a wake-up signal. This trigger signal can be included in the multi-byte DNN output. In addition to the trigger signal, the DNN output may also contain feature data related to the input image, which the woken-up and operational SoC can use to perform operations related to the feature data.

[0072] It should be noted that although the SoC is used as an example of a component that receives a trigger signal to wake up in the description herein, those skilled in the art will recognize other examples of components that can receive a trigger signal to wake up, such as CPUs, MCUs, DSPs, host processors, etc. The wake-up of such processors can correspond to the wake-up of electronic devices.

[0073] In the S630-A, in response to receiving a trigger signal to wake it up, the SoC switches the CIS, which is kept on in the S610-A, from DNN mode to Viewing mode to obtain images or videos through the same CIS for regular image capture. At this time, the SoC can use I... 2 C sends a switching signal to the CIS to enter the observation mode, for example, to provide the CIS with power for regular image capture and / or to activate the interface for the sensor to transmit signals to external memory or processor.

[0074] In the S640-A, the CIS enters Viewing mode (second mode) to capture images and output images or videos. For example, the CIS can capture images based on instructions issued by the SoC based on feature data contained in the DNN output, or the CIS can capture images directly after being switched to Viewing mode.

[0075] Figure 6B illustrates a schematic diagram of a processing sequence 600-B corresponding to an image processing triggering method for waking up an electronic device using a non-DNN according to an embodiment of the present disclosure. Although the description herein refers to waking up an electronic device using a non-DNN, those skilled in the art will understand that non-DNN can refer to any image processing algorithm that detects a target for waking up the electronic device without using a machine learning-based model.

[0076] In the S610-B, when the electronic device is in standby or sleep mode, the CIS remains powered on in image processing mode (first mode) to capture image data. In this mode, the CIS consumes less power than it would for regular image capture, capturing lower-resolution image data while maintaining low power consumption. The CIS sends the captured image data to the on-chip processor for use by an image processing algorithm (not involving an AI model) to determine whether to wake up the electronic device. Because this sensor, which includes the image processing algorithm, can function as part of a camera, additional sensors or other devices are no longer required to generate wake-up signals (e.g., interrupt signals to an external SoC), thus saving hardware overhead and cost.

[0077] In the S620-B, the image processing algorithm in the sensor's on-chip processor processes the received input image to detect targets contained within it. For example, a gesture can be detected using the method shown in Figure 5B. Alternatively, a predetermined object can be detected in the image using feature matching or similar methods. After detecting a target for waking up the electronic device, the on-chip processor sends a wake-up signal to a processor outside the sensor (e.g., a SoC). Similarly, CPUs (Central Processing Units), MCUs (Microcontrollers), DSPs (Digital Signal Processors), host processors, etc., can also serve as examples of components receiving wake-up signals to restore the electronic device from a sleep state to an active state.

[0078] In the S630-B, in response to receiving a wake-up signal, the SoC switches the CIS (CMOS Image Sensor) in the S610-B, which is currently enabled, from image processing mode to Viewing mode, so that it can capture images or videos using the same CIS for regular image capture. At this time, the SoC can... 2 C sends a switching signal to the sensor to enter observation mode, for example, to provide the sensor with power for regular image capture and / or to activate the sensor's interface for transmitting signals to external memory or processor.

[0079] In the S640-B, the CIS enters Viewing mode (second mode) to capture images and output images or videos. For example, the CIS can capture images based on instructions from the SoC using feature data from the CIS output, or the CIS can capture images directly after being switched to Viewing mode.

[0080] Figure 7A illustrates a schematic diagram of sensor mode switching corresponding to the DNN triggering method according to an embodiment of the present disclosure. Here, DNN is used as an example of a machine learning-based model for description; other AI models that generate wake-up signals based on images captured by the sensor are also possible.

[0081] The sensor shown in Figure 7A includes a photoelectric conversion unit composed of pixels and a DNN module, serving as an example of an on-chip processor. The photoelectric conversion unit converts incident light signals into electrical signals. The DNN module receives the converted electrical signals as input images from the photoelectric conversion unit and processes the input images to determine whether to wake up the SoC. The DNN module here is a lightweight DNN, such as one implemented using a pruned MobileNet SSD, which has less computation and parameters, consumes less power, and can be easily implemented in existing image sensors, thus facilitating low-cost implementation.

[0082] As shown in Figure 7A, regardless of whether the electronic device is in standby or normal operation, the sensor in the electronic device remains on and can be in two modes: DNN mode in standby mode and Viewing mode in normal operation mode. In DNN mode, the sensor consumes less power because it does not need to communicate with external memory or processor and can only capture lower-resolution images for DNN processing. In Viewing mode, the sensor consumes normal power for regular image capture. It should be noted that image capture in this article includes at least one of still image capture and video capture.

[0083] Before the sensor detects a target, the electronic device is in a low-power standby state, and the SoC is in a sleep state, but the sensor remains on in a low-power state to search for the target. In this case, the photoelectric conversion unit captures image data and inputs the image data as the DNN. When the DNN determines that there is no need to wake up the controller, it can stop producing output and continue processing the input image from the photoelectric conversion unit to detect the presence of a predetermined target (e.g., a gesture).

[0084] When the DNN detects a target based on the input image from the photoelectric conversion unit, the DNN outputs a trigger signal as a wake-up signal to wake up the SoC (e.g., outputting an interrupt signal to the SoC), and can also output DNN results (e.g., feature data) to cause the SoC to take appropriate action. The trigger signal can be passed to the SoC via a GPO (Group Policy Object), and the DNN results can be passed via I... 2 C or I 3 The C (Inter-Integrated Circuit Communication Interface) is passed to the SoC. Of course, the trigger signal and the result can also be passed to the SoC as a whole.

[0085] Upon receiving a trigger signal, the SoC is woken up and enters normal operating mode under normal power consumption. At this time, the SoC can switch the sensor from DNN mode to Viewing mode, allowing the sensor to output an image stream to the SoC. In Viewing mode, the DNN can stop working, and the photoelectric conversion unit can send the data obtained from image capture to the SoC for processing. The SoC can receive image output from the sensor in Viewing mode and analyze the image output using a large model to perform corresponding operations. For example, the SoC can recognize gestures, faces, predetermined objects, or scenes from images captured by the photoelectric conversion unit to instruct relevant components to take pictures, tell the user what is in front of them, access relevant applications, and analyze scenes using a large model.

[0086] Figure 7B shows a schematic diagram of sensor mode switching corresponding to the image processing triggering method according to an embodiment of the present disclosure.

[0087] The sensor shown in Figure 7B includes a photoelectric conversion unit composed of pixels and an image processing module, as an example of an on-chip processor. This module employs, for example, conventional image processing algorithms to identify objects or targets included in the image. Similar to Figure 7A, the sensor in the electronic device remains on regardless of whether the electronic device is in standby or normal operation, and can be in two modes: an image processing mode in standby mode and a viewing mode in normal operation mode. In image processing mode, the sensor consumes less power because it does not need to communicate with external memory or processor and can only capture lower-resolution images for processing by the image processing module. In viewing mode, the sensor consumes normal power for regular image capture.

[0088] Before the sensor detects a target, the electronic device is in a low-power standby state, and the SoC is in a sleep state, but the sensor remains on at low power to search for the target. In this case, the photoelectric conversion unit captures image data and inputs the image data as an input image to the image processing module. When the image processing module determines that it does not need to wake up the SoC, it can stop generating output and continue processing the input image from the photoelectric conversion unit to detect the presence of a predetermined target (e.g., a gesture). A low-power island may exist in the SoC, which serves as a component for waking up the SoC by receiving a wake-up signal from the sensor.

[0089] When the image processing module detects a target based on the input image from the photoelectric conversion unit, it outputs a trigger signal (e.g., an interrupt signal to the SoC) as a wake-up signal, and may also output processing results (e.g., feature data) to cause the SoC to take corresponding actions. The trigger signal can be transmitted to the SoC via the GPO, and the processing results can be transmitted via I... 2 C or I 3 The signal is transmitted to the SoC via C or SPI (Serial Peripheral Interface). Of course, the trigger signal and the processing result can also be transmitted to the SoC as a whole.

[0090] In response to a trigger signal received by the low-power island, the SoC is woken up and enters normal operating mode under normal power consumption. At this time, the SoC can switch the sensor from image processing mode to Viewing mode, allowing the sensor to output an image stream to the SoC. In Viewing mode, the image processing module can stop working, and the photoelectric conversion unit can send the data obtained from image capture to the SoC (e.g., for rich applications implemented on the SoC) for processing (e.g., gesture recognition, facial recognition, etc.). The SoC can receive the image output from the sensor in Viewing mode and analyze the image output using a large model to perform corresponding operations.

[0091] Figure 8 illustrates an example scenario of an electronic device application according to an embodiment of the present disclosure. The example in Figure 8 uses AI glasses as an example. These AI glasses are smart glasses that can be woken up by image processing algorithms and / or machine learning-based models. When a sensor detects a predetermined user gesture, the AI ​​glasses, which are in standby mode, are automatically woken up. After being woken up, the AI ​​glasses can communicate with a mobile terminal to launch relevant applications, such as enabling the mobile terminal to perform translation, reading support, route navigation, payment, etc., based on the captured information transmitted by the AI ​​glasses. The AI ​​glasses can also communicate with remote servers (this communication may or may not be via a mobile terminal), such as cloud servers or edge servers (e.g., edge boxes). These remote servers may contain more powerful large-scale models capable of providing AI assistants to the users of the AI ​​glasses, thereby offering various query, analysis, and other services. Users can use gestures to wake up dormant electronic devices via image processing algorithms and / or machine learning-based models to activate desired functions. For example, a user can use gestures to take photos while cycling; a user can use gestures to record videos while hiking; a user can use gestures to record videos while attending classes or meetings; a user can use gestures to translate what the sensors capture without interrupting other processing; a user can use gestures to activate their AI assistant to help summarize what they are reading; a user can use gestures to activate their AI assistant to help debug the code they are writing; a user can use gestures to activate route navigation, for example, in situations where the user has visual impairments or is in traffic jams; a user can use gestures to have the sensors scan QR codes and make payments; and so on. The operations that AI glasses awakened through image processing algorithms and / or machine learning-based models can perform are not limited to these, and those skilled in the art will understand that other operations to meet user needs are also possible.

[0092] It should be noted that the various services provided to users do not necessarily require the AI ​​glasses to be activated to communicate with mobile terminals and / or remote servers, and the AI ​​glasses themselves may also process information.

[0093] The technology disclosed herein can be applied to various devices, such as smart glasses, AR devices with cameras, smart cameras, or devices with smart cameras. The specific form of the device is not limited to this disclosure, as long as it can be used to wake up the electronic device (e.g., its processor) by outputting a wake-up signal based on a captured image using a sensor containing an on-chip processor.

[0094] The foregoing has described various exemplary devices and methods according to embodiments of this disclosure. It should be understood that the operation or function of these devices can be combined with each other to achieve more or fewer operations or functions than described. Similarly, the operational steps of the methods can be combined with each other in any suitable order to similarly achieve more or fewer operations than described.

[0095] It should be understood that the machine-executable instructions in a machine-readable storage medium or program product according to embodiments of this disclosure can be configured to perform operations corresponding to the above-described device and method embodiments. Embodiments of the machine-readable storage medium or program product will be clear to those skilled in the art when referring to the above-described device and method embodiments, and therefore will not be described again. Machine-readable storage media and program products used to carry or include the above-described machine-executable instructions also fall within the scope of this disclosure. Such storage media may include, but are not limited to, floppy disks, optical disks, magneto-optical disks, memory cards, memory sticks, etc.

[0096] Furthermore, it should be understood that the aforementioned series of processes and devices can also be implemented via software and / or firmware. In the case of software and / or firmware implementation, programs constituting the software are installed from a storage medium or network onto a computer with a dedicated hardware architecture, such as the example electronic device 1300 shown in FIG. 9. When various programs are installed, the device is capable of performing various functions, etc. FIG. 9 is a block diagram illustrating an example of a schematic configuration of an electronic device according to an embodiment of the present disclosure.

[0097] In Figure 9, the Central Processing Unit (CPU) 1301 performs various processes according to the program stored in the Read-Only Memory (ROM) 1302 or the program loaded into the Random Access Memory (RAM) 1303 from the Storage Section 1308. The RAM 1303 also stores data required as needed when the CPU 1301 performs various processes.

[0098] CPU 1301, ROM 1302 and RAM 1303 are connected to each other via bus 1304. Input / output interface 1305 is also connected to bus 1304.

[0099] The following components can be connected to the input / output interface 1305: input section 1306, including a keypad, etc.; output section 1307, including a display and speakers, etc.; storage section 1308, including a memory card, etc.; and communication section 1309, including a network interface card (such as a LAN card), modem, etc. The communication section 1309 can perform communication processing via a network (such as the Internet). A sensor (not shown) can also be connected to the input / output interface 1305.

[0100] As needed, drive 1310 is also connected to input / output interface 1305. Removable media 1311, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on drive 1310 as needed, so that computer programs read from them can be installed into storage section 1308 as needed.

[0101] When the above series of processes are implemented by software, the program constituting the software is installed from a network such as the Internet or a storage medium such as removable media 1311.

[0102] Those skilled in the art will understand that such storage media are not limited to the removable medium 1311 shown, which stores a program and is distributed separately from the device to provide the program to the user. Examples of removable media 1311 include magnetic disks (including floppy disks (registered trademark)), optical disks (including optical disc read-only memory (CD-ROM) and digital versatile disks (DVD)), magneto-optical disks (including mini-disk (MD) (registered trademark)), and semiconductor memories. Alternatively, the storage medium may be ROM 1302, a hard disk included in storage section 1308, etc., containing a program and distributed to the user along with the device containing them.

[0103] Figure 10 is a block diagram illustrating another example of a schematic configuration of an electronic device 1600 to which the technology of this disclosure can be applied. The electronic device 1600 includes a processor 1601, a memory 1602, a storage device 1603, an external connection interface 1604, a camera device 1606, a sensor 1607, a microphone 1608, an input device 1609, a display device 1610, a speaker 1611, a wireless communication interface 1612, one or more antenna switches 1615, one or more antennas 1616, a bus 1617, a battery 1618, and an auxiliary controller 1619. In one implementation, the electronic device 1600 herein may correspond to a device with a camera, such as smart glasses.

[0104] Processor 1601 may be, for example, a CPU or a System-on-a-Chip (SoC), and controls the application layer and other functions of electronic device 1600. Memory 1602 includes RAM and ROM, and stores data and programs executed by processor 1601. Storage device 1603 may include storage media such as semiconductor memory and hard disk. External connection interface 1604 is an interface for connecting external devices (such as memory cards and Universal Serial Bus (USB) devices) to electronic device 1600.

[0105] The camera device 1606 includes an image sensor (such as a charge-coupled device (CCD) and complementary metal-oxide-semiconductor (CMOS)) and generates captured images. The sensor 1607 may include a set of sensors, such as a measurement sensor, a gyroscope sensor, a magnetometer sensor, and an accelerometer sensor. The microphone 1608 converts sound input to the electronic device 1600 into an audio signal. The input device 1609 includes, for example, a touch sensor, keypad, keyboard, buttons, or switches configured to detect touches on the screen of the display device 1610 and receives operations or information input from the user. The display device 1610 includes a screen (such as a liquid crystal display (LCD) and an organic light-emitting diode (OLED) display) and displays the output image of the electronic device 1600. The speaker 1611 converts the audio signal output from the electronic device 1600 into sound.

[0106] Wireless communication interface 1612 supports any cellular communication scheme (such as LTE, LTE-Advanced, and NR) and performs wireless communication. Wireless communication interface 1612 typically includes, for example, a BB processor 1613 and RF circuitry 1614. BB processor 1613 can perform, for example, encoding / decoding, modulation / demodulation, and multiplexing / demultiplexing, and performs various types of signal processing for wireless communication. Meanwhile, RF circuitry 1614 can include, for example, mixers, filters, and amplifiers, and transmits and receives wireless signals via antenna 1616. Wireless communication interface 1612 can be a single chip module on which BB processor 1613 and RF circuitry 1614 are integrated. As shown in Figure 10, wireless communication interface 1612 can include multiple BB processors 1613 and multiple RF circuits 1614. Although Figure 10 shows an example where wireless communication interface 1612 includes multiple BB processors 1613 and multiple RF circuits 1614, wireless communication interface 1612 can also include a single BB processor 1613 or a single RF circuitry 1614.

[0107] In addition to cellular communication schemes, wireless communication interface 1612 can support other types of wireless communication schemes, such as short-range wireless communication schemes, near-field communication schemes, and wireless local area network (LAN) schemes. In this case, wireless communication interface 1612 may include a BB processor 1613 and RF circuitry 1614 for each wireless communication scheme.

[0108] Each of the antenna switches 1615 switches the connection destination of the antenna 1616 among multiple circuits (e.g., circuits for different wireless communication schemes) included in the wireless communication interface 1612.

[0109] Each of the antennas 1616 includes one or more antenna elements (such as multiple antenna elements included in a MIMO antenna) and is used for transmitting and receiving wireless signals through the wireless communication interface 1612. As shown in Figure 10, the electronic device 1600 may include multiple antennas 1616. Although Figure 10 shows an example in which the electronic device 1600 includes multiple antennas 1616, the electronic device 1600 may also include a single antenna 1616.

[0110] Furthermore, the electronic device 1600 may include an antenna 1616 for each wireless communication scheme. In this case, the antenna switch 1615 may be omitted from the configuration of the electronic device 1600.

[0111] Bus 1617 connects processor 1601, memory 1602, storage device 1603, external connection interface 1604, camera device 1606, sensor 1607, microphone 1608, input device 1609, display device 1610, speaker 1611, wireless communication interface 1612, and auxiliary controller 1619 to each other. Battery 1618 supplies power to the various blocks of electronic device 1600 shown in FIG. 10 via feeders, which are partially shown as dashed lines in the figure. Auxiliary controller 1619 operates the minimum necessary functions of electronic device 1600, for example, in sleep mode.

[0112] Figure 11 is a block diagram illustrating yet another example of a schematic configuration of an electronic device 1720 according to an embodiment of the present disclosure. The electronic device 1720 includes one or more of the following: a processor 1721, a memory 1722, a Global Positioning System (GPS) module 1724, a sensor 1725, a data interface 1726, a content player 1727, a storage medium interface 1728, an input device 1729, a display device 1730, a speaker 1731, a wireless communication interface 1733, one or more antenna switches 1736, one or more antennas 1737, and a battery 1738. In one implementation, the electronic device 1720 herein may correspond to a device with a camera, such as smart glasses.

[0113] The processor 1721 can be, for example, a CPU or a SoC, and controls various functions of the electronic device 1720. The memory 1722 includes RAM and ROM, and stores data and programs executed by the processor 1721.

[0114] GPS module 1724 uses GPS signals received from GPS satellites to measure the location (such as latitude, longitude, and altitude) of electronic device 1720. Sensor 1725 may include a set of sensors, such as a gyroscope sensor, a geomagnetic sensor, and an air pressure sensor, to collect relevant information. Sensor 1725 may also include an image sensor, for example, for waking electronic device 1720 from sleep mode based on captured image data. Data interface 1726 is connected to, for example, a data network 1741 via a terminal not shown.

[0115] Content player 1727 reproduces content stored on a storage medium (such as a memory card), which is inserted into storage medium interface 1728. Input device 1729 includes, for example, a touch sensor, button, or switch configured to detect touch on the screen of display device 1730, and receives operations or information input from the user. Display device 1730 includes a screen such as an LCD or OLED display and displays images or reproduced content for various functions. Speaker 1731 outputs sound or reproduced content for various functions.

[0116] The wireless communication interface 1733 supports any cellular communication scheme (such as LTE, LTE-Advanced, and NR) and performs wireless communication. The wireless communication interface 1733 typically includes, for example, a BB processor 1734 and RF circuitry 1735. The BB processor 1734 can perform, for example, encoding / decoding, modulation / demodulation, and multiplexing / demultiplexing, and performs various types of signal processing for wireless communication. Meanwhile, the RF circuitry 1735 can include, for example, mixers, filters, and amplifiers, and transmits and receives wireless signals via antenna 1737. The wireless communication interface 1733 can also be a chip module on which the BB processor 1734 and RF circuitry 1735 are integrated. As shown in Figure 11, the wireless communication interface 1733 can include multiple BB processors 1734 and multiple RF circuits 1735. Although Figure 11 shows an example where the wireless communication interface 1733 includes multiple BB processors 1734 and multiple RF circuits 1735, the wireless communication interface 1733 can also include a single BB processor 1734 or a single RF circuitry 1735.

[0117] In addition to cellular communication schemes, the wireless communication interface 1733 can support other types of wireless communication schemes, such as short-range wireless communication schemes, near-field communication schemes, and wireless LAN schemes. In this case, for each wireless communication scheme, the wireless communication interface 1733 may include a BB processor 1734 and an RF circuit 1735.

[0118] Each of the antenna switches 1736 switches the connection destination of the antenna 1737 among multiple circuits (such as circuits for different wireless communication schemes) included in the wireless communication interface 1733.

[0119] Each of the antennas 1737 includes one or more antenna elements (such as multiple antenna elements included in a MIMO antenna) and is used for transmitting and receiving wireless signals through the wireless communication interface 1733. As shown in Figure 11, the electronic device 1720 may include multiple antennas 1737. Although Figure 11 shows an example in which the electronic device 1720 includes multiple antennas 1737, the electronic device 1720 may also include a single antenna 1737.

[0120] Furthermore, the electronic device 1720 may include an antenna 1737 for each wireless communication scheme. In this case, the antenna switch 1736 can be omitted from the configuration of the electronic device 1720.

[0121] Battery 1738 may be a rechargeable battery and may supply power to various blocks of electronic device 1720 shown in FIG11 via a feeder, which is partially shown as a dashed line in the figure.

[0122] Exemplary embodiments of the present disclosure have been described above with reference to the accompanying drawings; however, the present disclosure is by no means limited to the examples described above. Various changes and modifications can be made by those skilled in the art within the scope of the appended claims, and it should be understood that such changes and modifications naturally fall within the technical scope of the present disclosure.

[0123] For example, the multiple functions included in one unit in the above embodiments can be implemented by separate devices. Alternatively, the multiple functions implemented by multiple units in the above embodiments can be implemented by separate devices respectively. In addition, one of the above functions can be implemented by multiple units. Needless to say, such a configuration is included within the scope of the present disclosure.

[0124] In this specification, the steps described in the flowchart include not only processes executed sequentially in the stated order, but also processes executed in parallel or individually, rather than necessarily sequentially. Furthermore, even within the steps of sequential processing, needless to say, the order can be appropriately altered.

[0125] While this disclosure and its advantages have been described in detail, it should be understood that various changes, substitutions, and modifications can be made without departing from the spirit and scope of this disclosure as defined by the appended claims. Furthermore, the terms "comprising," "including," or any other variations thereof used in embodiments of this disclosure are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

Claims

1. A method for waking up an electronic device, comprising: Image data is captured via sensors when the electronic device is in sleep mode; as well as The electronic device is woken up by a wake-up signal output by the sensor based on the captured image data.

2. The method according to claim 1, wherein, The sensor includes a photoelectric conversion unit and an on-chip processor. The photoelectric conversion unit is configured to capture image data by converting incident light signals into electrical signals, and The on-chip processor is configured to output a wake-up signal based on captured image data.

3. The method according to claim 2, wherein, The on-chip processor is configured to output a wake-up signal in response to the detection of a specific target from captured image data, the specific target being pre-specified for waking up the electronic device.

4. The method according to claim 3, wherein, The specific target includes at least one of the following: gestures and predetermined objects.

5. The method according to claim 2, wherein, The on-chip processor includes a machine learning-based model whose input includes captured image data and whose output includes a wake-up signal.

6. The method according to claim 5, wherein, The model includes a deep neural network (DNN) configured to process captured image data to generate a processing result including a trigger signal related to waking up the electronic device, wherein the trigger signal is used as the wake-up signal when a predetermined condition is met.

7. The method according to claim 6, wherein, The DNN is implemented based on a pruned MobileNet SSD.

8. The method according to claim 5, wherein, The output of the model also includes feature data related to the image data, which is used to enable the electronic device to provide corresponding operations to the user after being woken up.

9. The method according to claim 8, wherein, The corresponding operation includes one or more of the following: Image capture is performed using the sensor; Provide information for decision-making; Provide information to the user terminal to launch the relevant application; and Communicate with a remote server to enable AI services.

10. The method according to claim 2, wherein, The sensor has a first mode when the electronic device is in a sleep state and a second mode after the electronic device is woken up. In the first mode, the photoelectric conversion unit outputs the captured image data to the on-chip processor, so that the on-chip processor determines whether to output a wake-up signal based on the captured image data. In the second mode, the photoelectric conversion unit outputs the captured image data to a device external to the sensor. In the first mode, the sensor consumes less power than in the second mode.

11. The method of claim 10, further comprising: After waking up the electronic device, the sensor is switched to the second mode so that the photoelectric conversion unit stops outputting captured image data to the on-chip processor and instead outputs it to a device outside the sensor.

12. The method of claim 11, further comprising: When the sensor is in the second mode, the on-chip processor is no longer used.

13. The method of claim 12, wherein stopping the use of the on-chip processor includes selectively stopping the use of the on-chip processor.

14. The method according to claim 1, further comprising: Based on the feature data related to the captured image data output by the sensor, the electronic device can provide corresponding operations to the user after being woken up.

15. The method according to claim 1, wherein, The sensor is part of the electronic device, which is smart glasses, an AR device with a camera, a smart camera, or a device with a smart camera.

16. A sensor, comprising: The photoelectric conversion unit is configured to capture image data by converting incident light signals into electrical signals; as well as The on-chip processor is configured to output a wake-up signal based on captured image data to wake up the electronic device.

17. The sensor according to claim 16, wherein, The on-chip processor is configured to output a wake-up signal in response to the detection of a specific target from captured image data, the specific target being pre-specified for waking up the electronic device.

18. The sensor according to claim 16, wherein, The on-chip processor includes a machine learning-based model whose input includes captured image data and whose output includes a wake-up signal.

19. The sensor of claim 18, wherein the model comprises a deep neural network (DNN) configured to process captured image data to generate a processing result including a trigger signal relating to waking up the electronic device, wherein the trigger signal is used as the wake-up signal when a predetermined condition is met.

20. The sensor according to claim 19, wherein, The DNN is implemented based on a pruned MobileNet SSD.

21. The sensor according to claim 18, wherein, The output of the model also includes feature data related to the image data, which is used to enable the electronic device to provide corresponding operations to the user after being woken up.

22. The sensor according to claim 16, wherein, The sensor has a first mode when the electronic device is in a sleep state and a second mode after the electronic device is woken up. In the first mode, the photoelectric conversion unit outputs the captured image data to the on-chip processor, so that the on-chip processor determines whether to output a wake-up signal based on the captured image data. In the second mode, the photoelectric conversion unit outputs the captured image data to a device external to the sensor. In the first mode, the sensor consumes less power than in the second mode.

23. The sensor according to claim 22, wherein, In the second mode, the photoelectric conversion unit stops outputting the captured image data to the on-chip processor.

24. The sensor according to claim 23, wherein, In the second mode, the on-chip processor stops working.

25. An electronic device, comprising: The sensor according to any one of claims 16 to 24; as well as A device configured to receive image data output from the sensor.

26. A device for waking up an electronic device, comprising: processor; as well as The memory includes computer program instructions, wherein the memory and computer program instructions are configured to cause the device to perform the following operations via the processor: Image data captured by sensors included in an electronic device is processed using machine learning-based models or image processing operations; and The electronic device is activated based on the processing result.

27. A computer-readable storage medium having stored thereon computer program instructions, which, when executed by a processing device, cause the processing device to perform the following operations: Image data captured by sensors included in an electronic device is processed using machine learning-based models or image processing operations; and The electronic device is activated based on the processing result.