Method for assisting with object recognition, and electronic device and sensor
By integrating a photoelectric conversion unit and a digital processing module into smart glasses, and using digital processing and deep neural network algorithms to detect the location of the QR code, generating a large-area image and transmitting the recognition image, the problem of short QR code recognition distance and manual positioning in smart glasses is solved, improving recognition convenience and reducing power consumption.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SONY SEMICON SOLUTIONS CORP
- Filing Date
- 2026-01-15
- Publication Date
- 2026-07-30
AI Technical Summary
Existing smart glasses devices have a limited range when recognizing QR codes, requiring manual positioning of the QR code, which affects the user experience.
By integrating a photoelectric conversion unit and a digital processing module into smart glasses, the digital processing module is used to perform image processing and deep neural network algorithms to detect the location information of the QR code, generate a large-area image, and transmit the recognition image to the recognition device.
It increases the distance for QR code recognition, avoids the tedious operation of manual positioning, improves the user experience, and reduces the power consumption of electronic devices.
Smart Images

Figure CN2026072749_30072026_PF_FP_ABST
Abstract
Description
Methods, electronic devices, and sensors for assisting in target identification
[0001] Cross-references to related applications
[0002] This disclosure claims priority and benefits to Chinese Patent Application No. 202510123844.3, filed with the State Intellectual Property Office of the People's Republic of China on January 24, 2025, the entire contents of which are hereby incorporated herein by reference. Technical Field
[0003] This disclosure relates to the field of information processing, and more specifically, to methods, electronic devices, sensors, program products, apparatuses, and storage media that utilize electronic devices to assist in the identification of targets in the field of information processing. Background Technology
[0004] QR codes are ubiquitous in our daily lives, enabling various functions such as ordering food, making reservations, and processing payments. The typical way to use QR codes is by scanning them with a smartphone. However, this process requires multiple steps: taking out the phone, opening the relevant application, and then changing the distance or angle to scan the QR code, which requires a significant amount of human intervention.
[0005] To simplify this process and minimize user workload, wearable devices such as smart (AI) glasses can be used to scan QR codes. These devices can automatically launch the corresponding application and navigate directly to the designated interface, such as a menu or order page. This reduces manual steps, increases efficiency, and significantly improves the overall user experience.
[0006] Currently, some brands of smart glasses offer QR code scanning functionality. However, due to the limitations of these devices' camera resolution, the glasses cannot scan QR codes from a distance like smartphones. Instead, they can only scan QR codes at a relatively close distance. This significantly impacts the user experience, as users are forced to unconsciously bring their heads close to the QR code to complete the scan. Summary of the Invention
[0007] Therefore, there is a need to provide a new target recognition method that can improve the target recognition distance of electronic devices such as smart glasses and enhance the user experience.
[0008] One aspect of this disclosure relates to a method for assisting in target identification using an electronic device, the electronic device including a sensor and a processor, the method comprising: the sensor capturing a first image of the target and detecting position information of the target based on the first image; the sensor capturing a second image of the target and sending the second image and the position information to the processor, the position of the target in the first image corresponding to its position in the second image; and the processor processing the second image according to the position information and generating a third image for identifying the target, the third image corresponding to a region of interest including the target and its surroundings.
[0009] Another aspect of this disclosure relates to an electronic device including a sensor and a processor. The sensor includes a photoelectric conversion unit and a digital processing module. The sensor is configured to capture a first image of a target via the photoelectric conversion unit and to detect position information of the target based on the first image received from the photoelectric conversion unit via the digital processing module. The sensor is also configured to capture a second image of the target via the photoelectric conversion unit and to send the second image and the position information to the processor. The position of the target in the first image corresponds to its position in the second image. The processor is further configured to process the second image according to the position information and generate a third image for identifying the target, the third image corresponding to a region of interest including the target and its surroundings.
[0010] Another aspect of this disclosure relates to a sensor comprising: a photoelectric conversion unit configured to convert a received light signal into an image signal; and a digital processing module configured to process the image signal received from the photoelectric conversion unit to obtain location information of a target, the location information being configured to be used by an electronic device including the sensor to generate an identification image for identifying the target.
[0011] Another aspect of this disclosure relates to a program product comprising computer instructions configured to perform the following operations when executed by a processor: processing image signals from an electronic device to obtain location information of a target. The location information is configured to be used by the electronic device to generate an identification image for identifying the target.
[0012] Another aspect of this disclosure relates to an apparatus for assisting in target identification, comprising: a processor; and a memory storing computer instructions. The computer instructions are configured to, when executed by the processor, perform the following operations: process image signals from an electronic device to obtain location information of the target. The location information is configured to be used by the electronic device to generate an identification image for identifying the target.
[0013] Another aspect of this disclosure relates to a computer-readable storage medium having stored thereon computer program instructions configured to perform the following operations when executed by a processor: processing image signals from an electronic device to obtain location information of the target. The location information is configured to be used by the electronic device to generate an identification image for identifying the target.
[0014] The above overview is provided to summarize some exemplary embodiments to provide a basic understanding of the aspects of the subject matter described herein. Therefore, the features described above are merely examples and should not be construed as narrowing the scope or spirit of the subject matter described herein in any way. Other features, aspects, and advantages of the subject matter described herein will become apparent from the following detailed description, taken in conjunction with the accompanying drawings. Attached Figure Description
[0015] A better understanding of this disclosure can be obtained by considering the following detailed description of the embodiments in conjunction with the accompanying drawings. The same or similar reference numerals are used in the drawings to denote the same or similar parts. The drawings, together with the following detailed description, are incorporated in and form a part of this specification to illustrate embodiments of the disclosure and explain the principles and advantages of the disclosure. Wherein:
[0016] Figure 1 is a schematic diagram illustrating smart glasses as an example of an electronic device in the related technology.
[0017] Figure 2 is a schematic diagram illustrating an example of a scenario in the related technology where smart glasses assist a smartphone in scanning a QR code in front of a user.
[0018] Figure 3 is a schematic diagram illustrating an example configuration of an electronic device according to an embodiment of the present disclosure.
[0019] Figure 4 is a flowchart illustrating a method for identifying a target using an electronic device-assisted identification device according to an embodiment of the present disclosure.
[0020] Figure 5 is a schematic diagram showing examples of the original image, the large-area-ratio image, and the merged image.
[0021] Figure 6 is a flowchart illustrating another method for identifying a target using an electronic device-assisted identification device according to an embodiment of the present disclosure.
[0022] Figure 7 is a schematic diagram illustrating an example of a scenario in which smart glasses assist a smartphone in scanning a QR code in front of a user according to an embodiment of the present disclosure.
[0023] Figure 8 is a block diagram illustrating an example of a schematic configuration of an electronic device according to an embodiment of the present disclosure.
[0024] Figure 9 is a block diagram illustrating another example of a schematic configuration of an electronic device to which the techniques of this disclosure can be applied.
[0025] Figure 10 is a block diagram illustrating yet another example of a schematic configuration of an electronic device according to an embodiment of the present disclosure.
[0026] While the embodiments described in this disclosure may be readily modified and alternatively implemented, specific embodiments thereof are shown by way of example in the accompanying drawings and are described in detail herein. However, it should be understood that the drawings and the detailed description thereof are not intended to limit the embodiments to the specific forms disclosed, but rather are intended to cover all modifications, equivalents, and alternatives that fall within the spirit and scope of the claims. Detailed Implementation
[0027] The following description illustrates representative applications of the devices and methods described herein. These examples are provided merely to provide context and aid in understanding the described embodiments. Therefore, it will be apparent to those skilled in the art that the embodiments described below can be practiced without some or all of the specific details provided. In other instances, well-known process steps have not been described in detail to avoid unnecessarily obscuring the described embodiments. Other applications are also possible, and the scope of this disclosure is not limited to these examples.
[0028] Referring first to Figure 1, a schematic diagram of smart glasses, an example of an electronic device in the related art, is shown in Figure 1.
[0029] Smart glasses can integrate electronic components such as processors, memory, cameras, microphones, and displays. When a user wears smart glasses, they remain in standby mode to conserve power. After the user wakes the smart glasses, the processor starts up from sleep mode and can then provide various services, including virtual reality (VR), augmented reality (AR), or QR code recognition, based on information from the microphone, camera, etc., thereby expanding the user's capabilities and sensory experience. As shown on the left side of Figure 1, a camera can be mounted at the end of the temple of the smart glasses, as indicated by the circled area. The right side of Figure 1 shows the camera as seen from the front of the smart glasses.
[0030] Figure 2 is a schematic diagram illustrating an example of a scenario in the related technology where smart glasses assist a smartphone in scanning a QR code in front of a user.
[0031] Currently, to enable smart glasses to assist smartphones in scanning QR codes in front of users, users can generate trigger signals through buttons, voice, or gestures to wake up the smart glasses. Note that, as shown in Figure 2, the QR code on the whiteboard is represented by a thick black outline in the recognition image.
[0032] Specifically, in response to trigger signals generated by button clicks, user voice, or user gestures, the processor in the smart glasses, such as a SoC (System-on-a-Chip), MCU, or AP, sends control commands to the smart glasses' image sensor via I2C (Integrated Circuit Bus) to wake up the image sensor. Once woken up, the image sensor will capture the image in front of it.
[0033] At this time, the image sensor continuously captures images of the QR code in front of it to output images or videos. The image sensor is in streaming mode at this time.
[0034] Meanwhile, the smart glasses' processor transmits digital image data from the image sensor to a recognition device such as a smartphone via a wired interface (such as MIPI, CSI, or USB) or a wireless interface (such as WIFI, Bluetooth, Zigbee, or cellular communication technology).
[0035] After receiving image data from the smart glasses, the recognition device opens the relevant application (APP) to scan or recognize targets (e.g., QR codes) in the image.
[0036] Tests have revealed that when the raw images captured by smart glasses are sent directly to a recognition device for QR code scanning, there are certain distance limitations.
[0037] Taking a currently available pair of smart glasses as an example, when the smart glasses capture a 5x5 cm QR code at a distance of approximately 1.2 meters, a smartphone cannot recognize the QR code in the raw image received from the smart glasses. However, if the location of the QR code is manually determined and a region of interest (ROI) containing the QR code is set before transmitting the image to the smartphone, the recognition distance increases to approximately 3.2 meters, thus improving the recognition distance by about 2 meters. Clearly, this manual positioning also has the disadvantages of being cumbersome and negatively impacting the user experience.
[0038] To address the aforementioned issues, this disclosure proposes a scheme for identifying targets using electronic devices as an auxiliary identification device. This not only reduces the distance limitations of target identification but also avoids the tedious manual positioning process and improves the user experience, thereby enhancing the convenience and intelligence of target identification.
[0039] Figure 3 is a schematic diagram illustrating a configuration example of an electronic device (e.g., smart glasses) 10 according to an embodiment of the present disclosure.
[0040] As shown in Figure 3, the electronic device 10 may include a processor 11 and an image sensor 12.
[0041] In embodiments according to this disclosure, the image sensor 12 may include a photoelectric conversion unit 21 and a digital processing module 22. The photoelectric conversion unit can convert light signals received from the outside into electrical signals, i.e., image signals. The digital processing module 22 receives the electrical signals from the photoelectric conversion unit and can perform various processing on the electrical signals. For example, the processing by the digital processing module 22 may include cropping or digital upscaling, binning, DNN algorithms, etc., as described below.
[0042] The functions and configurations of the various components of the electronic device will be described in detail below in conjunction with methods 100 and 200 for identifying targets using electronic devices to assist identification devices.
[0043] Figure 4 is a flowchart illustrating a method 100 for identifying a target using an electronic device-assisted identification device according to an embodiment of the present disclosure. Although the following description uses smart glasses for scanning QR codes as an example of the electronic device 10 and a smartphone as an example of the identification device, the present disclosure is not limited thereto. Those skilled in the art will readily understand that the electronic device can be other information processing devices with image sensors, such as AR devices with cameras (head-mounted or wrist-worn AR devices), smart cameras, or devices with smart cameras (e.g., smart home appliances, smart toys, etc.). Furthermore, those skilled in the art will readily understand that the identification device can be a personal mobile terminal, tablet computer, personal computer, other smart devices with target recognition capabilities, etc.
[0044] According to this disclosure, before providing an image for target recognition to a smartphone, the smart glasses need to detect the location information of a QR code, serving as a target example, within the image. Subsequently, the smart glasses acquire a recognition image based on the detected location information and provide the recognition image to the smartphone for scanning or recognition.
[0045] Specifically, firstly, in step S110, the smart glasses capture an original image of a QR code located in front of the user wearing the smart glasses via the photoelectric conversion unit 21 in its image sensor 12. According to embodiments of this disclosure, the image sensor 12 may correspond to the camera of the smart glasses or be integrated as a component into the camera. Note that the original image at this time corresponds to the "first image" according to this disclosure.
[0046] It should be noted that the image sensor 12 of the smart glasses is currently operating in positioning mode. To switch the image sensor 12 to positioning mode, the user wearing the smart glasses can trigger the smart glasses via a button, voice, or gesture. In response to this trigger, the processor 11 of the smart glasses is awakened and thus switches from sleep mode to active mode. Next, the processor 11 sends a control command to the image sensor 12 to switch it to positioning mode. After receiving the control command, the image sensor 12 switches to positioning mode. At the same time, the processor 11 returns to sleep mode.
[0047] Next, in step S111, the image sensor 12 processes the original image by the digital processing module 22 with a predetermined area ratio increase coefficient, and thus obtains an image with a large area ratio.
[0048] Typically, the distance between the smart glasses and the QR code is too far, resulting in the QR code occupying too small a proportion in the original image captured by the smart glasses. Consequently, the image sensor 12 cannot effectively detect the QR code and its location information in the original image. Therefore, in order to better focus on the QR code and improve the efficiency and accuracy of subsequent processing, the image sensor 12 can process the original image through the digital processing module 22 to obtain an image that significantly increases the area proportion of the QR code in the entire image. The image obtained at this time will be referred to below as the large-area-proportion image. Note that this large-area-proportion image corresponds to the "fourth image" according to this disclosure.
[0049] Specifically, in step S111, the image sensor 12 can increase the area ratio of the QR code in the entire image by using a predetermined area ratio enhancement factor through the digital processing module 22, thereby obtaining a large area ratio image. Examples of digital processing methods include cropping, digital magnification, etc. Other software methods will also be conceived by those skilled in the art, as long as the software method can increase the area ratio of the QR code in the entire image.
[0050] According to this disclosure, when the location information of the target is detected for the first time, the predetermined area proportion enhancement factor can be set to 1. That is, no cropping or digital magnification is performed on the original image at this time, and the large area proportion image at this time corresponds to the original image.
[0051] Next, in step S112, the image sensor 12 detects the location information of the QR code in the large area proportion image.
[0052] Therefore, the image sensor 12 may integrate a digital processing module 22 for image processing and detecting the location information of the QR code in the image. In this case, the "image sensor" in the sense of this disclosure no longer only has an image capture function, but also has a digital processing function for processing images.
[0053] Image sensor 12 can use digital processing module 22 to detect the position of QR code in a large-area image using algorithms such as edge detection, dilation, and erosion. These common algorithms are well known to those skilled in the art and will not be described in detail herein.
[0054] Preferably, a Deep Neural Network (DNN) algorithm can also be used to perform QR code location information detection. Compared with common algorithms such as edge detection, dilation, and erosion algorithms, the DNN algorithm can locate the QR code faster and more accurately without needing to perform a pixel-by-pixel search across the entire image. Furthermore, the DNN algorithm has a faster search speed and higher robustness, because common algorithms for detecting QR code edges are more susceptible to environmental influences and have lower recognition efficiency and accuracy.
[0055] According to embodiments of this disclosure, the DNN algorithm can be implemented based on a DNN model. In this case, the DNN model can be integrated into the digital processing module 22. The DNN model in the digital processing module 22 receives a large-area proportion image as input. The digital processing module 22 uses the DNN model to detect the position of the QR code in the large-area proportion image.
[0056] The DNN model will now be described in detail.
[0057] DNN models can be obtained through pre-training. A DNN model can be trained to output the location information of a target based on a large proportion of an image as input. The training of the DNN model can employ any suitable supervised, unsupervised, semi-supervised, or reinforcement learning technique. For example, diverse QR code images can be collected in advance, labeled, and data-augmented for the DNN model to learn from.
[0058] DNN models can be implemented using various structures, such as one or more convolutional layers, one or more fully connected layers, etc., as long as they can determine whether a target exists in the input image and its location within the input image based on the input image fed into the DNN model. According to embodiments of this disclosure, a DNN model for target detection can be used to output information about whether a target has been detected and its possible location. For example, a classifier containing fully connected layers can be used to implement the DNN model to detect whether a target is present in the input image and determine its location. Another example is the implementation of a DNN model based on a pruned MobileNet SSD (Single Shot MultiBox Detector). That is, a DNN model is constructed by pruning an existing MobileNet SSD model to detect a target at a certain location in the input image. The pruned DNN model can have a simpler network structure, less computation and fewer model parameters, and faster processing speed, which is beneficial for use in resource-constrained devices such as smart glasses for target detection and localization. Of course, the DNN model implemented based on the pruned MobileNet SSD (Single Shot MultiBox Detector) is merely an example. DNN models based on other similar object detection algorithms such as the YOLO (You Only Look Once) series or the Fast-RCNN series can also be used. This disclosure does not impose any particular restrictions on the implementation of the DNN model, as long as these models can achieve object detection in the input image.
[0059] Specifically, the DNN model generates multiple anchor boxes at various locations in the input image. These anchor boxes are typically generated on different feature layers to ensure the model can detect objects at different scales. The DNN model determines which anchor box corresponds to the target based on its overlap with the ground truth bounding boxes. If the overlap exceeds a certain threshold, the anchor box is considered to correspond to the target. Furthermore, the DNN model uses regression to predict the offset of the anchor box corresponding to the target, thus obtaining the target's bounding box. At this point, the DNN model has obtained the location information of the QR code in the original image, i.e., the information of the bounding box coordinates containing the QR code. Conversely, if the overlap of all anchor boxes is below the aforementioned threshold, detection fails, meaning the target was not detected in the original image.
[0060] Next, in step S113, it is determined whether the image sensor 12 has detected the target in the image with a large area ratio and thus obtained the target's location information. If the determination result is positive, method 100 proceeds to step S114.
[0061] In step S114, in response to the acquisition of location information, the processor 11 is awakened. At this time, the processor 11 sends a control command to the image sensor 12 to switch the image sensor 12 to streaming mode. The image sensor 12 of the smart glasses recaptures the original image of the QR code via the photoelectric conversion unit 21, and transmits the captured original image and the aforementioned location information to the processor 11. Note that the original image at this time corresponds to the "second image" according to this disclosure.
[0062] Since the user wearing smart glasses is constantly looking at the QR code during target recognition, the position of the QR code in the original image captured at this time basically corresponds to the position in the original image captured in step S110.
[0063] Next, in step S115, the processor 11 processes the original image captured in step S114 based on the position information received from the image sensor 12 to obtain a recognition image for recognition by the recognition device. Note that the recognition image at this time corresponds to the "third image" according to this disclosure.
[0064] According to this disclosure, the recognition image is an image corresponding to a region of interest (ROI) containing the target, such that the area of the QR code in the recognition image is greater than its area in the original image. In the technical sense of this disclosure, the ROI is the area including the QR code and its surroundings. The QR code may be centrally positioned within this area. It should be understood that the area of the QR code within the ROI should be higher than a certain threshold, enabling the backend device to scan the QR code in the recognition image or to recognize the information contained within the QR code. For example, the area of the QR code within the ROI should be at least greater than 1%, particularly greater than 2%, preferably greater than 5%, and most preferably greater than 10%, depending on the clarity of the QR code content in the recognition image. This disclosure does not impose any particular limitation on the value of this area percentage, as long as the backend device can scan the QR code in the recognition image or recognize the information contained within the QR code.
[0065] Image sensor 12 can use digital processing methods via digital processing module 22 to acquire a recognition image that increases the area of the QR code within the entire image. Examples of digital processing methods include cropping, digital magnification, etc. Other software methods will also be conceived by those skilled in the art, as long as the software method can increase the area of the QR code within the entire image.
[0066] Next, after acquiring the recognition image, in step S116, the processor 11 of the smart glasses transmits the recognition image to the smartphone via a wired interface (such as MIPI, CSI or USB) or a wireless interface (such as WIFI, Bluetooth, Zigbee or cellular communication technology, etc.) for the smartphone to scan or recognize.
[0067] After receiving the image from the smart glasses, the smartphone can open the relevant application to scan the QR code in the image or identify the information contained in the QR code.
[0068] On the other hand, if the determination result in step S113 is negative, then method 100 proceeds to step S117.
[0069] In step S117, it is determined whether the detection termination condition is met. For example, the digital processing module 22 of the image sensor 12 may be set with a threshold for the area ratio enhancement factor. If the area ratio enhancement factor reaches or exceeds the threshold, the detection termination condition can be determined to be met. Alternatively, the digital processing module 22 of the image sensor 12 may be set with a predetermined detection time. If the predetermined detection time is reached, the detection termination condition can be determined to be met. As another example, the digital processing module 22 of the image sensor 12 may be set with multiple area ratio enhancement factors (e.g., 1, 4, 8, and 16). If all area ratio enhancement factors have been used, the detection termination condition can be determined to be met. Of course, those skilled in the art can conceive of other detection termination conditions as needed, and this disclosure does not impose any particular limitations on the detection termination conditions.
[0070] If the determination result in step S117 is negative, then method 100 proceeds to step S118.
[0071] In step S118, the digital processing module 22 of the image sensor 12 changes the area ratio enhancement factor. For example, the digital processing module 22 of the image sensor 12 increases the area ratio enhancement factor from the original 1 to 4.
[0072] Subsequently, method 100 returns to step S111. In step S111, the image sensor 12 can acquire a new large-area-ratio image by using the digital processing module 22 with a new area-ratio enhancement coefficient. At this time, the area ratio of the QR code in the entire image is further increased.
[0073] With a new area occupancy enhancement factor of 4, the size of the large-area-occupancy image is 1 / 4 of the original image size. In other words, compared to the original image, the QR code occupies a larger area in the large-area-occupancy image acquired at this point.
[0074] Here, the larger the area occupancy enhancement factor, the smaller the size of the large-area-occupancy image; conversely, the smaller the area occupancy enhancement factor, the larger the size of the large-area-occupancy image. Therefore, a larger area occupancy enhancement factor means that a greater distance can be seen.
[0075] On the other hand, if the determination result in step S117 is positive, then method 100 proceeds to step S119.
[0076] In step S119, if the image sensor 12 ultimately fails to detect the QR code, it can return to a sleep state. At this time, the processor 11 remains in a sleep state. Therefore, QR code location detection can also reduce accidental activation by voice, buttons, or gestures.
[0077] Note that when using a DNN algorithm based on a DNN model to detect the location information of a target, the input image (e.g., the large area image mentioned in step S111) input into the DNN model needs to match the DNN model.
[0078] Typically, in order to make the input image match the DNN model, the image sensor 12 can perform a merging process on the input image through the digital processing module 22 before detecting the location information of the target in the input image, so that the input image matches the DNN model. The input image at this time can also be called a "merged image".
[0079] Merging is typically achieved by combining multiple adjacent pixels into a single pixel. Besides enabling model matching, merging helps the model focus more efficiently on the target QR code and avoids interference from irrelevant areas.
[0080] For example, image merging processing can include horizontal merging (Binning in X), vertical merging (Binning in Y), and combinations thereof.
[0081] In horizontal pixel merging, multiple pixels are combined into one pixel along the horizontal direction. For example, merging two adjacent horizontal pixels into one Pixel halves the image width. Obviously, the horizontal merging factor is not limited to 2 and can be greater than 2. On the other hand, in vertical pixel merging, multiple pixels are combined into one pixel along the vertical direction. For example, merging two adjacent vertical pixels into one Pixel halves the image height. Obviously, the vertical merging factor is not limited to 2 and can be greater than 2. Merging can also be applied simultaneously in both the horizontal and vertical directions, resulting in a faster reduction in image resolution.
[0082] Suppose an image has a resolution of 640x480 pixels. If this image undergoes a pixel merging process, where every two adjacent pixels are merged into one, the image resolution will become 320x240 pixels. Specifically, the pixel values of each 2x2 pixel block in the image are merged in some way (such as averaging or summing) to generate a new pixel.
[0083] Furthermore, the merging coefficient can also vary gradually or stepwise along the merging direction. For example, the merging coefficient can be chosen to be larger in the peripheral parts of the image, while it can be chosen to be smaller in the center of the image.
[0084] After performing the merging process, the image sensor 12 can detect the location information of the QR code in the input image through the digital processing module 22.
[0085] Therefore, according to the method 100 of this disclosure, by using electronic devices to assist in the identification of targets, not only can the distance limitation of the above-mentioned target identification be reduced, but also the tedious operation of manual positioning can be avoided and the user experience can be improved, thereby enhancing the convenience and intelligence of target identification.
[0086] Furthermore, according to method 100 of this disclosure, target position detection is performed in an image sensor, which is a low-power device. The processor remains in a sleep state during the position detection process performed by the image sensor and is only awakened after the image sensor successfully detects the target position. Therefore, the electronic device according to this disclosure can assist the recognition device in target recognition with lower power consumption. Compared to the prior art, this disclosure significantly reduces the power consumption of the electronic device.
[0087] Figure 5 shows examples of the original image, the large-area proportion image, and the merged image. Note that in Figure 5, the large-area proportion image and the merged image are appropriately enlarged to more clearly show the QR code represented by the thick outline.
[0088] As shown in Figure 5, the area occupied by the QR code in the large-area image is significantly larger than its area occupied in the original image. For example, in the recognition image shown in Figure 5, the QR code occupies approximately 7% of the area in the region of interest.
[0089] Furthermore, the size of the recognition image is much smaller than the size of the original image, which significantly increases the proportion of the QR code in the recognition image. For example, the QR code can be located in the center of the recognition image. As described above, the processor 11 can acquire the recognition image in a similar manner to acquiring images with a large area proportion.
[0090] Figure 6 is a flowchart illustrating another method 200 for identifying a target using an electronic device-assisted identification device according to an embodiment of the present disclosure.
[0091] Specifically, firstly, in step S210, the smart glasses capture an original image of a QR code located in front of the user wearing the smart glasses via the photoelectric conversion unit 21 in its image sensor 12. Note that the original image at this time corresponds to the "first image" according to this disclosure.
[0092] It should be noted that the image sensor 12 of the smart glasses is currently operating in positioning mode. To switch the image sensor 12 to positioning mode, the user wearing the smart glasses can trigger the smart glasses via a button, voice, or gesture. In response to this trigger, the processor 11 of the smart glasses is awakened and thus switches from sleep mode to active mode. Next, the processor 11 sends a control command to the image sensor 12 to switch it to positioning mode. After receiving the control command, the image sensor 12 switches to positioning mode. At the same time, the processor 11 returns to sleep mode.
[0093] Next, in step S211, the image sensor 12 processes the original image with multiple different area ratio enhancement factors by the digital processing module 22, thereby generating multiple large area ratio images. For example, the image sensor 12 can process the original image with area ratio enhancement factors of 1, 4, 8, and 16 respectively, thereby generating four large area ratio images. Note that the large area ratio images at this time correspond to the "fifth image" according to this disclosure.
[0094] Next, the image sensor 12 stitches together the aforementioned multiple large-area-ratio images into a single stitched image via the digital processing module 22. Note that this stitched image corresponds to the "sixth image" according to this disclosure.
[0095] Next, in step S212, the image sensor 12 detects the target's position information in the stitched image through the digital processing module 22. The detection process at this point is similar to that in step S112 of method 100 described above, but the only difference is that the image being detected is a stitched image, not a large-area image. Therefore, a detailed description of the detection process at this point is omitted.
[0096] Next, in step S213, it is determined whether the image sensor 12 has detected the target in the stitched image and thus obtained the target's location information. If the determination result is positive, method 200 proceeds to step S214.
[0097] Similar to step S114 of method 100 described above, in step S214, the processor 11 is woken up in response to the acquisition of location information. At this time, the processor 11 sends a control command to switch the image sensor 12 to streaming mode. The image sensor 12 of the smart glasses recaptures the original image of the QR code via the photoelectric conversion unit 21, and transmits the captured original image and the aforementioned location information to the processor 11. Note that the original image at this time corresponds to the "second image" according to this disclosure.
[0098] Since the user wearing smart glasses is constantly looking at the QR code during target recognition, the position of the QR code in the original image captured at this time basically corresponds to the position in the original image captured in step S210.
[0099] Next, similar to step S115 of method 100 described above, in step S215, in response to the processor 11 processing the original image captured in step S214 based on the position information received from the image sensor 12, an identification image is obtained for identification by the identification device. Note that the identification image at this time corresponds to the "third image" according to this disclosure.
[0100] Next, after acquiring the recognition image, similar to step S116 of method 100 above, in step S216, the processor 11 of the smart glasses transmits the recognition image to the smartphone via a wired interface (such as MIPI, CSI or USB) or a wireless interface (such as WIFI, Bluetooth, Zigbee or cellular communication technology, etc.) for the smartphone to scan or recognize.
[0101] After receiving the image from the smart glasses, the smartphone can open the relevant application to scan the QR code in the image or identify the information contained in the QR code.
[0102] On the other hand, if the determination result in step S213 is negative, then method 200 proceeds to step S217.
[0103] Similar to step S119 of method 100 described above, in step S217, if the image sensor 12 ultimately fails to detect the QR code, the image sensor 12 can return to a sleep state. At this time, the processor 11 continues to remain in a sleep state. Therefore, QR code location detection can also reduce accidental activation by voice, buttons, or gestures.
[0104] Note that when using a DNN algorithm based on a DNN model to detect the location information of a target, the input image (e.g., the stitched image mentioned in step S211) input into the DNN model needs to match the DNN model.
[0105] Typically, to enable the input image to match the DNN model, before detecting the location information of the target in the input image, the image sensor 12 can perform the aforementioned merging process on the input image via the digital processing module 22 to make the input image match the DNN model. The input image at this time can also be referred to as the "merged image".
[0106] After performing the merging process, the image sensor 12 can detect the location information of the QR code in the input image through the digital processing module 22.
[0107] Therefore, according to the method 200 of this disclosure, by using electronic devices to assist in the identification of targets, not only can the distance limitation of the above-mentioned target identification be reduced, but also the tedious operation of manual positioning can be avoided and the user experience can be improved, thereby improving the convenience and intelligence of target identification.
[0108] Furthermore, according to method 200 of this disclosure, target position detection is performed in an image sensor, which is a low-power device. The processor remains in a sleep state during the position detection process performed by the image sensor and is only awakened after the image sensor successfully detects the target position. Therefore, the electronic device according to this disclosure can assist the recognition device in target recognition with lower power consumption. Compared to the prior art, this disclosure significantly reduces the power consumption of the electronic device.
[0109] Figure 7 is a schematic diagram illustrating an example of a scenario in which smart glasses assist a smartphone in scanning a QR code in front of a user, according to the present disclosure.
[0110] Similar to related technologies, in order to use smart glasses to assist smartphones in scanning QR codes in front of users, users can generate trigger signals through buttons, voice, or gestures to wake up the smart glasses.
[0111] Specifically, in response to trigger signals generated by button clicks, user voice, or user gestures, the processor 11 of the smart glasses, such as a SoC (System-on-a-Chip), MCU, or AP, is awakened and sends control commands to the image sensor 12 of the smart glasses via I2C (Integrated Circuit Bus) to cause the image sensor 12 to enter positioning mode. At the same time, the processor 11 enters sleep mode again.
[0112] In location mode, image sensor 12 captures the original image of the QR code in front and detects the location information of the QR code in the original image based on the original image.
[0113] After acquiring the location information of the QR code, the image sensor 12 can transmit the acquired location information to the processor 11.
[0114] At this time, in response to the receipt of location information, the processor 11 is woken up and sends a control command to switch the smart glasses from positioning mode to streaming mode.
[0115] In streaming mode, the processor 11 processes the raw image captured by the image sensor 12 based on location information and obtains a recognition image. As described above, the recognition image may correspond to a region of interest.
[0116] The processor 11 then sends the acquired recognition image to a smartphone, which serves as an example of a recognition device, for the smartphone to scan or recognize.
[0117] Finally, after receiving the image for recognition, the smartphone can open the relevant application to scan the QR code in the image or to recognize the information contained in the QR code.
[0118] The foregoing has described various exemplary devices and methods according to embodiments of this disclosure. It should be understood that the operation or function of these devices can be combined with each other to achieve more or fewer operations or functions than described. Similarly, the operational steps of the methods can be combined with each other in any suitable order to similarly achieve more or fewer operations than described.
[0119] It should be understood that the machine-executable instructions in a machine-readable storage medium or program product according to embodiments of this disclosure can be configured to perform operations corresponding to the above-described device and method embodiments. Embodiments of the machine-readable storage medium or program product will be clear to those skilled in the art when referring to the above-described device and method embodiments, and therefore will not be described again. Machine-readable storage media and program products used to carry or include the above-described machine-executable instructions also fall within the scope of this disclosure. Such storage media may include, but are not limited to, floppy disks, optical disks, magneto-optical disks, memory cards, memory sticks, etc.
[0120] Furthermore, it should be understood that the aforementioned series of processes and devices can also be implemented via software and / or firmware. In the case of implementation via software and / or firmware, programs constituting the software are installed from a storage medium or network onto a computer with a dedicated hardware architecture, such as the example electronic device 1300 shown in FIG8. When various programs are installed, the device is capable of performing various functions, etc. FIG8 is a block diagram illustrating an example of a schematic configuration of an electronic device according to an embodiment of the present disclosure.
[0121] In Figure 8, the Central Processing Unit (CPU) 1301 performs various processes according to the program stored in the Read-Only Memory (ROM) 1302 or the program loaded from the storage section 1308 into the Random Access Memory (RAM) 1303. The RAM 1303 also stores data as needed when the CPU 1301 performs various processes.
[0122] CPU 1301, ROM 1302 and RAM 1303 are connected to each other via bus 1304. Input / output interface 1305 is also connected to bus 1304.
[0123] The following components can be connected to the input / output interface 1305: input section 1306, including a keypad, etc.; output section 1307, including a display and speakers, etc.; storage section 1308, including a memory card, etc.; and communication section 1309, including a network interface card (such as a LAN card), modem, etc. The communication section 1309 can perform communication processing via a network (such as the Internet). A sensor (not shown) can also be connected to the input / output interface 1305.
[0124] As needed, drive 1310 is also connected to input / output interface 1305. Removable media 1311, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on drive 1310 as needed, so that computer programs read from them can be installed into storage section 1308 as needed.
[0125] When the above series of processes are implemented by software, the program constituting the software is installed from a network such as the Internet or a storage medium such as removable media 1311.
[0126] Those skilled in the art will understand that such storage media are not limited to the removable medium 1311 shown, which stores a program and is distributed separately from the device to provide the program to the user. Examples of removable media 1311 include magnetic disks (including floppy disks (registered trademark)), optical disks (including optical disc read-only memory (CD-ROM) and digital versatile disks (DVD)), magneto-optical disks (including mini-disk (MD) (registered trademark)), and semiconductor memories. Alternatively, the storage medium may be ROM 1302, a hard disk included in storage section 1308, etc., containing a program and distributed to the user along with the device containing them.
[0127] Figure 9 is a block diagram illustrating another example of a schematic configuration of an electronic device 1600 to which the technology of this disclosure can be applied. The electronic device 1600 includes a processor 1601, a memory 1602, a storage device 1603, an external connection interface 1604, a camera device 1606, a sensor 1607, a microphone 1608, an input device 1609, a display device 1610, a speaker 1611, a wireless communication interface 1612, one or more antenna switches 1615, one or more antennas 1616, a bus 1617, a battery 1618, and an auxiliary controller 1619. In one implementation, the electronic device 1600 herein may correspond to a device with a camera, such as smart glasses.
[0128] Processor 1601 may be, for example, a CPU or a System-on-a-Chip (SoC), and controls the application layer and other functions of electronic device 1600. Memory 1602 includes RAM and ROM, and stores data and programs executed by processor 1601. Storage device 1603 may include storage media such as semiconductor memory and hard disk. External connection interface 1604 is an interface for connecting external devices (such as memory cards and Universal Serial Bus (USB) devices) to electronic device 1600.
[0129] The camera device 1606 includes an image sensor (such as a charge-coupled device (CCD) and complementary metal-oxide-semiconductor (CMOS)) and generates captured images. The sensor 1607 may include a set of sensors, such as a measurement sensor, a gyroscope sensor, a magnetometer sensor, and an accelerometer sensor. The microphone 1608 converts sound input to the electronic device 1600 into an audio signal. The input device 1609 includes, for example, a touch sensor, keypad, keyboard, buttons, or switches configured to detect touches on the screen of the display device 1610 and receives operations or information input from the user. The display device 1610 includes a screen (such as a liquid crystal display (LCD) and an organic light-emitting diode (OLED) display) and displays the output image of the electronic device 1600. The speaker 1611 converts the audio signal output from the electronic device 1600 into sound.
[0130] Wireless communication interface 1612 supports any cellular communication scheme (such as LTE, LTE-Advanced, and NR) and performs wireless communication. Wireless communication interface 1612 typically includes, for example, a BB processor 1613 and RF circuitry 1614. BB processor 1613 can perform, for example, encoding / decoding, modulation / demodulation, and multiplexing / demultiplexing, and performs various types of signal processing for wireless communication. Meanwhile, RF circuitry 1614 can include, for example, mixers, filters, and amplifiers, and transmits and receives wireless signals via antenna 1616. Wireless communication interface 1612 can be a single chip module on which BB processor 1613 and RF circuitry 1614 are integrated. As shown in Figure 9, wireless communication interface 1612 can include multiple BB processors 1613 and multiple RF circuits 1614. Although Figure 9 shows an example where wireless communication interface 1612 includes multiple BB processors 1613 and multiple RF circuits 1614, wireless communication interface 1612 can also include a single BB processor 1613 or a single RF circuitry 1614.
[0131] In addition to cellular communication schemes, wireless communication interface 1612 can support other types of wireless communication schemes, such as short-range wireless communication schemes, near-field communication schemes, and wireless local area network (LAN) schemes. In this case, wireless communication interface 1612 may include a BB processor 1613 and RF circuitry 1614 for each wireless communication scheme.
[0132] Each of the antenna switches 1615 switches the connection destination of the antenna 1616 among multiple circuits (e.g., circuits for different wireless communication schemes) included in the wireless communication interface 1612.
[0133] Each of the antennas 1616 includes one or more antenna elements (such as multiple antenna elements included in a MIMO antenna) and is used for transmitting and receiving wireless signals through the wireless communication interface 1612. As shown in FIG9, the electronic device 1600 may include multiple antennas 1616. Although FIG9 shows an example in which the electronic device 1600 includes multiple antennas 1616, the electronic device 1600 may also include a single antenna 1616.
[0134] Furthermore, the electronic device 1600 may include an antenna 1616 for each wireless communication scheme. In this case, the antenna switch 1615 may be omitted from the configuration of the electronic device 1600.
[0135] Bus 1617 connects processor 1601, memory 1602, storage device 1603, external connection interface 1604, camera device 1606, sensor 1607, microphone 1608, input device 1609, display device 1610, speaker 1611, wireless communication interface 1612, and auxiliary controller 1619 to each other. Battery 1618 supplies power to the various blocks of electronic device 1600 shown in FIG9 via feeders, which are partially shown as dashed lines in the figure. Auxiliary controller 1619 operates the minimum necessary functions of electronic device 1600, for example, in sleep mode.
[0136] Figure 10 is a block diagram illustrating yet another example of a schematic configuration of an electronic device 1720 according to an embodiment of the present disclosure. The electronic device 1720 includes one or more of the following: a processor 1721, a memory 1722, a Global Positioning System (GPS) module 1724, a sensor 1725, a data interface 1726, a content player 1727, a storage medium interface 1728, an input device 1729, a display device 1730, a speaker 1731, a wireless communication interface 1733, one or more antenna switches 1736, one or more antennas 1737, and a battery 1738. In one implementation, the electronic device 1720 herein may correspond to a device with a camera, such as smart glasses.
[0137] The processor 1721 can be, for example, a CPU or a SoC, and controls various functions of the electronic device 1720. The memory 1722 includes RAM and ROM, and stores data and programs executed by the processor 1721.
[0138] GPS module 1724 uses GPS signals received from GPS satellites to measure the location (such as latitude, longitude, and altitude) of electronic device 1720. Sensor 1725 may include a set of sensors, such as a gyroscope sensor, a geomagnetic sensor, and an air pressure sensor, to collect relevant information. Sensor 1725 may also include an image sensor, for example, for waking electronic device 1720 from sleep mode based on captured image data. Data interface 1726 is connected to, for example, a data network 1741 via a terminal not shown.
[0139] Content player 1727 reproduces content stored on a storage medium (such as a memory card), which is inserted into storage medium interface 1728. Input device 1729 includes, for example, a touch sensor, button, or switch configured to detect touch on the screen of display device 1730, and receives operations or information input from the user. Display device 1730 includes a screen such as an LCD or OLED display and displays images or reproduced content for various functions. Speaker 1731 outputs sound or reproduced content for various functions.
[0140] The wireless communication interface 1733 supports any cellular communication scheme (such as LTE, LTE-Advanced, and NR) and performs wireless communication. The wireless communication interface 1733 typically includes, for example, a BB processor 1734 and RF circuitry 1735. The BB processor 1734 can perform, for example, encoding / decoding, modulation / demodulation, and multiplexing / demultiplexing, and performs various types of signal processing for wireless communication. Meanwhile, the RF circuitry 1735 can include, for example, mixers, filters, and amplifiers, and transmits and receives wireless signals via antenna 1737. The wireless communication interface 1733 can also be a chip module on which the BB processor 1734 and RF circuitry 1735 are integrated. As shown in Figure 10, the wireless communication interface 1733 can include multiple BB processors 1734 and multiple RF circuits 1735. Although Figure 10 shows an example where the wireless communication interface 1733 includes multiple BB processors 1734 and multiple RF circuits 1735, the wireless communication interface 1733 can also include a single BB processor 1734 or a single RF circuitry 1735.
[0141] In addition to cellular communication schemes, the wireless communication interface 1733 can support other types of wireless communication schemes, such as short-range wireless communication schemes, near-field communication schemes, and wireless LAN schemes. In this case, for each wireless communication scheme, the wireless communication interface 1733 may include a BB processor 1734 and an RF circuit 1735.
[0142] Each of the antenna switches 1736 switches the connection destination of the antenna 1737 among multiple circuits (such as circuits for different wireless communication schemes) included in the wireless communication interface 1733.
[0143] Each of the antennas 1737 includes one or more antenna elements (such as multiple antenna elements included in a MIMO antenna) and is used for transmitting and receiving wireless signals through the wireless communication interface 1733. As shown in Figure 10, the electronic device 1720 may include multiple antennas 1737. Although Figure 10 shows an example in which the electronic device 1720 includes multiple antennas 1737, the electronic device 1720 may also include a single antenna 1737.
[0144] Furthermore, the electronic device 1720 may include an antenna 1737 for each wireless communication scheme. In this case, the antenna switch 1736 can be omitted from the configuration of the electronic device 1720.
[0145] Battery 1738 may be a rechargeable battery and may supply power to various blocks of electronic device 1720 shown in FIG10 via feeders, which are partially shown as dashed lines in the figure.
[0146] Although the present disclosure has been described above with the example of using smart glasses as an electronic device to assist in the identification of targets, those skilled in the art will understand that the techniques according to the present disclosure can also be applied, for example, to perform secondary feature information amplification and other processing such as locating a distant face and then performing facial information recognition. Specifically, a surveillance camera with a smart camera can be used to assist in the identification of facial features of people within the field of view of the surveillance camera.
[0147] As will be appreciated from the description herein, embodiments of this disclosure can be configured as follows:
[0148] (1) A method for assisting in target identification using an electronic device, the electronic device comprising a sensor and a processor, the method comprising:
[0149] The sensor captures a first image of the target and detects the target's location information based on the first image;
[0150] The sensor captures a second image of the target and sends the second image and the location information to the processor, wherein the target's position in the first image corresponds to its position in the second image; and
[0151] The processor processes the second image based on the location information and generates a third image for identifying the target, the third image corresponding to a region of interest including the target and its surroundings.
[0152] (2) According to the method described in (1) above, wherein, when detecting the location information of the target based on the first image,
[0153] The sensor processes the first image to obtain a fourth image of the target, wherein the target's area in the fourth image is greater than its area in the first image, and
[0154] The sensor detects the location information of the target based on the fourth image.
[0155] (3) According to the method described in (2) above, the location information is detected using a DNN algorithm.
[0156] (4) According to the method described in (3) above, the DNN algorithm is implemented based on pruned MobileNet SSD, YOLO series or Fast-RCNN series.
[0157] (5) The method according to (3) or (4) above, wherein, when detecting the location information of the target based on the fourth image,
[0158] The sensor first performs a merging process on the fourth image to make the fourth image match the DNN algorithm, and then detects the location information of the target based on the merged fourth image.
[0159] (6) The method according to any one of (2) to (4) above, wherein the sensor processes the first image by cropping or digital magnification to obtain the fourth image of the target.
[0160] (7) When detecting the location information of the target based on the first image according to any one of (1) to (6) above,
[0161] The sensor processes the first image to obtain multiple fifth images of the target, wherein the area of the target in each of the fifth images is different and is greater than or equal to the area of the target in the first image.
[0162] The sensor performs stitching processing on the plurality of fifth images to obtain a sixth image of the target, and
[0163] The sensor detects the target's location information based on the sixth image.
[0164] (8) According to the method described in (7) above, the location information is detected using a DNN algorithm.
[0165] (9) According to the method described in (8) above, wherein the DNN algorithm is implemented based on pruned MobileNet SSD, YOLO series or Fast-RCNN series.
[0166] (10) The method according to (8) or (9) above, wherein, when detecting the location information of the target based on the sixth image,
[0167] The sensor first performs a merging process on the sixth image to match the DNN algorithm, and then detects the location information of the target based on the merged sixth image.
[0168] (11) The method according to any one of (7) to (10) above, wherein the sensor processes the first image by cropping or digital magnification with different cropping coefficients or digital magnification coefficients to obtain the plurality of fifth images.
[0169] (12) The method according to any one of (1) to (11) above, wherein the processor processes the second image by cropping or digital magnification to obtain the third image.
[0170] (13) The method according to any one of (1) to (12) above further includes:
[0171] Before the sensor captures a first image of the target and detects the target's location information based on the first image: the processor is awakened from sleep mode by an external trigger signal; the processor sends control commands to the sensor to prepare it for detecting the target's location information, while the processor returns to sleep mode; and
[0172] When the sensor sends the second image and the location information to the processor: wake the processor from the sleep state.
[0173] (14) The method according to any one of (1) to (13) above, wherein,
[0174] The electronic device is smart glasses, an AR device with a camera, a smart camera, or a device with a smart camera.
[0175] (15) An electronic device comprising a sensor and a processor, said sensor including a photoelectric conversion unit and a digital processing module.
[0176] The sensor is configured to capture a first image of the target via the photoelectric conversion unit, and the digital processing module detects the target's location information based on the first image received from the photoelectric conversion unit.
[0177] The sensor is configured to capture a second image of the target via the photoelectric conversion unit, and send the second image and the position information to the processor. The position of the target in the first image corresponds to its position in the second image.
[0178] The processor is further configured to process the second image based on the location information and generate a third image for identifying the target, the third image corresponding to a region of interest including the target and its surroundings.
[0179] (16) The electronic device according to (15) above, wherein the digital processing module is configured to, when detecting the location information of the target based on the first image,
[0180] The first image is processed to obtain a fourth image of the target, wherein the area of the target in the fourth image is greater than its area in the first image, and
[0181] The location information of the target is detected based on the fourth image.
[0182] (17) The electronic device according to (16) above, wherein the digital processing module includes a DNN module.
[0183] (18) The electronic device according to (17) above, wherein the DNN module is implemented based on pruned MobileNet SSD, YOLO series or Fast-RCNN series.
[0184] (19) The electronic device according to (17) or (18) above, wherein the digital processing module is configured to, when detecting the location information of the target based on the fourth image,
[0185] First, a merging process is performed on the fourth image to make the fourth image match the DNN algorithm of the DNN module, and the location information of the target is detected based on the fourth image after the merging process.
[0186] (20) The electronic device according to any one of (16) to (18) above, wherein the digital processing module is configured to process the first image by cropping or digital magnification to obtain the fourth image of the target.
[0187] (21) In any one of (15) to (20) above, the digital processing module is configured to, when detecting the location information of the target based on the first image,
[0188] The first image is processed to obtain multiple fifth images of the target, wherein the area of the target in each of the fifth images is different and is greater than or equal to the area of the target in the first image.
[0189] The plurality of fifth images are stitched together to obtain a sixth image of the target.
[0190] The location information of the target is detected based on the sixth image.
[0191] (22) The electronic device according to (21) above, wherein the digital processing module includes a DNN module.
[0192] (23) The electronic device according to (22) above, wherein the DNN module is implemented based on pruned MobileNet SSD, YOLO series or Fast-RCNN series.
[0193] (24) The electronic device according to (22) or (23) above, wherein the digital processing module is configured to, when detecting the location information of the target based on the sixth image,
[0194] First, a merging process is performed on the sixth image to make the sixth image match the DNN algorithm of the DNN module, and the location information of the target is detected based on the merged sixth image.
[0195] (25) The electronic device according to any one of (21) to (24) above, wherein the digital processing module is configured to process the first image by cropping or digital magnification with different cropping factors or digital magnification factors to obtain the plurality of fifth images.
[0196] (26) The electronic device according to any one of (15) to (25) above, wherein the processor is configured to process the second image by cropping or digital magnification to obtain the third image.
[0197] (27) The electronic device according to any one of (15) to (26) above, wherein,
[0198] Before the sensor captures a first image of the target and detects the target's location information based on the first image: the processor wakes up from sleep mode via an external trigger signal; the processor sends control commands to the sensor to prepare it for detecting the target's location information, while the processor returns to sleep mode; and
[0199] When the sensor sends the second image and the location information to the processor: the processor is awakened from the sleep state.
[0200] (28) The electronic device according to any one of (15) to (27) above, wherein,
[0201] The electronic device is smart glasses, an AR device with a camera, a smart camera, or a device with a smart camera.
[0202] (29) A sensor comprising:
[0203] A photoelectric conversion unit, configured to convert received optical signals into image signals; and
[0204] A digital processing module is configured to process the image signal received from the photoelectric conversion unit to obtain target location information, the location information being configured to be used by an electronic device including the sensor to generate an identification image for identifying the target.
[0205] (30) A program product comprising computer instructions configured to perform the following operations when executed by a processor:
[0206] Image signals from an electronic device are processed to obtain location information of a target, which is configured to be used by the electronic device to generate an identification image for identifying the target.
[0207] (31) An apparatus for assisting in the identification of a target, comprising:
[0208] Processor; and
[0209] A memory storing computer instructions configured to perform the following operations when executed by the processor:
[0210] Image signals from an electronic device are processed to obtain location information of the target, the location information being configured to be used by the electronic device to generate an identification image for identifying the target.
[0211] (32) A computer-readable storage medium having stored thereon computer program instructions configured to perform the following operations when executed by a processor:
[0212] Image signals from an electronic device are processed to obtain location information of the target, the location information being configured to be used by the electronic device to generate an identification image for identifying the target.
[0213] Exemplary embodiments of the present disclosure have been described above with reference to the accompanying drawings; however, the present disclosure is by no means limited to the examples described above. Various changes and modifications can be made by those skilled in the art within the scope of the appended claims, and it should be understood that such changes and modifications naturally fall within the technical scope of the present disclosure.
[0214] For example, the multiple functions included in one unit in the above embodiments can be implemented by separate devices. Alternatively, the multiple functions implemented by multiple units in the above embodiments can be implemented by separate devices respectively. In addition, one of the above functions can be implemented by multiple units. Needless to say, such a configuration is included within the scope of the present disclosure.
[0215] In this specification, the steps described in the flowchart include not only processes executed sequentially in the stated order, but also processes executed in parallel or individually, rather than necessarily sequentially. Furthermore, even within the steps of sequential processing, needless to say, the order can be appropriately altered.
[0216] While this disclosure and its advantages have been described in detail, it should be understood that various changes, substitutions, and modifications can be made without departing from the spirit and scope of this disclosure as defined by the appended claims. Furthermore, the terms "comprising," "including," or any other variations thereof used in embodiments of this disclosure are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
Claims
1. A method for assisting in target identification using an electronic device, the electronic device comprising a sensor and a processor, the method comprising: The sensor captures a first image of the target and detects the target's location information based on the first image; The sensor captures a second image of the target and sends the second image and the location information to the processor, wherein the position of the target in the first image corresponds to its position in the second image; as well as The processor processes the second image based on the location information and generates a third image for identifying the target, the third image corresponding to a region of interest including the target and its surroundings.
2. The method according to claim 1, wherein, When detecting the location information of the target based on the first image, The sensor processes the first image to obtain a fourth image of the target, wherein the target's area in the fourth image is greater than its area in the first image, and The sensor detects the location information of the target based on the fourth image.
3. The method according to claim 2, wherein, The location information is detected using a DNN algorithm.
4. The method according to claim 3, wherein, The DNN algorithm is implemented based on pruned MobileNet SSD, YOLO series, or Fast-RCNN series.
5. The method according to claim 3, wherein, When detecting the location information of the target based on the fourth image, The sensor first performs a merging process on the fourth image to make the fourth image match the DNN algorithm, and then detects the location information of the target based on the merged fourth image.
6. The method according to claim 2, wherein, The sensor processes the first image by cropping or digitally magnifying it to obtain the fourth image of the target.
7. The method according to claim 1, wherein when detecting the location information of the target based on the first image, The sensor processes the first image to obtain multiple fifth images of the target, wherein the area of the target in each of the fifth images is different and is greater than or equal to the area of the target in the first image. The sensor performs stitching processing on the plurality of fifth images to obtain a sixth image of the target, and The sensor detects the target's location information based on the sixth image.
8. The method according to claim 7, wherein, The location information is detected using a DNN algorithm.
9. The method according to claim 8, wherein, The DNN algorithm is implemented based on pruned MobileNet SSD, YOLO series, or Fast-RCNN series.
10. The method according to claim 8, wherein, When detecting the location information of the target based on the sixth image, The sensor first performs a merging process on the sixth image to match the DNN algorithm, and then detects the location information of the target based on the merged sixth image.
11. The method according to claim 7, wherein, The sensor processes the first image by cropping or digitally magnifying it with different cropping or digital magnification factors to obtain the plurality of fifth images.
12. The method according to claim 1, wherein, The processor processes the second image by cropping or digitally enlarging it to obtain the third image.
13. The method according to claim 1, further comprising: Before the sensor captures a first image of the target and detects the target's location information based on the first image: the processor is woken up from sleep state by a trigger signal from an external source; The processor sends control commands to the sensor to prepare it to detect the location information of the target, while the processor returns to the sleep state. as well as When the sensor sends the second image and the location information to the processor: wake the processor from the sleep state.
14. The method according to claim 1, wherein, The electronic device is smart glasses, an AR device with a camera, a smart camera, or a device with a smart camera.
15. An electronic device comprising a sensor and a processor, said sensor including a photoelectric conversion unit and a digital processing module. in, The sensor is configured to capture a first image of the target via the photoelectric conversion unit, and the digital processing module detects the target's position information based on the first image received from the photoelectric conversion unit. The sensor is configured to capture a second image of the target via the photoelectric conversion unit, and send the second image and the position information to the processor. The position of the target in the first image corresponds to its position in the second image. The processor is further configured to process the second image based on the location information and generate a third image for identifying the target, the third image corresponding to a region of interest including the target and its surroundings.
16. The electronic device according to claim 15, wherein, The digital processing module is configured to, when detecting the location information of the target based on the first image, The first image is processed to obtain a fourth image of the target, wherein the area of the target in the fourth image is greater than its area in the first image, and The location information of the target is detected based on the fourth image.
17. The electronic device according to claim 16, wherein, The digital processing module includes a DNN module.
18. The electronic device according to claim 17, wherein, The DNN module is implemented based on pruned MobileNet SSD, YOLO series, or Fast-RCNN series.
19. The electronic device according to claim 17, wherein, The digital processing module is configured to, when detecting the location information of the target based on the fourth image, First, a merging process is performed on the fourth image to make the fourth image match the DNN algorithm of the DNN module, and the location information of the target is detected based on the fourth image after the merging process.
20. The electronic device according to claim 16, wherein, The digital processing module is configured to process the first image by cropping or digital magnification to obtain the fourth image of the target.
21. The electronic device of claim 15, wherein the digital processing module is configured to, when detecting the location information of the target based on the first image, The first image is processed to obtain multiple fifth images of the target, wherein the area of the target in each of the fifth images is different and is greater than or equal to the area of the target in the first image. The plurality of fifth images are stitched together to obtain a sixth image of the target. The location information of the target is detected based on the sixth image.
22. The electronic device according to claim 21, wherein, The digital processing module includes a DNN module.
23. The electronic device according to claim 22, wherein, The DNN module is implemented based on pruned MobileNet SSD, YOLO series, or Fast-RCNN series.
24. The electronic device according to claim 22, wherein, The digital processing module is configured to, when detecting the location information of the target based on the sixth image, First, a merging process is performed on the sixth image to make the sixth image match the DNN algorithm of the DNN module, and the location information of the target is detected based on the merged sixth image.
25. The electronic device according to claim 21, wherein, The digital processing module is configured to process the first image by cropping or digitally magnifying it with different cropping or magnification factors to obtain the plurality of fifth images.
26. The electronic device according to claim 15, wherein, The processor is configured to process the second image by cropping or digitally enlarging to obtain the third image.
27. The electronic device according to claim 15, wherein, Before the sensor captures a first image of the target and detects the target's location information based on the first image: the processor wakes up from sleep mode by a trigger signal from an external source; The processor sends control commands to the sensor to prepare it to detect the location information of the target, while the processor returns to the sleep state. as well as When the sensor sends the second image and the location information to the processor: the processor is awakened from the sleep state.
28. The electronic device according to claim 15, wherein, The electronic device is smart glasses, an AR device with a camera, a smart camera, or a device with a smart camera.
29. A sensor comprising: A photoelectric conversion unit is configured to convert received optical signals into image signals; and A digital processing module is configured to process the image signal received from the photoelectric conversion unit to obtain target location information, the location information being configured to be used by an electronic device including the sensor to generate an identification image for identifying the target.
30. A program product comprising computer instructions configured to perform the following actions when executed by a processor: Image signals from an electronic device are processed to obtain location information of a target, which is configured to be used by the electronic device to generate an identification image for identifying the target.
31. An apparatus for assisting in target identification, comprising: processor; as well as A memory storing computer instructions configured to perform the following operations when executed by the processor: Image signals from an electronic device are processed to obtain location information of the target, the location information being configured to be used by the electronic device to generate an identification image for identifying the target.
32. A computer-readable storage medium having stored thereon computer program instructions configured to perform the following operations when executed by a processor: Image signals from an electronic device are processed to obtain location information of the target, the location information being configured to be used by the electronic device to generate an identification image for identifying the target.