Image recognition method, intelligent glasses and readable storage medium
By combining the image acquisition module and the eye movement information acquisition module in the smart glasses, the target area is determined using eye movement gaze information and high resolution acquisition, the problem of high power consumption of smart glasses is solved, and low power consumption and high accuracy image recognition is achieved.
Patent Information
- Application Number
- CN202510453919.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-22
AI Technical Summary
Existing smart glasses consume high power when recognizing images, resulting in serious heat generation and shorter battery life.
By setting up an image acquisition module and an eye movement information acquisition module in the smart glasses, the low-resolution image is first acquired, the target area is determined using eye movement gaze information, and then the image of the area is collected at high resolution for identification, avoiding the acquisition of full-size high-resolution images.
It reduces the power consumption of smart glasses for image recognition, improves recognition accuracy and extends battery life.
Smart Images

Figure CN120355898A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of smart wearable devices, and more specifically, to an image recognition method, smart glasses, and readable storage medium. Background Art
[0002] Currently, with the development of electronic information technology, images can be collected through smart glasses, and then image recognition can be performed on the collected images. However, the power consumption required for current image recognition is relatively high. Summary of the Invention
[0003] This application proposes an image recognition method, smart glasses, and readable storage medium to improve the above-mentioned defects.
[0004] In a first aspect, an embodiment of this application provides an image recognition method, which is applied to a processing module of smart glasses. The smart glasses further include an image acquisition module and an eye movement information acquisition module. The processing module is respectively connected to the image acquisition module and the eye movement information acquisition module. The method includes: obtaining a first image with a first resolution based on the image acquisition module; determining a target area in the first image based on the eye movement fixation information collected by the eye movement information acquisition module; determining the image acquisition area of the image acquisition module through the target area; performing image acquisition on the image acquisition area by the image acquisition module at a second resolution to obtain a second image, where the second resolution is greater than the first resolution; and performing image recognition on the second image through an image recognition model.
[0005] In a second aspect, an embodiment of this application further provides a pair of smart glasses, including: a processing module, an image acquisition module, and an eye movement information acquisition module. The processing module is respectively connected to the image acquisition module and the eye movement information acquisition module; and the processing module is configured to perform image recognition based on the method described in the first aspect.
[0006] In a third aspect, an embodiment of this application further provides a computer-readable storage medium, in which program code is stored, and the program code can be called by a processor to execute the method described in the first aspect.
[0007] The image recognition method, smart glasses, and readable storage medium provided by this application obtain a first image with a first resolution based on the image acquisition module; determine a target area in the first image based on the eye movement fixation information collected by the eye movement information acquisition module; determine the image acquisition area of the image acquisition module through the target area; perform image acquisition on the image acquisition area by the image acquisition module at a second resolution to obtain a second image, where the second resolution is greater than the first resolution; perform image recognition on the second image through an image recognition model. If high-resolution images are collected and then combined with the annotation of eye movement fixation information for image recognition, the power consumption will be relatively high. In the solution of this application, first, a target area corresponding to the eye movement fixation information is determined through a first image with a lower resolution, and then the image acquisition area is determined by this target area, so as to only collect the second image corresponding to the high-resolution target acquisition area, and then perform image recognition on the second image. That is to say, the second image with a higher resolution is not a full-size image. Therefore, the entire processing process does not involve full-size high-resolution images, resulting in lower power consumption required for image recognition through smart glasses.
[0008] Other features and advantages of the embodiments of this application will be described in the subsequent specification, and part of them will become obvious from the specification, or be understood by implementing the embodiments of this application. The objectives and other advantages of the embodiments of this application can be achieved and obtained through the structures specifically pointed out in the written specification, claims, and drawings. Brief Description of the Drawings
[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0010] Figure 1 Shows the structural block diagram of the smart glasses provided by the embodiments of this application;
[0011] Figure 2 Shows the method flow chart of the image recognition method provided by the embodiments of this application;
[0012] Figure 3 Shows the application scenario diagram of the image recognition method provided by the embodiments of this application;
[0013] Figure 4 Shows the method flow chart of the image recognition method provided by another embodiment of this application;
[0014] Figure 5The figure shows a schematic diagram of determining a target area provided by an embodiment of the present application;
[0015] Figure 6 The figure shows a schematic diagram of determining a target area provided by another embodiment of the present application;
[0016] Figure 7 The figure shows a schematic diagram of determining a target area provided by yet another embodiment of the present application;
[0017] Figure 8 The figure shows a flowchart of a method for an image recognition method provided by yet another embodiment of the present application;
[0018] Figure 9 The figure shows a structural block diagram of a computer-readable storage medium provided by an embodiment of the present application. Detailed implementation manners
[0019] In order to enable those skilled in the art of the present technology to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Usually, the components of the embodiments of the present application described and shown in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application to be protected, but only represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0020] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, the terms "first", "second", etc. are only used for differential description and cannot be understood as indicating or implying relative importance.
[0021] Currently, with the development of electronic information technology, images can be collected through smart glasses, and then image recognition is performed on the collected images. However, the current power consumption required for image recognition is relatively high. How to reduce the power consumption required for image recognition through smart glasses is an urgent problem to be solved.
[0022] Currently, an image acquisition module for collecting images can be set on the smart glasses, and then the images can be collected through this image acquisition module, and then image recognition is performed on the images.
[0023] However, the inventors found in their research that if image recognition is directly performed on the images collected by the image acquisition module, the required power consumption is relatively high, which may also cause problems such as serious overheating and shortened battery life of the smart glasses.
[0024] Therefore, in order to overcome the above defects, the present application provides an image recognition method, a smart glasses and a readable storage medium.
[0025] Please refer to Figure 1 , Figure 1 which shows the structural block diagram of the smart glasses provided by the embodiment of the present application, specifically showing the smart glasses 100.
[0026] In some embodiments, the smart glasses 100 may be an augmented reality glasses device. In the smart glasses 100 shown in Figure 1 , a processing module 110, an eye movement information acquisition module 120 and an image acquisition module 130 are provided. Among them, the processing module 110 is respectively connected to the eye movement information acquisition module 120 and the image acquisition module 130.
[0027] The eye movement information acquisition module 120 can be used to acquire eye movement fixation information. The eye movement fixation information is the eye movement fixation information of the user wearing the smart glasses. Among them, the eye movement fixation information represents the fixation behavior of the user when observing a scene, an image, text or a dynamic picture through the visual attention distribution data recorded by tracking the eye movement.
[0028] The image acquisition module 130 may include at least one camera. The image acquisition module 130 can be used to acquire images. In the embodiment provided by the present application, the image acquisition module 130 can be used to acquire images with different resolutions, for example, images with a first resolution and images with a second resolution can be acquired respectively. Thus, the smart glasses 100 can combine the acquired images and eye movement fixation information to perform image recognition. For a detailed introduction, please refer to the subsequent embodiments.
[0029] Please refer to Figure 2 , Figure 2 which shows the method flow chart of the image recognition method provided by the embodiment of the present application. The image recognition method can be applied to the processing module in the smart glasses shown in Figure 1 . Specifically, the method may include step S110 to step S150.
[0030] As can be seen from the foregoing introduction, the smart glasses are provided with a processing module, an image acquisition module and an eye movement information acquisition module. The processing module is respectively connected to the image acquisition module and the eye movement information acquisition module.
[0031] Step S110: Obtain a first image with a first resolution based on the image acquisition module.
[0032] It can be understood that, in order to subsequently determine the image area that the user is gazing at or observing in the image in combination with the eye movement gaze information, an image can be first obtained through the image acquisition module. And in order to reduce the computational complexity required to subsequently determine the image area that the user is gazing at or observing in the image in combination with the acquired image and the eye movement gaze information, thereby reducing power consumption, for some embodiments, an image with a lower resolution can be obtained, that is, a first image with a first resolution can be obtained based on the image acquisition module, and then the image area that the user is gazing at or observing can be determined in combination with the first image with a lower resolution.
[0033] Exemplarily, the resolution of the image acquisition module can be set, and then the image acquisition module can directly acquire an image at the first resolution, thereby obtaining a first image with the first resolution.
[0034] Exemplarily, it can also be set to directly acquire an image at the maximum resolution supported by the image acquisition module, thereby obtaining an image with a higher resolution, and then perform an operation to reduce the resolution of the image with a higher resolution to obtain a first image with the first resolution.
[0035] It should be noted that the first resolution is the lower resolution, that is, the first resolution is less than the maximum resolution of the image that the image acquisition module can support for acquisition.
[0036] Step S120: Determine a target area in the first image based on the eye movement gaze information collected by the eye movement information acquisition module.
[0037] In addition, the eye movement gaze information can also be collected by the eye movement information acquisition module, and then a target area can be determined in the first image based on the eye movement gaze information collected by the eye movement information acquisition module. Among them, the eye movement gaze information can be used to represent the position coordinates of the user's gaze area. Exemplarily, the position coordinates can be represented by a first position in a first coordinate system, where the first coordinate system can be a coordinate system established based on the eye movement information acquisition module.
[0038] Furthermore, the coordinate system established based on the image acquisition module can be used as a second coordinate system. Since the first image is obtained by the image acquisition module, the positions of each pixel point in the first image can be represented by the respective position points in the second coordinate system. Thus, to determine the target area in the first image based on the eye movement gaze information, first, a second position corresponding to the eye movement gaze information can be determined in the second coordinate system, and then the target area can be determined in the first image based on the second position. For a detailed introduction, reference can be made to the subsequent embodiments.
[0039] Among them, the target area can be used to represent the area where the user gazes in the first image. Therefore, the target area is the area where the user's attention is located. Among them, the user described in each embodiment of this application is the user wearing the smart glasses.
[0040] Exemplarily, please refer to Figure 3 , Figure 3 , which shows an application scenario diagram of the image recognition method provided by the embodiment of this application, that is, the image recognition scenario 300. In Figure 3 The shown image recognition scenario 300 includes a user 320 wearing smart glasses 310. In the smart glasses 310, an image acquisition module ( Figure 3 not shown in the figure) and an eye movement information acquisition module ( Figure 3 not shown in the figure) can be set. Optionally, in order to balance the weights of the two temples of the smart glasses 310, the smart glasses 310 can also be configured with a counterweight module ( Figure 3 not shown in the figure).
[0041] Among them, the process of determining the target area in this step can also be called image segmentation. It can be understood that image segmentation requires relatively high computing power, while the computing power that the hardware configured in the smart glasses can provide is generally not too high. Therefore, in some embodiments, the eye movement gaze information collected by the eye movement information acquisition module can be obtained first; then the eye movement gaze information and the first image are sent to an electronic device connected to the smart glasses to instruct the electronic device to determine the target area in the first image and send the target area to the smart glasses; thus, the smart glasses can receive the target area. Among them, the electronic device generally has higher computing power compared to the smart glasses, and the electronic device is communicatively connected to the smart glasses. Thus, the smart glasses can send the collected eye movement gaze information and the first image to the electronic device. After the electronic device determines the target area in the first image, it sends the target area to the smart glasses. Thus, the smart glasses can obtain the target area with lower power consumption and higher efficiency.
[0042] Among them, the electronic device can be a smartphone, and the smart glasses can be directly connected to the smartphone; in addition, the electronic device can also be a cloud server, so that the smart device can establish a communication connection with the cloud server through the smartphone. Optionally, the smart glasses can also be configured with a cellular data network or a wireless network Wi-Fi, so that the smart glasses can directly communicate with the cloud server through the cellular data network or the wireless network Wi-Fi. The embodiments of this application do not make specific limitations.
[0043] Step S130: Determine the image acquisition area of the image acquisition module through the target area.
[0044] Further, after obtaining the target area, the image acquisition area of the image acquisition module can be determined based on the target area. It should be noted that the image acquisition module generally can control the exposure of the sensor to obtain the acquired image. By setting or adjusting the area of the sensor exposure, only the image of a partial area can be acquired.
[0045] Therefore, in the embodiment provided in this application, the target area has been obtained through the foregoing steps. At this time, the image acquisition module can be set based on the target area to determine the image acquisition area of the image acquisition module. Exemplarily, the area where the target area is located in the sensor of the image acquisition module can be set as the image acquisition area.
[0046] Step S140: Image acquisition is performed on the image acquisition area by the image acquisition module at a second resolution to obtain a second image, where the second resolution is greater than the first resolution.
[0047] Thus, after determining the image acquisition area, the image acquisition area can be acquired by the image acquisition module. In the embodiment provided in this application, the image acquisition area can be acquired by the image acquisition module at a second resolution to obtain a second image. Among them, the second resolution is a higher resolution, so the second resolution is greater than the first resolution. Exemplarily, the second resolution can be the maximum resolution of the image that the image acquisition module can support for acquisition.
[0048] Therefore, the second image obtained through this step is an image with a higher resolution, and the second image contains the area where the user's attention is located. Moreover, the second image is not a full-size second-resolution image acquired by the image acquisition module. Therefore, the power consumption can be reduced in the subsequent process of further processing the second image. Herein, full-size is used to represent the maximum size of the image that the image acquisition module can support for acquisition, that is, the image size corresponding to the entire area in the sensor of the image acquisition module.
[0049] In addition, since there may be a delay between the first moment when the first image is obtained and the second moment when the second image is obtained in the embodiments of the present application, the user wearing the smart glasses may move or shake between the first moment and the second moment, which may cause the attitude information of the image acquisition module to be different at the first moment and the second moment, thereby reducing the accuracy of the determined image acquisition area. Therefore, optionally, in some embodiments, the image acquisition area may also be adjusted based on the first attitude information of the image acquisition module corresponding to the first image and the second attitude information of the image acquisition module corresponding to the second image, and then image acquisition is performed on the adjusted image acquisition area to obtain a second image. For a detailed introduction, reference may be made to the subsequent embodiments.
[0050] Step S150: Perform image recognition on the second image through an image recognition model.
[0051] The second image is obtained through the foregoing steps, so that image recognition can be performed on the second image through an image recognition model. After performing image recognition on the second image, a recognition result can be obtained.
[0052] Exemplarily, the image recognition model can recognize the second image to give information related to the second image, and this related information is the recognition result. It should be noted that the recognition result can be text.
[0053] Exemplarily, the user can also input query information for the second image, so that the image recognition model can use the second image and the query information as input parameters together to recognize the second image and obtain a recognition result. For example, the smart glasses can be configured with a microphone, so that the user can utter voice content. After the smart glasses obtain the voice content, the voice content is converted into text content through voice-to-text conversion, and then the text content and the second image are used as input parameters of the image recognition model to obtain a recognition result.
[0054] It can be understood that the image recognition module can be a model obtained through pre-training. The image recognition model can be processed through quantization, pruning, etc., so that it is easier to be deployed to the terminal device, that is, deployed to the smart glasses, and then the smart glasses can directly call the image recognition module.
[0055] Optionally, in order to reduce the power consumption of the smart glasses and improve the efficiency of performing image recognition on the second image, the image recognition model can also be deployed in an electronic device outside the smart glasses, where the smart glasses are communicatively connected to the electronic device. Thus, the smart glasses can send the second image to the electronic device, and then the electronic device performs image recognition on the second image through the image processing model to obtain a recognition result, and then sends the recognition result back to the smart glasses.
[0056] The image recognition method provided by this application acquires a first image with a first resolution based on the image acquisition module; determines a target area in the first image based on the eye movement fixation information collected by the eye movement information acquisition module; determines the image acquisition area of the image acquisition module through the target area; acquires an image of the image acquisition area with a second resolution through the image acquisition module to obtain a second image, where the second resolution is greater than the first resolution; and performs image recognition on the second image through an image recognition model. If image recognition is performed by collecting a higher-resolution image and then combining the annotation of eye movement fixation information, the power consumption will be relatively high. In the solution of this application, first, a target area corresponding to the eye movement fixation information is determined through a first image with a lower resolution, and then the image acquisition area is determined by this target area, so as to only acquire a second image corresponding to the high-resolution target acquisition area, and then perform image recognition on the second image. That is to say, the second image with a higher resolution is not a full-size image. Therefore, the entire processing process does not involve full-size high-resolution images, resulting in lower power consumption required for image recognition through smart glasses. In addition, the second image determined by combining the eye movement fixation information can also improve the accuracy of subsequent image recognition.
[0057] Please refer to Figure 4 , Figure 4 which shows the flowchart of the image recognition method provided by the embodiment of this application. This image recognition method can be applied to the processing module in the smart glasses shown in Figure 1 . Specifically, this method may include steps S210 to S2120.
[0058] Step S210: Acquire a fourth image with a second resolution based on the image acquisition module.
[0059] Step S220: Perform a resolution reduction operation on the fourth image to obtain the first image.
[0060] In some embodiments, the processing module configured in the smart glasses may include an Image Signal Processor (ISP). Thus, a fourth image with a second resolution can be directly acquired through the image acquisition module, that is, a fourth image with a higher resolution is obtained. Then, a resolution reduction operation is performed on the fourth image to obtain the first image. Exemplarily, the resolution reduction operation can be performed on the fourth image through the image signal processor to obtain the first image.
[0061] Exemplarily, the resolution reduction operation may include a downsampling operation.
[0062] Step S230: Obtain the eye movement fixation information collected by the eye movement information acquisition module, where the eye movement fixation information includes a first position of the eye movement fixation point in a first coordinate system, and the first coordinate system is a coordinate system established based on the eye movement information acquisition module.
[0063] Step S240: Obtain the conversion parameters for mapping the first coordinate system to a second coordinate system, where the second coordinate system is a coordinate system established based on the image acquisition module.
[0064] Step S250: Convert the first position to a second position in the second coordinate system based on the conversion parameters.
[0065] For some embodiments, the eye movement fixation information collected by the eye movement information acquisition module can be used. The eye movement fixation information includes a first position of the eye movement fixation point in a first coordinate system. As can be known from the foregoing introduction, the first coordinate system is a coordinate system established based on the eye movement information acquisition module.
[0066] That is to say, through the eye movement information acquisition module, the first position in the first coordinate system can be obtained, and this first position can be used to represent the position where the user's attention is located in the first coordinate system.
[0067] Furthermore, the conversion parameters for mapping the first coordinate system to a second coordinate system can be obtained, where the second coordinate system is a coordinate system established based on the image acquisition module. Exemplarily, the conversion parameters can be represented by P1, the parameters of the first coordinate system can be represented by P2, the rotation matrix for transforming from the first coordinate system to the second coordinate system can be represented by R 12 and the translation vector for transforming from the first coordinate system to the second coordinate system can be represented by T 12 and the internal parameter matrix of the eye movement information acquisition module can be represented by K. Thus, the conversion parameters for mapping the first coordinate system to the second coordinate system are taken as P1 = K * P2 [R 12 |T 12 .
[0068] Therefore, after obtaining the conversion parameters, the first position can be converted to a second position in the second coordinate system based on the conversion parameters. That is to say, in the second coordinate system, the position where the user's attention is located can be represented by the second position.
[0069] Step S260: Determine the target area in the first image based on the second position.
[0070] After obtaining the second position, the target area can be further determined in the first image based on the second position. It can be understood that the second position may only correspond to a relatively small area in the first image. Therefore, it is necessary to combine the middle position of the second position in the first image to determine the target area.
[0071] For some embodiments, step S260 may include step S261 and step S262.
[0072] Step S261: Determine the target object corresponding to the second position in the first image.
[0073] Step S262: Determine the area corresponding to the target object in the first image as the target area.
[0074] First, the target object corresponding to the second position can be determined in the first image. Exemplarily, the middle position corresponding to the second position can be determined in the first image, and then the corresponding target object can be determined according to the middle position.
[0075] Optionally, an object recognition model can be used to determine the corresponding target object according to the middle position determined in the first image.
[0076] Please refer to Figure 5 , Figure 5 which shows a schematic diagram of determining the target area provided by an embodiment of the present application. In Figure 5 , the first image 510 is shown, and the middle position corresponding to the second position in the first image is 520. Thus, the first image 510 marked with the middle position 520 can be used as the input parameter of the object recognition model to obtain the target object as the sign 530.
[0077] Furthermore, the area corresponding to the target object in the first image can be used as the target area. Please continue to refer to Figure 5 the first image 510 in. At this time, the area 540 corresponding to the sign 530 in the first image 510 can be used as the target area. It should be noted that Figure 5 the dashed box 540 shown in is used to represent the target area.
[0078] Exemplarily, the area corresponding to the target object in the first image can be the smallest rectangular area including the target object.
[0079] Among them, the method of determining the target area through step S261 and step S262 can be called the dynamic cropping method, that is, the size of the obtained target area is associated with the size of the target object.
[0080] It can be understood that after determining the area corresponding to the target object in the first image, the proportion of this area in the first image can be further obtained. If the proportion is small, this area can be directly used as the target area; if the proportion is large, in order to improve the accuracy of subsequent image recognition, the entire area of the first image can be used as the target area at this time. Thus, the subsequently determined image acquisition area is essentially the maximum size of the image that the image acquisition module can support for acquisition, that is, the image size corresponding to the entire area in the sensor of this image acquisition module.
[0081] Exemplarily, a threshold proportion can be set. If the proportion in the first image is less than or equal to this threshold proportion, this area can be directly used as the target area; if the proportion in the first image is greater than this threshold proportion, the entire area of the first image can be used as the target area. It should be noted that the entire area of the first image is used to represent all the picture areas of this first image.
[0082] Please refer to Figure 6 , Figure 6 which shows a schematic diagram of determining the target area provided by an embodiment of the present application. In Figure 6 , a first image 610 is shown, and the area 620 corresponding to the target object determined through eye movement fixation information in the first image. If the proportion of this area in the first image 610 is greater than the threshold proportion, the entire area of the first image 610 can be determined as the target area.
[0083] The above method of using the entire area of the first image as the target area can also be called an uncropped method. Through this uncropped method, more background or environmental information in the second image can be retained during subsequent image recognition, and in some scenarios, the recognition will be more accurate.
[0084] For some other embodiments, step S260 may further include steps S263 to S265.
[0085] Step S263: Obtain a reference area of a preset size.
[0086] Step S264: Determine an intermediate position corresponding to the second position in the first image.
[0087] Step S265: Use the intermediate position as the center of the reference area, and the image area corresponding to the reference area in the first image as the target area.
[0088] As can be seen from the foregoing introduction, the second position may only correspond to a relatively small area in the first image. Therefore, a reference area can be determined, and the size of the reference area can be a preset size set in advance. Thus, an intermediate position corresponding to the second position is determined in the first image, and then the intermediate position is used as the center of the reference area, and the image area corresponding to the reference area in the first image is used as the target area.
[0089] That is to say, in the first image, a target area can be determined with the intermediate position as the center and the preset size of the reference area as the area size.
[0090] Exemplarily, please refer to Figure 7 , Figure 7 which shows a schematic diagram of determining the target area provided by an embodiment of the present application. In Figure 7 , a first image 710 is shown, and the intermediate position corresponding to the second position in the first image is 720. Further, the intermediate position 720 can be used as the center of the reference area 730, and the image area corresponding to the reference area 730 in the first image 710 is used as the target area.
[0091] Among them, the method of determining the target area through steps S263 to S265 can be called a fixed cropping method, that is, the size of the obtained target area is fixed.
[0092] Exemplarily, if the maximum size of the image supported by the image acquisition module used corresponds to a pixel value of 12M, the pixel value corresponding to the preset size of the reference area can be set to 6M.
[0093] It can be seen that the method of determining the target area through a reference area of a preset size can determine the target area without identifying the target object, which can reduce power consumption to a certain extent.
[0094] Optionally, the above steps S240 to S260 can also be executed and completed by an electronic device having a connection relationship with the electronic device. Specifically, after obtaining the eye movement fixation information in step S230, the eye movement fixation information and the first image are sent to the electronic device connected to the smart glasses to instruct the electronic device to determine the target area in the first image and send the target area to the smart glasses; thus, the smart glasses can receive the target area. The electronic device can be a smart phone or a cloud server. Optionally, when the electronic device is a cloud server, the smart glasses can also be connected to the cloud server through the smart phone, so as to realize communication with the cloud server.
[0095] Step S270: Determine the image acquisition area of the image acquisition module through the target area.
[0096] Among them, step S270 has been introduced in detail in the foregoing embodiments and will not be elaborated here.
[0097] Step S280: Obtain the first attitude information of the image acquisition module corresponding to the first image and the second attitude information of the image acquisition module corresponding to the second image.
[0098] Step S290: Determine the transformation parameters for transforming from the first attitude information to the second attitude information.
[0099] Step S2100: Adjust the image acquisition area based on the transformation parameters to obtain a target acquisition area.
[0100] The image acquisition module performs image acquisition on the target acquisition area at a second resolution to obtain a second image.
[0101] It can be understood that there may be a delay between the first moment when the first image is obtained and the second moment when the second image is obtained. Therefore, the user wearing the smart glasses may move or shake between the first moment and the second moment, which may cause the attitude information of the image acquisition module to be different at the first moment and the second moment, thereby reducing the accuracy of the determined image acquisition area.
[0102] Therefore, in some embodiments, the first attitude information of the image acquisition module corresponding to the first image and the second attitude information of the image acquisition module corresponding to the second image can be obtained. Then, based on the first attitude information and the second attitude information, the image acquisition area is adjusted, and then image acquisition is performed on the adjusted image acquisition area to obtain a second image.
[0103] Exemplarily, an inertial measurement unit (IMU) can also be provided in the smart glasses, and the inertial measurement unit is connected to the processing module. Thus, the first attitude information and the second attitude information can be collected through the inertial measurement unit.
[0104] Therefore, if the difference between the second attitude information and the first attitude information is large, it indicates that the user has had a large position change or shake between the first moment and the second moment; on the contrary, if the difference between the second attitude information and the first attitude information is small, it indicates that the user has only had a small position change or shake between the first moment and the second moment.
[0105] Furthermore, the transformation parameters for transforming from the first attitude information to the second attitude information can be determined. Then, based on the transformation parameters, the image acquisition area is adjusted to obtain a target acquisition area.
[0106] Then, the image acquisition module acquires an image of the target acquisition area at a second resolution to obtain a second image. Thus, in the embodiment of the present application, the second image is obtained by the image acquisition module acquiring the image acquisition area after the transformation parameter adjustment, that is, the obtained second image is corrected by the first pose information and the second pose information, and has a high accuracy.
[0107] Optionally, in order to improve the accuracy of the obtained second image and reduce the negative impact caused by the user's shaking or displacement. In some embodiments, the image acquisition module of the smart glasses may further include two cameras. Thus, the first image at the first resolution can be acquired by one camera, and the fifth image at the second resolution can be acquired by the other camera simultaneously. Since the first image and the second image are acquired by controlling the two cameras simultaneously, there will be no change in the perspective of the acquired images due to the user's shaking or displacement between the first image and the second image. Further, after the image acquisition area is obtained, the fifth image is directly cropped based on the image acquisition area to obtain the cropped fifth image as the second image. Exemplarily, the cropping operation may be to determine the area corresponding to the image acquisition area in the fifth image, and then crop and remove the image area outside the corresponding area, and retain the corresponding area as the second image.
[0108] It should be noted that the image acquisition area at this time is actually the image acquisition area corresponding to the other camera, that is, the image acquisition area corresponding to the camera for acquiring the fifth image.
[0109] Optionally, in some embodiments, the image acquisition module of the smart glasses may further include a camera capable of realizing dual-stream output. The camera capable of realizing dual-stream output can acquire two images with different resolutions simultaneously. Thus, the camera can be set to acquire the first image at the first resolution and the fifth image at the second resolution simultaneously. Then, similar to the aforementioned cropping operation of the fifth image based on the image acquisition area, the fifth image is cropped based on the image acquisition area corresponding to the dual-stream output camera to obtain the second image, which will not be elaborated here. By using the camera with dual-stream output, the cost of one camera can be saved, and at the same time, parameter calibration between multiple cameras is not required, saving man-hours and improving efficiency.
[0110] Step S2110: Send the second image to the electronic device connected to the smart glasses to instruct the electronic device to perform image recognition on the second image through the image recognition model, and send the recognition result of the image recognition to the smart glasses.
[0111] Step S2120: Receive the recognition result.
[0112] It can be understood that image recognition of the second image by an image recognition model generally consumes a large amount of computing resources, and a large amount of computing resources generally leads to a large amount of energy consumption. Therefore, in order to reduce the power consumption of the smart glasses and improve the efficiency of image recognition of the second image, the image recognition model can also be deployed in an electronic device outside the smart glasses, where the smart glasses are communicatively connected to the electronic device. Thus, the smart glasses can send the second image to the electronic device, and then the electronic device performs image recognition on the second image through the image processing model to obtain a recognition result, and then sends the recognition result back to the smart glasses. The smart glasses directly receive the recognition result.
[0113] In the embodiments provided in the present application, the target area can be determined by an electronic device connected to the smart glasses and image recognition of the second image is performed to obtain a recognition result, reducing the power consumption and heat generation of the smart glasses and improving the battery life. Furthermore, the second image is obtained by the image acquisition module acquiring the image acquisition area after the transformation parameter adjustment, that is, the obtained second image is corrected by the first pose information and the second pose information and has a high accuracy.
[0114] Please continue to refer to Figure 1 , Figure 1 The processing module 110 shown in can be used to perform image recognition based on the image recognition methods provided in the foregoing method embodiments.
[0115] Optionally, Figure 1 An inertial measurement unit 140 can also be provided in the smart glasses 100 shown in . The inertial measurement unit 140 is connected to the processing module 110, and the inertial measurement unit 140 is used to collect the first pose information corresponding to the first image and the second pose information corresponding to the second image. For a detailed introduction, reference can be made to the foregoing method embodiments, which will not be elaborated here.
[0116] In some embodiments, the smart glasses can include a front frame and two temple arms connected to the frame. Thus, the eye movement information acquisition module 120 can be disposed on the front frame; the control module 110, the image acquisition module 130 are disposed on the same temple arm; and a counterweight module 150 is disposed on the other temple arm to maintain the balance of the smart glasses 100.
[0117] It can be understood that in order to implement the above-mentioned method functions of the smart glasses 100, electronic devices such as a circuit board and a battery are also provided in the smart glasses, which are not shown here one by one.
[0118] Optionally, a memory can also be provided in the wearable device 100 ( Figure 1is not shown), is connected to the memory and the processing module 110.
[0119] Among them, the processing module 110 may include one or more processing cores. The processing module 110 connects various parts within the entire wearable device 100 through various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory, and by calling data stored in the memory, it executes various functions of the wearable device 100 and processes data. Optionally, the processing module 110 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processing module 110 may integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, computer programs, etc.; the GPU is responsible for rendering and drawing the display content; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processing module 110 and may be implemented separately through a communication chip. Specifically, the methods described in the foregoing embodiments may be executed by one or more processing modules 110.
[0120] For some embodiments, the memory may include a random access memory (RAM) and may also include a read-only memory. The memory can be used to store instructions, programs, codes, code sets, or instruction sets. The memory may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for implementing at least one function, instructions for implementing the following method embodiments, etc. The data storage area may also store data created during the use of the wearable device 100.
[0121] Please refer to Figure 8 , Figure 8 shows a schematic flowchart of the image recognition method provided by the embodiment of the present application. In Figure 8 , it specifically includes steps S310 to S390.
[0122] Step S310: The smart glasses collect a first image with a first resolution and eye movement fixation information.
[0123] Step S320: Send the first image with the first resolution and the eye gaze information to the smartphone.
[0124] Step S330: The smartphone determines the target area in the first image.
[0125] Step S340: The smartphone sends the target area to the smart glasses.
[0126] Step S350: The smart glasses determine the image acquisition area of the image acquisition module through the target area, and perform image acquisition on the image acquisition area at the second resolution through the image acquisition module to obtain a second image.
[0127] Step S360: The smart glasses send the second image to the smartphone.
[0128] Step S370: The smartphone sends the second image to the cloud server.
[0129] Step S380: The cloud server performs image recognition on the second image to obtain a recognition result.
[0130] Step S390: The cloud server sends the recognition result to the smart glasses through the smartphone.
[0131] For the specific introductions of the smart glasses, smartphone, cloud server in the above steps and the methods of each step, reference can be made to the foregoing method embodiments, which will not be elaborated here.
[0132] Please refer to Figure 9 , which shows a structural block diagram of a computer-readable storage medium provided by an embodiment of the present application. Program code is stored in the computer-readable medium 900, and the program code can be called by a processor to execute the method described in the foregoing method embodiments.
[0133] The computer-readable storage medium 900 may be an electronic memory such as a flash memory, EEPROM (electrically erasable programmable read-only memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium 900 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 900 has a storage space for the program code 910 that executes any method step in the above method. These program codes can be read from or written into one or more computer program products. The program code 910 can be compressed in an appropriate form, for example.
[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. An image recognition method, characterized in that, A processing module applied to smart glasses, the smart glasses further comprising an image acquisition module and an eye movement information acquisition module, the processing module being respectively connected to the image acquisition module and the eye movement information acquisition module, the method comprising: Obtaining a first image with a first resolution based on the image acquisition module; Determining a target area in the first image based on the eye movement fixation information collected by the eye movement information acquisition module; Determining an image acquisition area of the image acquisition module through the target area; Performing image acquisition on the image acquisition area by the image acquisition module at a second resolution to obtain a second image, wherein the second resolution is greater than the first resolution; Performing image recognition on the second image through an image recognition model.
2. The method according to claim 1, wherein The determining a target area in the first image based on the eye movement fixation information collected by the eye movement information acquisition module includes: Obtaining the eye movement fixation information collected by the eye movement information acquisition module, wherein the eye movement fixation information includes a first position of the eye movement fixation point in a first coordinate system, and the first coordinate system is a coordinate system established based on the eye movement information acquisition module; Obtaining a conversion parameter for mapping the first coordinate system to a second coordinate system, wherein the second coordinate system is a coordinate system established based on the image acquisition module; Converting the first position to a second position in the second coordinate system based on the conversion parameter; Determining a target area in the first image based on the second position.
3. The method according to claim 2, wherein The determining a target area in the first image based on the second position includes: Determining a target object corresponding to the second position in the first image; Determining the area corresponding to the target object in the first image as the target area.
4. The method according to claim 2, characterized in that The determining a target area in the first image based on the second position includes: Obtaining a reference area with a preset size; Determining an intermediate position corresponding to the second position in the first image; Taking the intermediate position as the center of the reference area, and taking the image area corresponding to the reference area in the first image as the target area.
5. The method according to claim 1, wherein The performing image acquisition on the image acquisition area by the image acquisition module at a second resolution to obtain a second image includes: Obtaining first attitude information of the image acquisition module corresponding to the first image and second attitude information of the image acquisition module corresponding to the second image; Determining a transformation parameter for transforming from the first attitude information to the second attitude information; Adjusting the image acquisition area based on the transformation parameter to obtain a target acquisition area; Performing image acquisition on the target acquisition area by the image acquisition module at a second resolution to obtain a second image.
6. The method according to claim 1, characterized in that, The determining a target area in the first image based on the eye movement fixation information collected by the eye movement information acquisition module includes: Obtaining the eye movement fixation information collected by the eye movement information acquisition module; Send the eye movement fixation information and the first image to an electronic device connected to the smart glasses to instruct the electronic device to determine a target area in the first image and send the target area to the smart glasses; Receive the target area.
7. The method according to claim 1, characterized in that, The performing image recognition on the second image by an image recognition model includes: Send the second image to an electronic device connected to the smart glasses to instruct the electronic device to perform image recognition on the second image by an image recognition model and send the recognition result of the image recognition to the smart glasses; Receive the recognition result.
8. The method according to claim 1, wherein The acquiring the first image with a first resolution based on the image acquisition module includes: Acquire a fourth image based on the image acquisition module at a second resolution; Perform a resolution reduction operation on the fourth image to obtain the first image.
9. An intelligent glasses, characterized in that, Comprising a processing module, an image acquisition module, and an eye movement information acquisition module, the processing module is respectively connected to the image acquisition module and the eye movement information acquisition module; The processing module is configured to perform image recognition based on the method according to any one of claims 1-8.
10. The smart glasses according to claim 9, characterized in that, The smart glasses further include an inertial measurement unit, and the inertial measurement unit is connected to the processing module; The inertial measurement unit is configured to acquire first attitude information corresponding to the first image and second attitude information corresponding to the second image.
11. A computer-readable storage medium, characterized in that, Program code is stored in the computer-readable storage medium, and the program code can be called by a processor to execute the method according to any one of claims 1-8.