Image acquisition method and device of intelligent glasses, intelligent glasses and storage medium

Through a multimodal analysis model, the photo control instructions and the initial image are comprehensively analyzed to determine the target focus object and its pixel coordinates, and the photo parameters of the smart glasses are adjusted. This solves the problems of high hardware cost and inaccurate focusing in the existing technology and achieves high-quality image acquisition.

CN120751256APending Publication Date: 2025-10-03GOERTEK INC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511264089.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-05
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing smart glasses photography solutions rely on adding eye-tracking hardware, which increases hardware costs and design complexity. Performance is affected by ambient light and individual user differences, and lacks stability and universality, making it impossible to meet users' needs for precise focusing on specific targets.

Method used

Through the multimodal analysis model, the camera control instructions and the initial image are correlated and analyzed to determine the user's target focus object and its pixel coordinates. The camera parameters of the smart glasses are adjusted according to these coordinates to achieve precise focus.

Benefits of technology

Without increasing additional hardware costs, it effectively overcomes the stability issues caused by environmental and individual differences, achieves precise positioning of the target object, and greatly improves shooting quality and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751256A_ABST
    Figure CN120751256A_ABST
Patent Text Reader

Abstract

The invention discloses an image acquisition method and device for intelligent glasses, the intelligent glasses and a storage medium, and the method comprises the steps: obtaining an initial image in response to a photographing control instruction sent by a user; inputting the photographing control instruction and the initial image into a multi-modal analysis model for correlation analysis, and determining a target focusing object and a corresponding target pixel coordinate; and adjusting photographing parameters of the intelligent glasses according to the target pixel coordinates, and performing image acquisition on the target focusing object according to the adjusted photographing parameters. According to the mode, the target focusing object expected by the user and the corresponding pixel coordinates are output through the multi-modal analysis model, so that the intelligent glasses are controlled to carry out parameter adjustment and complete image acquisition, the stability problem caused by environment and individual differences is effectively solved on the premise of not increasing extra hardware cost, and the user experience is improved. Accurate positioning of the target object is achieved, and the overall shooting quality and the user experience are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of wearable devices, and in particular to an image acquisition method and device for smart glasses, smart glasses, and a storage medium. Background Art

[0002] Existing smartglasses photography solutions have numerous limitations: First, they rely on eye-tracking hardware to achieve focus, which not only increases hardware cost and design complexity, but also suffers from performance impacted by factors like ambient lighting and individual user differences, resulting in limited stability and adaptability. Second, they employ a fixed focal length or focus at the center of the field of view, limiting focus accuracy and failing to meet users' precise focus requirements on specific targets. Summary of the Invention

[0003] The main purpose of this application is to provide an image acquisition method and device for smart glasses, smart glasses and storage medium, aiming to solve the technical problem in the prior art that smart glasses cannot meet the requirements of efficient, stable and precise focusing for taking pictures.

[0004] To achieve the above objectives, the present application proposes an image acquisition method for smart glasses, the method comprising: Responding to a photo-taking control instruction sent by a user, acquiring an initial image; Inputting the photographing control instruction and the initial image into a multimodal analysis model for correlation analysis to determine the user's target focus object and the target pixel coordinates of the target focus object on the initial image; Adjusting the photographing parameters of the smart glasses according to the target pixel coordinates to obtain adjusted photographing parameters; An image of the target focus object is captured according to the adjusted photographing parameters.

[0005] In one embodiment, the step of adjusting the photographing parameters of the smart glasses according to the target pixel coordinates to obtain the adjusted photographing parameters includes: Determine a target photographing area on the initial image according to the target pixel coordinates; Determining target photographing parameters of the smart glasses when performing image acquisition according to the target photographing area; The photographing parameters of the smart glasses are adjusted according to the target photographing parameters to obtain adjusted photographing parameters.

[0006] In one embodiment, the step of determining target photographing parameters of the smart glasses when performing image acquisition according to the target photographing area includes: Determining target exposure parameters of the smart glasses during image acquisition according to regional brightness information of the target photographing area; Determining target focus parameters of the smart glasses when performing image acquisition according to the target photographing area; Target photographing parameters are determined according to the target exposure parameters and the target focus parameters.

[0007] In one embodiment, the step of determining the target exposure parameter of the smart glasses during image acquisition based on the regional brightness information of the target photographing area includes: Determine the brightness value of each pixel coordinate in the target photographing area according to the area brightness information of the target photographing area; Calculate the average brightness of each pixel to determine the average brightness of the target area. The target exposure parameters of the smart glasses during image acquisition are determined according to the average brightness value and the target area brightness of the target photographing area.

[0008] In one embodiment, the step of inputting the photographing control instruction and the initial image into a multimodal analysis model for correlation analysis to determine the user's target focus object and the target pixel coordinates of the target focus object on the initial image includes: Inputting the photographing control instruction and the initial image into a multimodal analysis model, and performing feature extraction on the initial image according to a visual encoder in the multimodal analysis model to obtain visual features of the initial image; performing feature extraction on the photographing control instruction according to the speech encoder in the multimodal analysis model to obtain speech features of the photographing control instruction; Performing a fusion analysis on the visual features and the speech features according to the multimodal attention fusion mechanism in the multimodal analysis model to determine a target fusion feature; The target fusion feature is analyzed according to the discriminator in the multimodal analysis model to determine the target focus object of the user and the target pixel coordinates of the target focus object on the initial image.

[0009] In one embodiment, before the step of inputting the photographing control instruction and the initial image into a multimodal analysis model for correlation analysis to determine the user's target focus object and the target pixel coordinates of the target focus object on the initial image, the step further includes: Repairing bad pixels in the initial image according to the bad pixel coordinates of the camera module in the smart glasses to obtain a repaired initial image; Performing black level correction on the repaired initial image according to the optical black area of ​​the camera module to obtain a corrected initial image; Performing shadow correction on the corrected initial image to obtain a preprocessed initial image; The step of inputting the photographing control instruction and the initial image into a multimodal analysis model for correlation analysis to determine the target focus object of the user and the target pixel coordinates of the target focus object on the initial image includes: The photographing control instruction and the pre-processed initial image are input into a multimodal analysis model for correlation analysis to determine the user's target focus object and the target pixel coordinates of the target focus object on the initial image.

[0010] In one embodiment, the step of capturing an image of the target focus object according to the adjusted photographing parameters includes: Determining a target exposure parameter of the camera module in the smart glasses and a target focus parameter of the camera module according to the adjusted photographing parameters; Controlling the camera module to focus the lens according to the target focus parameter; The camera module after focusing is controlled to capture an image of the target focus object according to the target exposure parameters.

[0011] In addition, to achieve the above-mentioned purpose, the present application also proposes an image acquisition device for smart glasses, the image acquisition device for smart glasses comprising: An acquisition module, configured to acquire an initial image in response to a photo control instruction sent by a user; an analysis module, configured to input the photographing control instruction and the initial image into a multimodal analysis model for correlation analysis, and determine the target focus object of the user and the target pixel coordinates of the target focus object on the initial image; An adjustment module, configured to adjust the photographing parameters of the smart glasses according to the target pixel coordinates to obtain adjusted photographing parameters; An acquisition module is used to acquire an image of the target focus object according to the adjusted photographing parameters.

[0012] In addition, to achieve the above-mentioned purpose, the present application also proposes a pair of smart glasses, which include: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the computer program is configured to implement the steps of the image acquisition method of the smart glasses as described above.

[0013] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the image acquisition method of the smart glasses as described above are implemented.

[0014] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the image acquisition method of the smart glasses as described above.

[0015] This application obtains an initial image in response to a photo control instruction sent by the user; inputs the photo control instruction and the initial image into a multimodal analysis model for correlation analysis to determine the target focus object of the user and the target pixel coordinates of the target focus object on the initial image; adjusts the photo parameters of the smart glasses according to the target pixel coordinates to obtain adjusted photo parameters; and performs image capture of the target focus object according to the adjusted photo parameters. In the above manner, the photo control instruction and the initial image are comprehensively analyzed by the multimodal analysis model, and the target focus object expected by the user and its corresponding pixel coordinates are output, thereby controlling the smart glasses to adjust parameters and complete image capture. Without increasing additional hardware costs, it effectively overcomes the stability problems caused by environmental and individual differences, achieves accurate positioning of the target object, and greatly improves the overall shooting quality and user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0018] Figure 1 A flowchart of the first embodiment of the image acquisition method for smart glasses of the present application is provided; Figure 2 This is a brief schematic diagram of the image acquisition method of the smart glasses provided in Example 1 of this application; Figure 3 This is a schematic diagram of an initial image of the image acquisition method for smart glasses provided in Example 1 of the present application; Figure 4 This is a schematic diagram of secondary image acquisition of the image acquisition method for smart glasses provided in Example 1 of the present application; Figure 5 A flowchart of the second embodiment of the image acquisition method for smart glasses of the present application is provided; Figure 6 A flowchart of the third embodiment of the image acquisition method for smart glasses of the present application is provided; Figure 7 This is a schematic diagram of the module structure of the image acquisition device of the smart glasses according to an embodiment of the present application; Figure 8 Schematic diagram of the device structure of the hardware operating environment involved in the image acquisition method of smart glasses in the embodiment of the present application.

[0019] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0020] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.

[0021] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0022] The main solution of the embodiment of the present application is: in response to a photo control instruction sent by a user, an initial image is obtained; the photo control instruction and the initial image are input into a multimodal analysis model for correlation analysis to determine the target focus object of the user and the target pixel coordinates of the target focus object on the initial image; the photo parameters of the smart glasses are adjusted according to the target pixel coordinates to obtain adjusted photo parameters; and image capture of the target focus object is performed according to the adjusted photo parameters.

[0023] Existing smart glasses photography solutions have many limitations: First, they rely on the addition of eye-tracking hardware to achieve focus. This solution not only increases hardware cost and design complexity, but its performance is also affected by factors such as ambient light and individual differences among users, resulting in insufficient stability and universality. Second, they use a fixed focal length or field of view center focus, which limits focus accuracy and cannot meet users' needs for precise focus on specific targets. Third, traditional multi-camera fusion or pre-shot image selection solutions require users to manually select the target area. Not only is the interactive link lengthy, it can also easily interrupt the photo-taking process, resulting in a fragmented user experience.

[0024] This application provides a solution that uses a multimodal analysis model to comprehensively analyze the photo control instructions and the initial image, and outputs the user's desired target focus object and its corresponding pixel coordinates, thereby controlling the smart glasses to adjust parameters and complete image acquisition. Without increasing additional hardware costs, it effectively overcomes the stability problems caused by environmental and individual differences, achieves precise positioning of the target object, and greatly improves the overall shooting quality and user experience.

[0025] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, etc., or a pair of smart glasses capable of performing the above functions. The following uses smart glasses as an example to illustrate this embodiment and the following embodiments.

[0026] Based on this, the embodiment of the present application provides an image acquisition method for smart glasses, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the image acquisition method for smart glasses of this application.

[0027] In this embodiment, the image acquisition method of the smart glasses includes steps S10 to S40: Step S10: acquiring an initial image in response to a photographing control instruction sent by the user.

[0028] It should be noted that the smart glasses in this embodiment refer to glasses that can be worn directly by the user and can take photos. Smart glasses include but are not limited to a camera module, an audio acquisition array, a processing unit, and the glasses themselves. The camera module is the core component used to capture images or videos and includes components such as a lens, a lens barrel, a CMOS (Complementary Metal-Oxide-Semiconductor Image Sensor) image sensor, and a focus motor. The audio acquisition array is a component used to collect user voice commands and ambient audio. It typically includes one or more microphones and can use beamforming technology to enhance voice signals in a specific direction and effectively suppress ambient noise interference. The processing unit is the computing core of the smart glasses and can be integrated into the glasses themselves or connected to the smart glasses via wired or wireless means. This embodiment does not limit this.

[0029] It is understood that a photo control command is a control signal sent by the user in natural language, including a description of the photo intention and the target object. Its expression includes but is not limited to voice commands, text commands sent by the user through a terminal connected to the smart glasses, etc. However, in this embodiment, to improve interaction efficiency, the expression form is voice commands as an example.

[0030] In a specific implementation, after the smart glasses are powered on and enter a low-power standby state, their built-in audio acquisition array continues to collect ambient audio. The processing unit in the smart glasses performs real-time streaming analysis of the collected ambient audio to detect whether the user issues a voice command containing the intention to take a photo; the voice command may be expressed in the form of a specific wake-up word and specific shooting content, or directly contains a natural language description of the shooting target. If an audio signal matching the preset expression is detected, it is confirmed that the photo control command sent by the user has been received.

[0031] In a specific implementation, when the smart glasses receive a photo control command sent by the user, they control the camera module in the smart glasses to capture the image in the current field of view, use it as the initial image, and trigger the subsequent multimodal association analysis process.

[0032] Step S20: Input the photographing control instruction and the initial image into a multimodal analysis model for correlation analysis to determine the target focus object of the user and the target pixel coordinates of the target focus object on the initial image.

[0033] It should be noted that the multimodal analysis model is a large artificial intelligence model that can simultaneously process and correlate multiple different types of data (such as text, voice, and images). Its core lies in deeply understanding the semantic associations between different modal information. The multimodal analysis model can be deployed in the cloud or locally on the smart glasses. However, in this embodiment, to eliminate network transmission delays, improve system real-time performance, and prevent user voice and image data from leaving the domain to ensure privacy and security, the lightweight multimodal analysis model is chosen to be deployed in the local processing unit of the smart glasses. The lightweight multimodal analysis model is optimized through compression technologies such as model pruning and quantization. While maintaining high accuracy, it significantly reduces computational complexity and storage overhead to adapt to the strict resource constraints of smart glasses in terms of computing power, memory, and power consumption.

[0034] It can be understood that the target focus object is the subject that the user expects the smart glasses to accurately focus on, and the target pixel coordinates are the specific location of the target focus object in the image coordinate system of the initial image.

[0035] In practice, after receiving a photo control command, the smart glasses use default photo parameters to capture an initial image of the current field of view. Because these default parameters aren't optimized for the user's actual shooting intent, the focus doesn't fall on the desired subject, resulting in suboptimal imaging. Therefore, the photo control command and the initial image are simultaneously fed into the multimodal analysis model.

[0036] It should be noted that the multimodal analysis model performs visual analysis on the initial image, identifies the visual elements contained therein and their corresponding pixel coordinates; then, through the semantic understanding of the photo control instructions and the correlation analysis of the image content, it accurately determines the target focus object pointed to by the user among all identified elements and outputs its corresponding target pixel coordinates.

[0037] Step S30: adjusting the photographing parameters of the smart glasses according to the target pixel coordinates to obtain adjusted photographing parameters.

[0038] It should be noted that camera parameters refer to the core parameter set that controls the operation of the camera module, including but not limited to focus parameters, exposure parameters, and color control parameters. Focus parameters are used to control the movement of the camera module's lens motor, adjusting the distance between the lens and the image sensor to ensure a clear image of the target object on the imaging surface. Exposure parameters control the exposure behavior of the image sensor, ensuring appropriate brightness and rich details in the area where the target object is located.

[0039] It is understood that the processing unit in the smart glasses can calculate the actual physical distance between the target focus object and the camera module based on the target pixel coordinates of the target focus object, thereby obtaining corresponding spatial information; and perform brightness analysis on the area where the target focus object is located to determine corresponding brightness information. Based on the spatial information, the focus parameters in the original photographic parameters of the smart glasses are adjusted, and based on the brightness information, the exposure parameters in the original photographic parameters are adjusted to obtain adjusted exposure parameters. In this embodiment, the adjusted photographic parameters include but are not limited to adjusted exposure parameters and adjusted focus parameters.

[0040] Step S40: capturing an image of the target focus object according to the adjusted photographing parameters.

[0041] It should be noted that the processing unit in the smart glasses sends the adjusted shooting parameters to the camera module. Based on the adjusted shooting parameters, the camera module drives the focus motor and sets the exposure parameters through AF (Auto Focus) and AE (Auto Exposure) technology, and re-captures the image under the current field of view, thereby obtaining an image with precise focus, optimized exposure and centered on the target focus object.

[0042] In a feasible implementation manner, before step S20, steps A11 to A13 may also be included: Step A11: repair the bad pixels of the initial image according to the bad pixel coordinates of the camera module in the smart glasses to obtain a repaired initial image.

[0043] It should be noted that after obtaining the initial image, it is necessary to use the ISP (Image Signal Processor) to preprocess the initial image, including but not limited to bad pixel repair, black level correction, lens shading correction, and automatic white balance operations, so as to obtain the preprocessed initial image.

[0044] It can be understood that the bad point coordinates refer to the location markers of performance abnormalities caused by manufacturing defects or long-term aging on the image sensor of the camera module pre-stored in the processing unit of the smart glasses.

[0045] In practice, to eliminate interference from fixed, abnormal pixels on the image sensor and prevent bad pixels from being misidentified as target features, the initial image requires bad pixel repair. Specifically, after acquiring the initial image, the processing unit determines the bad pixel type based on the pixel values ​​corresponding to the ring point coordinates and selects the appropriate repair algorithm to repair the bad pixels in the initial image. The coordinates of all bad pixels in the initial image are traversed, and after repairing each one individually, the repaired initial image is generated.

[0046] Step A12: performing black level correction on the repaired initial image according to the optical black area of ​​the camera module to obtain a corrected initial image.

[0047] It should be noted that the optical black area refers to a non-imaging area on the image sensor that does not receive any light and is pre-stored in the processing unit.

[0048] It is understood that in order to remove the gray / color cast in dark areas of the image caused by dark current and ensure that pure black areas appear truly black, the restored original image requires black level correction. Specifically, the processing unit reads the pixel data of the optically black areas in the restored original image and calculates the average pixel value of the optically black areas. For each pixel in the restored original image, the processing unit performs the operation "pixel value minus average pixel value" to complete the black level correction for each pixel. After all pixels are corrected, the corrected original image is obtained.

[0049] Step A13: performing shadow correction on the corrected initial image to obtain a pre-processed initial image.

[0050] It should be noted that in order to eliminate the uneven brightness caused by the assembly of the camera module and ensure that the target focus object can be clearly presented at any position, the corrected initial image needs to be subjected to lens shading correction. Specifically, the processing unit obtains the pre-stored shadow correction template corresponding to the camera module, and multiplies the brightness channel of the corrected initial image by the shadow correction template pixel by pixel, thereby completing the shadow correction of each pixel. After all pixels are corrected, the pre-processed initial image is obtained.

[0051] The step S20 further includes: inputting the photographing control instruction and the preprocessed initial image into a multimodal analysis model for correlation analysis to determine the user's target focus object and the target pixel coordinates of the target focus object on the initial image.

[0052] It should be noted that the photo control command and the pre-processed initial image are synchronously input into the multimodal analysis model, and the semantic understanding of the photo control command and the image content are correlated and analyzed based on the multimodal analysis model. The target focus object indicated by the user is accurately determined among all the identified elements, and its corresponding target pixel coordinates are output.

[0053] In a feasible implementation, step S40 may further include steps B11 to B13: Step B11: determining a target exposure parameter of the camera module and a target focus parameter of the camera module in the smart glasses according to the adjusted photographing parameters.

[0054] Step B12: controlling the camera module to focus the lens according to the target focus parameter.

[0055] Step B13: Control the focused camera module to capture an image of the target focused object according to the target exposure parameters.

[0056] It should be noted that the target exposure parameters refer to the hardware setting values ​​that can achieve the optimal brightness of the area where the target focus object is located, including but not limited to shutter speed and sensitivity; the target focus parameters refer to the hardware control values ​​that can make the target focus object have a clear image, including but not limited to the drive value of the focus motor.

[0057] It is understandable that the adjusted shooting parameters are extracted and the target exposure parameters and target focus parameters are parsed therefrom. The processing unit sends the target focus parameters to the driver chip of the focus motor, and the driver chip generates a drive signal based on the target focus parameters to control the focus motor to move the lens and complete the precise focus on the target focus object. Based on the focus lock, the processing unit sends the target exposure parameters to the image sensor, and the image sensor sets the parameters according to the target exposure parameters. After the settings are completed, the focused camera module captures the image in the current field of view.

[0058] In the specific implementation, the interaction process of smart glasses is as follows: Figure 2As shown, the audio acquisition array captures the user's photo control commands, and the camera module captures the initial image. The ISP preprocesses the initial image, and the processing unit simultaneously inputs the preprocessed initial image and the photo control commands into the multimodal analysis model. The multimodal analysis model then provides feedback on the target focus object and its corresponding target pixel coordinates. The processing unit then determines adjusted photo parameters based on the target pixel coordinates and transmits these adjusted photo parameters to the camera module, causing it to recapture the image according to the adjusted photo parameters, thereby achieving precise focus on the target object and improving image acquisition efficiency.

[0059] For example, the user sends a voice signal of "take a picture of the dog" as the photo control instruction. The initial image collected at this time is as follows: Figure 3 As shown in the figure, after the target pixel coordinates are output by the multimodal analysis model and the secondary focusing is performed using the adjusted shooting parameters, the image captured by the camera module is as follows: Figure 4 shown.

[0060] This embodiment obtains an initial image by responding to a photo control instruction sent by the user; inputs the photo control instruction and the initial image into a multimodal analysis model for correlation analysis to determine the user's target photographic object and the target pixel coordinates of the target photographic object on the initial image; adjusts the photo parameters of the smart glasses according to the target pixel coordinates to obtain adjusted photo parameters; and performs image capture of the target photographic object according to the adjusted photo parameters. In the above manner, the photo control instruction and the initial image are comprehensively analyzed by the multimodal analysis model, and the target photographic object desired by the user and its corresponding pixel coordinates are output, thereby controlling the smart glasses to adjust parameters and complete image capture. Without increasing additional hardware costs, the stability issues caused by environmental and individual differences are effectively overcome, and accurate positioning of the target object is achieved, greatly improving the overall shooting quality and user experience.

[0061] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above embodiment 1 can be referred to the above introduction and will not be described in detail later. Figure 5 , the step S30 further includes steps S31 to S33: Step S31: determining a target photographing area on the initial image according to the target pixel coordinates.

[0062] It should be noted that after obtaining the target pixel coordinates, a sub-region is divided on the initial image through an edge extension algorithm based on the target pixel coordinates, thereby obtaining a target photographing region that completely includes the target focus object.

[0063] It is understood that when dividing the sub-regions, several methods may be used, including but not limited to the following: 1. If there are multiple target pixel coordinates, a certain percentage (e.g., 5%) of pixels are expanded in all directions to form the target photographing area. 2. If the target pixel coordinates are single-point coordinates, a rectangular area of ​​a preset size is generated with that point as the center to form the target photographing area. During the region generation process using the above methods, if the expanded area exceeds the edge of the image, the edge of the image is used as the boundary of the target photographing area.

[0064] Step S32: determining target photographing parameters of the smart glasses when performing image acquisition according to the target photographing area.

[0065] It should be noted that the processing unit calculates the actual physical distance between the target focus object and the camera module based on the target photography area, thereby obtaining corresponding spatial information. It also performs brightness analysis on the target photography area to determine corresponding brightness information. Based on this spatial and brightness information, the ideal values ​​of various parameters for precise focus on the target object are determined, and these ideal values ​​of each parameter type are summarized to obtain the target photography parameters.

[0066] Step S33: Adjust the photographing parameters of the smart glasses according to the target photographing parameters to obtain adjusted photographing parameters.

[0067] It should be noted that various types of parameters in the original photographing parameters of the smart glasses are adjusted according to the target photographing parameters, thereby obtaining the adjusted photographing parameters.

[0068] In a feasible implementation, step S32 may further include steps C11 to C13: Step C11: determining target exposure parameters of the smart glasses during image acquisition according to the regional brightness information of the target photographing area.

[0069] Step C12: determining target focus parameters of the smart glasses when performing image acquisition according to the target photographing area.

[0070] Step C13: determining target photographing parameters according to the target exposure parameters and the target focus parameters.

[0071] It should be noted that regional brightness information includes, but is not limited to, the brightness values ​​of each pixel coordinate within the target capture area and the brightness distribution range of the target capture area. In this embodiment, based on a preset regional brightness-exposure parameter mapping table, the target exposure parameter for the smart glasses during image acquisition can be determined by matching the brightness characteristics in the regional brightness information with the aforementioned mapping table. Alternatively, the regional brightness information can be used to further determine the average brightness value of the region, and the average brightness value can be compared with a preset baseline brightness value to further determine the target exposure parameter based on the brightness difference between the two. In addition to the above methods, other methods can also be used to determine the target exposure parameter.

[0072] It is understood that the processing unit calculates the actual physical distance between the target focus object and the camera module based on the target photography area. Combined with the optical characteristics of the camera module, the processing unit determines the target focus parameters for the smart glasses during image acquisition using a pre-set physical distance-focus position mapping relationship. The target exposure parameters and target focus parameters are then combined to obtain the target photography parameters.

[0073] In a feasible implementation, step C11 may further include steps D11 to D13: Step D11 , determining the brightness value of each pixel coordinate in the target photographing area according to the area brightness information of the target photographing area.

[0074] Step D12: Calculate the average brightness value of each pixel coordinate to determine the average brightness value of the target photographing area.

[0075] Step D13: determining a target exposure parameter of the smart glasses during image acquisition according to the average brightness value and the target area brightness of the target photographing area.

[0076] It should be noted that the regional brightness information of the target photographing area is extracted to determine the brightness value of each pixel coordinate in the target photographing area, and further a weighted average calculation is performed on the brightness values ​​of all pixel coordinates to obtain the average brightness value of the target photographing area.

[0077] It should be understood that the target area brightness refers to a preset baseline brightness value (e.g., 128L_avg) that ensures clear imaging of the target area. The difference between the target area brightness and the average brightness is calculated, where the difference = target area brightness - average brightness. The exposure compensation value is further calculated using this difference: exposure compensation = (k × difference) / preset maximum brightness value, where k is the preset system gain factor, which ranges from 1 to 2.

[0078] In the specific implementation, after determining the exposure compensation value, the exposure parameter adjustment range of the camera module at the current exposure compensation value is determined through the mapping relationship between the exposure compensation value and the exposure parameter adjustment range. Based on the original exposure parameters, the exposure parameter adjustment range of the camera module at the current exposure compensation value is superimposed. Under the condition that the hardware constraints of the camera module are met, the target exposure parameters for the smart glasses during image acquisition are generated.

[0079] This embodiment determines a target photographing area on the initial image based on the target pixel coordinates; determines target photographing parameters for the smart glasses during image acquisition based on the target photographing area; and adjusts the photographing parameters of the smart glasses according to the target photographing parameters to obtain adjusted photographing parameters. This approach achieves refined optimization of photographing parameters and ensures image quality of the target focused object.

[0080] Based on the first embodiment and / or the second embodiment of the present application, in the third embodiment of the present application, the same or similar contents as those in the above embodiments can be referred to the above introduction and will not be described in detail later. Figure 6 , step S20, further comprising steps S21 to S24: In step S21 , the photographing control instruction and the initial image are input into a multimodal analysis model, and features of the initial image are extracted according to a visual encoder in the multimodal analysis model to obtain visual features of the initial image.

[0081] It should be noted that, in this embodiment, the multimodal analysis model includes but is not limited to a visual encoder, a speech encoder, a discriminator, and a multimodal attention fusion mechanism. Among them, the visual encoder is a feature extraction module based on deep learning, which is used to extract visual features from the initial image and map the original pixels to high-dimensional semantic features. The speech encoder is a neural network used to convert continuous speech signals into discrete semantic representations. The multimodal attention fusion mechanism is a feature interaction module based on attention weights to achieve cross-modal semantic alignment. The discriminator is a decoding module for target detection and positioning, which is responsible for identifying the target focus object from the target fusion features and locating its position in the image.

[0082] It can be understood that the photo control instructions and the pre-processed initial image are synchronously input into the multimodal analysis model, and the visual encoder in the model performs hierarchical feature processing on the initial image, thereby outputting structured data containing the image content of the initial image, including but not limited to the multiple elements existing in the image, the pixel coordinates of each element in the initial image, and the object attributes of each element.

[0083] Step S22: extracting features of the photographing control instruction according to the speech encoder in the multimodal analysis model to obtain speech features of the photographing control instruction.

[0084] It should be noted that the speech encoder in the multimodal analysis model processes the camera control command, including but not limited to the following steps: converting the speech signal in the camera control command into time-frequency features such as a mel-spectrogram; capturing the temporal dependencies in the speech through a recurrent neural network or self-attention mechanism; and outputting a fixed-dimensional speech feature vector that encodes the key words in the command. Finally, the speech encoder outputs the speech features of the camera control command, including but not limited to the semantic information in the command.

[0085] Step S23: performing a fusion analysis on the visual features and the speech features according to the multimodal attention fusion mechanism in the multimodal analysis model to determine a target fusion feature.

[0086] It should be noted that the multimodal attention fusion mechanism in the multimodal analysis model performs cross-attention calculations on visual and speech features, localizing visual areas related to the speech description and enhancing semantic keywords that match the visual content. Furthermore, a multi-head attention layer establishes dynamic weighted associations between visual and speech features to obtain feature interaction results. The cross-modal interaction results are then concatenated with the original features and passed through a fully connected layer to generate target fusion features. These target fusion features incorporate both the visual details of the target focused object and the semantic constraints of the camera control instructions.

[0087] Step S24 , analyzing the target fusion features according to the discriminator in the multimodal analysis model to determine the target focus object of the user and the target pixel coordinates of the target focus object on the initial image.

[0088] It should be noted that the discriminator in the multimodal analysis model decodes the target fusion features: it generates candidate target frames, then calculates the semantic matching score between each candidate frame and the speech features, and selects the candidate target with the highest confidence as the target focus object; finally, it fine-tunes the boundary coordinates of the target at the pixel level, and outputs the center point coordinates as the target pixel coordinates.

[0089] This embodiment inputs the photo control instruction and the initial image into a multimodal analysis model, extracts features from the initial image according to the visual encoder in the multimodal analysis model, and obtains the visual features of the initial image; extracts features from the photo control instruction according to the voice encoder in the multimodal analysis model, and obtains the voice features of the photo control instruction; performs fusion analysis on the visual features and the voice features according to the multimodal attention fusion mechanism in the multimodal analysis model to determine the target fusion features; analyzes the target fusion features according to the discriminator in the multimodal analysis model to determine the target focus object of the user and the target pixel coordinates of the target focus object on the initial image. Through the above methods, accurate analysis and definition of the target object are achieved, laying the foundation for subsequent image quality improvement.

[0090] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the image acquisition method of the smart glasses of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.

[0091] This application also provides an image acquisition device for smart glasses, please refer to Figure 7 , the image acquisition device of the smart glasses includes: The acquisition module 10 is configured to acquire an initial image in response to a photographing control instruction sent by a user.

[0092] The analysis module 20 is used to input the photographing control instruction and the initial image into a multimodal analysis model for correlation analysis to determine the user's target focus object and the target pixel coordinates of the target focus object on the initial image.

[0093] The adjustment module 30 is configured to adjust the photographing parameters of the smart glasses according to the target pixel coordinates to obtain adjusted photographing parameters.

[0094] The acquisition module 40 is used to acquire an image of the target focus object according to the adjusted photographing parameters.

[0095] Optionally, the adjustment module 30 is further configured to: Determine a target photographing area on the initial image according to the target pixel coordinates; Determining target photographing parameters of the smart glasses when performing image acquisition according to the target photographing area; The photographing parameters of the smart glasses are adjusted according to the target photographing parameters to obtain adjusted photographing parameters.

[0096] Optionally, the adjustment module 30 is further configured to: Determining target exposure parameters of the smart glasses during image acquisition according to regional brightness information of the target photographing area; Determining target focus parameters of the smart glasses when performing image acquisition according to the target photographing area; Target photographing parameters are determined according to the target exposure parameters and the target focus parameters.

[0097] Optionally, the adjustment module 30 is further configured to: Determine the brightness value of each pixel coordinate in the target photographing area according to the area brightness information of the target photographing area; Calculate the average brightness of each pixel to determine the average brightness of the target area. The target exposure parameters of the smart glasses during image acquisition are determined according to the average brightness value and the target area brightness of the target photographing area.

[0098] Optionally, the analysis module 20 is further configured to: Inputting the photographing control instruction and the initial image into a multimodal analysis model, and performing feature extraction on the initial image according to a visual encoder in the multimodal analysis model to obtain visual features of the initial image; performing feature extraction on the photographing control instruction according to the speech encoder in the multimodal analysis model to obtain speech features of the photographing control instruction; Performing a fusion analysis on the visual features and the speech features according to the multimodal attention fusion mechanism in the multimodal analysis model to determine a target fusion feature; The target fusion feature is analyzed according to the discriminator in the multimodal analysis model to determine the target focus object of the user and the target pixel coordinates of the target focus object on the initial image.

[0099] Optionally, the analysis module 20 is further configured to: Repairing bad pixels in the initial image according to the bad pixel coordinates of the camera module in the smart glasses to obtain a repaired initial image; Performing black level correction on the repaired initial image according to the optical black area of ​​the camera module to obtain a corrected initial image; Performing shadow correction on the corrected initial image to obtain a preprocessed initial image; The photographing control instruction and the pre-processed initial image are input into a multimodal analysis model for correlation analysis to determine the user's target focus object and the target pixel coordinates of the target focus object on the initial image.

[0100] Optionally, the acquisition module 40 is further configured to: Determining a target exposure parameter of the camera module in the smart glasses and a target focus parameter of the camera module according to the adjusted photographing parameters; Controlling the camera module to focus the lens according to the target focus parameter; The camera module after focusing is controlled to capture an image of the target focus object according to the target exposure parameters.

[0101] The image acquisition device for smart glasses provided in this application utilizes the image acquisition method for smart glasses described in the aforementioned embodiments, resolving the technical issue in prior art where smart glasses are unable to meet the requirements for efficient, stable, and precisely focused photography. Compared to prior art, the image acquisition device for smart glasses provided in this application provides the same beneficial effects as the image acquisition method for smart glasses described in the aforementioned embodiments. Other technical features of the image acquisition device for smart glasses are the same as those disclosed in the aforementioned embodiments and are not further detailed here.

[0102] The present application provides smart glasses, which include: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the image acquisition method of the smart glasses in the above-mentioned embodiment 1.

[0103] Reference below Figure 8 , which shows a schematic diagram of the structure of smart glasses suitable for implementing the embodiments of the present application. The smart glasses in the embodiments of the present application may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 8 The smart glasses shown are merely an example and should not limit the functions and scope of use of the embodiments of the present application.

[0104] like Figure 8As shown, smart glasses may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 1002 or programs loaded from a storage device 1003 into a random access memory (RAM) 1004. RAM 1004 also stores various programs and data required for the operation of the smart glasses. Processing device 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems may be connected to I / O interface 1006: input devices 1007, such as a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008, such as a liquid crystal display (LCD), speaker, vibrator, etc.; storage device 1003, such as a magnetic tape or hard disk; and communication device 1009. The communication device 1009 can allow the smart glasses to communicate with other devices wirelessly or wired to exchange data. Although the figure shows smart glasses with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have instead.

[0105] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0106] The smart glasses provided in this application utilize the image acquisition method for smart glasses in the aforementioned embodiments, resolving the technical issue in prior art where smart glasses are unable to meet the requirements for efficient, stable, and precisely focused photography. Compared to prior art, the beneficial effects of the smart glasses provided in this application are the same as those of the image acquisition method for smart glasses provided in the aforementioned embodiments, and the other technical features of the smart glasses are the same as those disclosed in the aforementioned embodiments, and are not further elaborated here.

[0107] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0108] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0109] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, and the computer-readable program instructions are used to execute the image acquisition method for smart glasses in the above-mentioned embodiment.

[0110] The computer-readable storage medium provided herein may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including, but not limited to, wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0111] The computer-readable storage medium may be included in the smart glasses, or may exist independently without being incorporated into the smart glasses.

[0112] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the smart glasses, the smart glasses: acquire an initial image in response to a photo control instruction sent by a user; input the photo control instruction and the initial image into a multimodal analysis model for correlation analysis to determine the user's target photographic object and the target pixel coordinates of the target photographic object on the initial image; adjust the photographic parameters of the smart glasses according to the target pixel coordinates to obtain adjusted photographic parameters; and capture an image of the target photographic object according to the adjusted photographic parameters.

[0113] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0114] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.

[0115] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.

[0116] The computer-readable storage medium provided in this application stores computer-readable program instructions (i.e., a computer program) for executing the image acquisition method for smart glasses described above. This computer-readable storage medium can address the technical issue of prior art smart glasses failing to meet the requirements for efficient, stable, and precisely focused photography. Compared to prior art, the beneficial effects of the computer-readable storage medium provided in this application are similar to those of the image acquisition method for smart glasses provided in the aforementioned embodiments, and are not further elaborated here.

[0117] The present application also provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the image acquisition method for smart glasses as described above.

[0118] The computer program product provided in this application can address the technical problem in existing smart glasses that cannot meet the requirements for efficient, stable, and accurately focused photography. Compared with existing technologies, the beneficial effects of the computer program product provided in this application are the same as those of the image acquisition method of smart glasses provided in the above embodiments, and will not be elaborated here.

[0119] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.

Claims

1. An image acquisition method for smart glasses, characterized in that: The method comprises: Responding to a photo-taking control instruction sent by a user, acquiring an initial image; Inputting the photographing control instruction and the initial image into a multimodal analysis model for correlation analysis to determine the user's target focus object and the target pixel coordinates of the target focus object on the initial image; Adjusting the photographing parameters of the smart glasses according to the target pixel coordinates to obtain adjusted photographing parameters; An image of the target focus object is captured according to the adjusted photographing parameters.

2. The method according to claim 1, wherein The step of adjusting the photographing parameters of the smart glasses according to the target pixel coordinates to obtain the adjusted photographing parameters includes: Determine a target photographing area on the initial image according to the target pixel coordinates; Determining target photographing parameters of the smart glasses when performing image acquisition according to the target photographing area; The photographing parameters of the smart glasses are adjusted according to the target photographing parameters to obtain adjusted photographing parameters.

3. The method according to claim 2, wherein The step of determining target photographing parameters of the smart glasses when performing image acquisition according to the target photographing area includes: Determining target exposure parameters of the smart glasses during image acquisition according to regional brightness information of the target photographing area; Determining target focus parameters of the smart glasses when performing image acquisition according to the target photographing area; Target photographing parameters are determined according to the target exposure parameters and the target focus parameters.

4. The method according to claim 3, wherein The step of determining the target exposure parameter of the smart glasses during image acquisition according to the regional brightness information of the target photographing area includes: Determine the brightness value of each pixel coordinate in the target photographing area according to the area brightness information of the target photographing area; Calculate the average brightness of each pixel to determine the average brightness of the target area. The target exposure parameters of the smart glasses during image acquisition are determined according to the average brightness value and the target area brightness of the target photographing area.

5. The method according to any one of claims 1 to 4, characterized in that The step of inputting the photographing control instruction and the initial image into a multimodal analysis model for correlation analysis to determine the user's target focus object and the target pixel coordinates of the target focus object on the initial image includes: Inputting the photographing control instruction and the initial image into a multimodal analysis model, and performing feature extraction on the initial image according to a visual encoder in the multimodal analysis model to obtain visual features of the initial image; performing feature extraction on the photographing control instruction according to the speech encoder in the multimodal analysis model to obtain speech features of the photographing control instruction; Performing a fusion analysis on the visual features and the speech features according to the multimodal attention fusion mechanism in the multimodal analysis model to determine a target fusion feature; The target fusion feature is analyzed according to the discriminator in the multimodal analysis model to determine the target focus object of the user and the target pixel coordinates of the target focus object on the initial image.

6. The method according to any one of claims 1 to 4, characterized in that Before the step of inputting the photographing control instruction and the initial image into a multimodal analysis model for correlation analysis to determine the user's target focus object and the target pixel coordinates of the target focus object on the initial image, the method further includes: Repairing bad pixels in the initial image according to the bad pixel coordinates of the camera module in the smart glasses to obtain a repaired initial image; Performing black level correction on the repaired initial image according to the optical black area of ​​the camera module to obtain a corrected initial image; Performing shadow correction on the corrected initial image to obtain a preprocessed initial image; The step of inputting the photographing control instruction and the initial image into a multimodal analysis model for correlation analysis to determine the target focus object of the user and the target pixel coordinates of the target focus object on the initial image includes: The photographing control instruction and the pre-processed initial image are input into a multimodal analysis model for correlation analysis to determine the user's target focus object and the target pixel coordinates of the target focus object on the initial image.

7. The method according to any one of claims 1 to 4, characterized in that The step of capturing an image of the target focus object according to the adjusted photographing parameters comprises: Determining a target exposure parameter of the camera module in the smart glasses and a target focus parameter of the camera module according to the adjusted photographing parameters; Controlling the camera module to focus the lens according to the target focus parameter; The camera module after focusing is controlled to capture an image of the target focus object according to the target exposure parameters.

8. An image acquisition device for smart glasses, characterized in that: The image acquisition device of the smart glasses includes: An acquisition module, configured to acquire an initial image in response to a photo control instruction sent by a user; an analysis module, configured to input the photographing control instruction and the initial image into a multimodal analysis model for correlation analysis, and determine the target focus object of the user and the target pixel coordinates of the target focus object on the initial image; An adjustment module, configured to adjust the photographing parameters of the smart glasses according to the target pixel coordinates to obtain adjusted photographing parameters; An acquisition module is used to acquire an image of the target focus object according to the adjusted photographing parameters.

9. A pair of smart glasses, characterized in that: The smart glasses include: a memory, a processor, and an image acquisition program for the smart glasses stored in the memory and executable on the processor, wherein the image acquisition program for the smart glasses is configured to implement the steps of the image acquisition method for the smart glasses according to any one of claims 1 to 7.

10. A storage medium, characterized in that: The storage medium stores an image acquisition program for the smart glasses, and when the image acquisition program for the smart glasses is executed by the processor, the steps of the image acquisition method for the smart glasses according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Image processing method and device, camera assembly, electronic equipment and medium

    CN113473028A

  • Shooting method and device, electronic equipment and readable storage medium

    CN118381987A

  • Image-based halo processing method and device, and storage medium

    CN120070289A

  • Focusing control system and method, wearable device, medium and program product

    CN120151651A

  • Sound-vision collaborative focusing method, device and equipment based on cross-modal distillation

    CN120434504A