Image acquisition method, image acquisition equipment and storage medium

By adopting pre-cache mode and multi-dimensional scoring model in image acquisition devices, low-latency and high-quality image frames are selected, and the image shooting delay problem caused by hardware and network factors in smart devices such as TWS headphones is solved, which improves the shooting efficiency and user experience of the device.

CN120343393APending Publication Date: 2025-07-18GEER TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510573473.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Image acquisition equipment has high image shooting delay due to low hardware performance, software processing, network transmission and dual-camera synchronization accuracy in smart devices such as TWS headsets.

Method used

The pre-cache mode is used to continuously collect and store image data. By calculating the time-stamp difference of the instruction response time and image data and image score, the target image frames that meet the shooting requirements are selected to reduce the delay caused by device factors interference.

Benefits of technology

Through the pre-cache mechanism and multi-dimensional scoring model, low-latency and optimal quality target image frames are selected to reduce shooting delay caused by interference from the device itself, and improve user experience and device performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343393A_ABST
    Figure CN120343393A_ABST
Patent Text Reader

Abstract

The invention discloses an image acquisition method, image acquisition equipment and a storage medium, and relates to the technical field of image processing, and the image acquisition method comprises the following steps: if a pre-caching mode is triggered, continuously acquiring and storing image data; responding to a shooting instruction, and determining a time difference between instruction response time and the timestamp of each frame of image data; determining a target image frame according to the time difference and an image score of the image data; and taking the target image frame as a target image corresponding to the shooting instruction. On the basis, the candidate frame closest to the actual intention of the user can be quickly positioned based on the time difference between the instruction response time and the image timestamp, the target image frame with the optimal quality is further screened out in combination with the image score, the target image frame serves as the image needing to be shot, the situation that the image delay is high due to time consumption of hardware processing is avoided, and the user experience is improved. And the shooting delay is effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technologies, and in particular, to an image acquisition method, an image acquisition device, and a storage medium. Background Art

[0002] Devices such as intelligent earphones integrated with cameras, such as TWS earphones (True Wireless Stereo) and intelligent glasses, have extensive application requirements in fields such as remote collaboration, security monitoring, sports recording, and online teaching.

[0003] During the image acquisition process of TWS earphones, due to interference from various factors such as hardware performance, software processing, network transmission, and low binocular synchronization accuracy, the image shooting delay is relatively high.

[0004] The above content is only used to assist in understanding the technical solution of this application, and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main purpose of this application is to provide an image acquisition method, an image acquisition device, and a storage medium, aiming to solve the technical problem of relatively high image shooting delay.

[0006] To achieve the above purpose, this application proposes an image acquisition method, and the method includes: If the pre-buffering mode is triggered, continuously acquire and store image data; In response to a shooting instruction, determine the time difference between the instruction response time and the timestamp of each frame of the image data; Determine a target image frame according to the time difference and the image score of the image data; Use the target image frame as the target image corresponding to the shooting instruction.

[0007] In an embodiment, after the step of determining a target image frame according to the time difference and the image score of the image data, the method further includes: If the shooting mode is single-sided shooting, execute using the target image frame as the target image corresponding to the shooting instruction; If the shooting mode is bilateral shooting, determine a combined matching frame according to the target image frame; Use the combined matching frame as the target image corresponding to the shooting instruction.

[0008] In an embodiment, the step of determining a combined matching frame according to the target image frame when the shooting mode is bilateral shooting includes: Determine a first image frame and a second image frame respectively corresponding to the target image frame, and a target time difference between the first image frame and the second image frame; Determine a joint score value between the image data corresponding to the first image frame and the second image frame according to the target time difference, and determine the joint score between the image data corresponding to the second image frame and the first image frame; Set the image frame corresponding to the minimum value of the joint score as the joint image frame.

[0009] In one embodiment, before the step of determining the target image frame according to the time difference and the image score of the image data, the method further includes: Obtain the device status information of the image acquisition device; Determine the shooting mode of the image acquisition device according to the shooting instruction and / or the device status information.

[0010] In one embodiment, the step of determining the target image frame according to the time difference and the image score of the image data includes: Obtain the stability score, environment score, and quality score corresponding to the image score; Determine the weight coefficients corresponding to the stability score, the environment score, the quality score, and the time difference; Calculate the target score of each frame of the image data according to the weight coefficients corresponding to the stability score, the environment score, the quality score, and the time difference respectively; Determine the target image frame according to the target score.

[0011] In one embodiment, before the step of determining the target image frame according to the time difference and the image score of the image data, the method further includes: Obtain the inertial sensing data and environmental sensing data corresponding to the image data; Determine the stability score according to the inertial sensing data, and determine the environment score according to the environmental sensing data; Determine the quality score according to the visual information of the image data, where the visual information includes at least one of contrast, clarity, and exposure.

[0012] In one embodiment, after the step of continuously collecting and storing image data if the pre-buffering mode is triggered, the method further includes: If the shooting instruction is not detected within the caching period, clear the image data cached in the cache address; Control the video module of the image acquisition device to enter the standby state.

[0013] In one embodiment, the pre-buffering mode is triggered when at least one of the following conditions is met: A pre - shooting instruction is received; A device operation instruction is detected; The preset cache time point is reached.

[0014] In addition, to achieve the above object, the present application also proposes an image acquisition device, which includes: a memory, a processor, and a computer program stored on the memory and operable on the processor. The computer program is configured to implement the steps of the image acquisition method as described above.

[0015] In addition, to achieve the above object, the present application also proposes a storage medium, which is a computer - readable storage medium. A computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the image acquisition method as described above.

[0016] One or more technical solutions proposed by the present application have at least the following technical effects: When triggering the pre - cache mode of the shooting process, continuously collect and store image data, then respond to the shooting instruction, determine the time difference between the instruction response time and the time stamp of each frame of image data, then select the target image frame according to the time difference and the image score of the stored image data, and finally use the target image frame as the target image corresponding to the shooting instruction. Based on this, before shooting, continuously cache the image data within a period of time, and then at the time of shooting, determine the image frame that meets the current shooting requirements through the pre - cached image data, and use this image frame as the target image. In this way, by combining the image score and the time difference, filter out the target image frame with low latency and optimal quality, and use the target image frame as the image to be shot, reducing the shooting delay caused by the interference of the device's own factors. Description of the Drawings

[0017] The accompanying drawings here are incorporated into the specification and form a part of the specification, showing the embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0018] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0019] Figure 1 It is a schematic flowchart provided for the first embodiment of the image acquisition method of the present application; Figure 2 It is a schematic flowchart provided for the second embodiment of the image acquisition method of the present application; Figure 3It is a schematic flowchart provided for the fourth embodiment of the image acquisition method of this application; Figure 4 It is a schematic flowchart of the image acquisition method provided by combining each embodiment of this application; Figure 5 It is a schematic diagram of the device structure of the hardware operating environment involved in the image acquisition method in the embodiments of this application.

[0020] The realization of the purpose, functional characteristics and advantages of this application will be further described with reference to the embodiments and the accompanying drawings. Specific Embodiments

[0021] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application.

[0022] During the image acquisition process of TWS earphones, due to various factors such as hardware performance, software processing, network transmission, and low dual-camera synchronization accuracy, the image shooting delay is relatively high.

[0023] Based on this, the main solution of the embodiments of this application is: if the pre-buffering mode is triggered, continuously acquire and store image data; Respond to the shooting instruction, and determine the time difference between the instruction response time and the timestamp of each frame of the image data; Determine the target image frame according to the time difference and the image score of the image data; Use the target image frame as the target image corresponding to the shooting instruction.

[0024] Specifically, before shooting, continuously cache the image data within a period of time. Subsequently, during shooting, determine the image frame that meets the current shooting requirements through the pre-cached image data, and use this image frame as the target image or shoot based on the acquisition parameters of the image frame. In this way, by combining the image score and the time difference, filter out the target image frame with lower delay and optimal quality, and use the target image frame as the image to be shot, reducing the shooting delay caused by the interference of the device's own factors.

[0025] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, or an image acquisition device that can implement the above functions. Among them, the image acquisition device includes earphones with a video module such as TWS earphones, a smart home system with a video module, an online education system, a security monitoring system, and other wearable devices that need to perform image shooting through the video module. Thereby improving the user experience and market competitiveness of these devices in actual use and promoting the development of intelligent wearable technology. Hereinafter, TWS earphones will be used as an example to illustrate this embodiment and the following embodiments.

[0026] To better understand the technical solution of the present application, the following will be described in detail in conjunction with the accompanying drawings of the specification and specific implementation manners.

[0027] Based on this, an embodiment of the present application provides an image acquisition method. Referring to Figure 1 , Figure 1 is a schematic flowchart of the first embodiment of the image acquisition method of the present application.

[0028] In this embodiment, the image acquisition method includes steps S10 to S40: Step S10, if the pre-buffering mode is triggered, continuously acquire and store image data.

[0029] In this embodiment, after the TWS earphone is awakened, the camera is immediately started and enters the pre-buffering mode. In the pre-buffering mode, the TWS earphone continuously acquires and stores image data within a preset time period, including the time stamp corresponding to each frame of image data.

[0030] Optionally, when the camera of the TWS earphone is in the standby state, when a pre-shooting instruction is received, or when a preset standby time point or a device operation instruction is reached, the pre-buffering mode is triggered. Among them, the pre-shooting instruction can be a voice instruction, such as a control instruction like "turn on the camera". When a device operation instruction is detected, such as when the TWS earphone changes from the shutdown state to the on state, the pre-buffering mode is triggered and image data is cached; when a preset standby time point is reached, such as it is pre-set that shooting is required 10 seconds after the earphone is turned on, the pre-buffering mode is entered 5 seconds after the power-on and image data is continuously acquired. In addition, if the above multiple conditions are met simultaneously, the pre-buffering mode can also be triggered.

[0031] Exemplarily, after the user wakes up the camera of the TWS earphone through a voice instruction of "ready to take a photo", the TWS earphone stores 150 frames of image data within the most recent 5 seconds through the memory buffer. When it exceeds the limit, the earliest frame is overwritten in chronological order, so as to achieve the effect of continuously acquiring and storing image data. By continuously acquiring and storing image data, in the subsequent shooting process, the image acquisition parameters corresponding to the best shooting image frame are selected based on the pre-buffered images for shooting, or the best shooting image frame is used as the shooting image, thereby reducing the shooting delay caused by interference from the device itself.

[0032] Optionally, to improve the battery life of the TWS earphones, if no user shooting instruction is detected after a preset period, unnecessary cached data will be cleared and the camera will be turned off. For example, if the preset period is 5 seconds, after continuously collecting image data in the most recent 5 seconds, if no shooting instruction is detected, the cache resources will be continuously released so that the current cached content is the image data from the previous 5 seconds to the present, or the camera will be directly turned off to minimize the device power consumption. Therefore, after step S10, if no shooting instruction is detected within the cache period, the image data stored in the cache address will be cleared, and then the video module of the image acquisition device will be controlled to enter the standby state.

[0033] Step S20: In response to the shooting instruction, determine the time difference between the instruction response time and the timestamp of each frame of the image data.

[0034] In this embodiment, after the user presses the shooting virtual button of the remote application through a mobile phone or issues a shooting instruction to the TWS earphones by voice, the TWS earphones record the instruction response time T at this moment, then traverse all the image frame data in the pre-cached queue, extract the timestamp t corresponding to each frame of the image, and then calculate the time difference Δt = T - t between the two.

[0035] For example, if the user issues a shooting instruction at 13:00:05.200, the earphones traverse the timestamps of 130 frames within the time range of 13:00:02.000 to 13:00:05.000 in the cache queue and calculate the difference between each frame and the instruction time.

[0036] Step S30: Determine the target image frame according to the time difference and the image score of the image data.

[0037] The image score can be the score value of the acquisition quality of the current frame of the image, or the score of the current frame of the image in multiple dimensions such as quality, stability, and environment. Among them, each score corresponds to a score weight value.

[0038] In this embodiment, the target score of each frame of the image can be calculated through the time difference and the image score of each frame of the image, and then the target image frame is determined based on the target score. Among them, the target score (also called the final score) of each frame of the image can be calculated through the time difference and the image score, or multiple sub-score items of the time difference and the image score, and then the image that meets the requirements is selected from the target scores as the target image frame.

[0039] As an optional implementation manner for calculating the target score, the calculation formula for the target score Score of each frame of the image is as follows: Score = α·ΔT + Ω·S_ima, Among them, α and Ω are scoring weight values, ΔT is the time difference, and S_ima is the image score.

[0040] If the image score S_ima includes multiple sub-scoring items, each sub-scoring item corresponds to a scoring weight value respectively.

[0041] It should be noted that in a normal shooting scenario, the greater the time difference, the lower the similarity between the pre-stored image frame and the image to be shot at the current moment. In the above formula, the target score is positively correlated with the time difference. Therefore, the smaller the target score of the image frame, the more the image frame meets the actual requirements. At the same time, the higher the image quality, the smaller the image score.

[0042] Optionally, in addition to calculating through the score value and actual parameters, a time priority coefficient corresponding to the time difference can also be set. For example, in the calculation formula Score = α·ΔT + Ω·S_ima, a time priority coefficient X (X < 0) is added, that is, Score = α·X·ΔT + Ω·S_ima. In this way, when the weight value is fixed, the influence of the time difference on the target score can be reduced through the time priority coefficient, so as to improve the effectiveness and accuracy of target image frame screening.

[0043] Exemplarily, in a certain image frame, the time difference ΔT = 50ms, the weight value is 0.6, the image score is 92, and the weight value is 0.4. The target score of this image frame is 0.6×50 + 0.4×92 = 66.8. Among them, this score value is the lowest among all image frames, so this image frame is selected as the target image frame.

[0044] Step S40, use the target image frame as the target image corresponding to the shooting instruction.

[0045] In this embodiment, the TWS earphone includes single-sided shooting and double-sided shooting, that is, shooting with the camera of the left ear or the right ear, or shooting simultaneously with the cameras of both ears. Therefore, when the shooting mode of the earphone is single-sided shooting, directly use the target image frame as the target image corresponding to the shooting instruction, so as to directly select the image that meets the requirements from the pre-buffered image frames, and avoid shooting delay caused by the interference of the device's own factors.

[0046] Optionally, the TWS earphones are provided with a wired feedback learning and adaptive optimization mechanism, which can use the automatically evaluated image quality data obtained after shooting or the actual feedback information of the user as input, and continuously update and optimize the parameters in the dynamic attention mechanism by using the online feedback learning algorithm. This enables the processing system of the earphones to adaptively adjust the weights of various features according to the actual usage scenario and the user's usage habits, dynamically optimize the system performance, and ensure that a high shooting quality and user experience can always be maintained in different application scenarios. At the same time, it enables the earphones to have long-term scene adaptation ability, and continuously improve the shooting accuracy and user experience over time.

[0047] It should be noted that the above parameters are only for explanation and are not intended to limit the present application.

[0048] This embodiment provides an image acquisition method. Through the cooperation of the pre-buffering mechanism of the image and the multi-dimensional scoring model, in the case where the hardware response delay objectively exists, the time difference between the shooting instruction and the pre-stored image is incorporated into the scoring process of the image data, and the target image frame is selected by comprehensively scoring the overall image. Then, in combination with the shooting mode of the image acquisition device and the actual requirements, the target image frame is used as the target image, thereby reducing the shooting delay caused by the interference of the device's own factors.

[0049] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as the above first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 2 , after step S30, it further includes steps S50 to S60: Step S50, if the shooting mode is bilateral shooting, determine the joint matching frame according to the target image frame.

[0050] When the shooting mode of the earphones is bilateral shooting, since the target image frame is determined based on the actual score and time difference of the image, the target image frames corresponding to the two earphone units on both sides are usually not the same frame. For example, the target image frame of the left ear is the image of the 20th frame, and the target image frame of the right ear is the image of the 10th frame. At this time, if the target image frames on both sides are used as the joint matching frame, the synchronization accuracy of the dual cameras will be relatively low, and it will be difficult to achieve accurate image stitching.

[0051] Therefore, in the case of bilateral shooting, it is necessary to calculate the joint score value between the image frames of the two earphones based on each frame of the image pre-stored in one earphone and each frame of the image pre-stored in the other earphone. For example, both the left and right earphones pre-store 20 frames of image data, and each frame has a target score. At this time, calculate the joint score of each frame with the image data on the other side, such as calculating up to 20 * 20 joint scores (ignoring duplicate calculations).

[0052] Optionally, the target image frame includes the image frame with the lowest target score value for both earphones on both sides. To reduce the computational complexity and improve the computational efficiency, the joint score value between the target image frame on one side and all the image frames on the other side can be calculated respectively, so that at most 2 * 20 - 2 joint scores need to be calculated.

[0053] Furthermore, when calculating the joint score, the closer the time difference between the two images is, the higher the probability that the two images are used as the joint image frames. Therefore, when calculating the joint score, the time difference between the two image frames also needs to be numerically corrected.

[0054] Specifically, as an optional implementation method for determining the joint matching frame, first determine the first image frame and the second image frame respectively corresponding to the target image frame, where the first image frame is the target image frame of one earphone unit on one side, and the second image frame is the target image frame of the other earphone unit on the other side. At the same time, determine the target time difference between the first image frame and the second image frame. Then, according to the target time difference, determine the joint score value between the image data corresponding to the first image frame and the second image frame, and at the same time determine the joint score value between the image data corresponding to the second image frame and the first image frame. Finally, set the image frame corresponding to the minimum value of the joint score value as the joint image frame. For example, the calculation formula for the joint score value is as follows: Joint score = Score_left + Score_right + λ·ΔT_LR, where Score_left is the target score of a certain frame of image data of the left earphone unit, Score_right is the target score of a certain frame of image data of the right earphone unit, ΔT_LR is the target time difference between the two frames of images to be calculated, and λ is the weight value of the target time difference. It can be understood that when calculating based on the above formula, usually, the joint score between the first image frame and all the image frames on the other side (the side where the second image frame is located) is calculated first, and then the joint score between the second image frame and all the image frames on the other side (the side where the first image frame is located) is calculated.

[0055] The smaller the target score value, the more it matches the shooting expectation. Therefore, the smaller the joint score value, the more the corresponding image frames on both sides match and the more they meet the actual shooting requirements. Finally, the image frame corresponding to the minimum value of the joint score needs to be set as the joint image frame. For example, if the joint score corresponding to the 8th frame on the left and the 9th frame on the right is the smallest, then these two frames of images are set as the joint image frames.

[0056] Step S60, use the joint matching frame as the target image corresponding to the shooting instruction.

[0057] In this embodiment, the combined image frames are the image frames on the left and right sides. Therefore, the combined image frames can be respectively used as the target images corresponding to the shooting instructions of the left and right cameras, so as to obtain images from the pre-stored images, and based on the pre-shooting mechanism, the user-specified moment can be captured in real time to reduce the shooting delay.

[0058] Optionally, after obtaining the target image, the image can be uploaded to the cloud or the user side such as a mobile phone, so as to splice the image through the cloud or the user side to reduce the headphone power consumption.

[0059] It should be noted that after obtaining the combined matching frames, dynamic time warping (DTW) optimization can be performed. For example, when the target time difference between the two is large, the dynamic time warping (DTW) algorithm is used to adjust the timestamps of the left and right image frames to accurately optimize the synchronization alignment of the left and right cameras and ensure that the output bilateral image times are highly consistent. The above parameters are only for explanation and are not intended to limit the present application.

[0060] This embodiment provides an image acquisition method. During bilateral shooting, the combined score between the image frames is calculated based on the target image frames of the two earphones, and the combined matching frames that meet the shooting requirements of the two earphones are determined through the combined score. Finally, the combined matching frames are used as the target images corresponding to the shooting instructions, so as to reduce the shooting delay caused by the interference of the device's own factors, capture the user-specified moment in real time based on the pre-shooting mechanism, and improve the image synchronization accuracy of the bilateral cameras.

[0061] Based on the second embodiment of the present application, in the third embodiment of the present application, the same or similar content as the above first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, after step S30, it is also necessary to determine the shooting mode of the image acquisition device so as to select the processing actions that meet the requirements based on the shooting mode.

[0062] Therefore, after step S30, it is necessary to obtain the device status information of the image acquisition device. The device status information includes the closed states of the two cameras of the earphones, and then the shooting mode of the image acquisition device is determined according to the shooting instruction and / or the device status information.

[0063] For example, if the shooting instruction requires shooting through the left ear / right ear, the shooting mode is single-sided shooting. If shooting is through both ears, the shooting mode is double-sided shooting. If the shooting instruction requires shooting through the left ear / right ear and the cameras of the left ear / right ear are in the on state and continuously collecting pre-buffered images, the shooting mode is single-sided shooting, and the same applies to double-sided shooting. When determining the shooting mode solely based on the device status information, if the camera of one of the earphones is in the on state, the shooting mode is single-sided shooting. If both are in the on state, the shooting mode is double-sided shooting. Further, if there is a conflict between the shooting instruction and the device status information, the actual requirements of the shooting instruction shall prevail.

[0064] Based on this, if the shooting mode is single shooting mode, step S40 is executed. If the shooting mode is double shooting mode, step S50 is executed.

[0065] It should be noted that the shooting mode of the image acquisition device can be determined before determining the target image frame, or after determining the target image frame. This application does not make any limitations in this regard.

[0066] This embodiment provides an image acquisition method. After determining the target image frame, by judging the shooting mode of the image acquisition device, the pre-stored target image frame is used as the shooting image corresponding to the shooting instruction in single-camera shooting, thereby reducing the delay caused during actual shooting. In double-sided shooting, based on the target image frame, a joint matching frame is selected and used as the shooting image corresponding to the shooting instruction, reducing the calibration comparison during double-camera shooting, thereby reducing the shooting delay.

[0067] Based on the first embodiment of this application, in the fourth embodiment of this application, the same or similar content as the above first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, in addition to determining the target image frame through the image score and time difference of the image data itself, in this embodiment, the target image frame is determined through multiple sub-score items of the time difference and image score, so as to screen the image frame through the scores in multiple dimensions, improve the accuracy or effectiveness of the obtained target image frame, and thus intelligently select the best image matching frame.

[0068] Specifically, in an optional implementation manner of calculating the target score, please refer to Figure 3 , step S20 further includes steps S21 to S24: Step S21, obtaining the stability score, environment score, and quality score corresponding to the image score.

[0069] In this embodiment, to select the best image matching frame, the stability score, environment score, and quality score corresponding to the image score can be obtained, so as to screen the image frames through the score items in multiple dimensions, improving the effectiveness and accuracy of the target image frame matching.

[0070] Specifically, the TWS earphone is provided with an inertial sensor and an environment sensor. Therefore, the stability score can be calculated from the inertial sensing data collected by the inertial sensor, and the environment score can be calculated from the environmental sensing data collected by the environment sensor. Therefore, the inertial sensing data and environmental sensing data corresponding to the image data can be obtained first, and then the stability score can be determined according to the inertial sensing data, and the environment score can be determined according to the environmental sensing data at the same time. It can be understood that the inertial sensing data usually includes acceleration data and angular velocity data, so the state (stationary or moving) of the earphone during shooting can be judged according to these data, and the stability of the earphone at the moment of shooting can be calculated. Among them, the lower the score, the higher the stability. The environmental sensor data usually includes parameters such as light brightness, and the environmental score can be calculated through the light brightness. The lower the score, the higher the environmental score. Among them, the pre-trained intelligent model can be used to process the environmental sensing data and inertial sensing data to obtain the corresponding environmental score and stability score.

[0071] Furthermore, the quality score can be determined through the visual information of the image data, where the visual information includes at least one of contrast, clarity, and exposure. The image quality of the image data symbolizes the acquisition effect of this frame of image. Since the lower the stability score and environment score, the better the effect, the lower the score value of the quality score, the higher the effect. Among them, the lightweight convolutional neural network (Convolutional Neural Network, CNN) or deep learning network can be used to estimate parameters such as the clarity and contrast of each frame of image data.

[0072] Optionally, in addition to determining the target image frame based on the stability score, environment score, and quality score at the same time, the target image can also be determined based on two of the sub-score items, or processed based on more dimensions of sub-score items. Therefore, this application only lists one implementation method, and the target image frame can also be determined through different score combinations. This application does not make any limitations here.

[0073] Step S22, determining the weight coefficients corresponding to the stability score, the environment score, the image quality score, and the time difference.

[0074] In this embodiment, each score and the time difference correspond to weight ratios, so as to calculate the subsequent comprehensive score based on the weight ratios. Among them, the image data can be analyzed and processed through a dynamic attention module, so that the module adaptively determines the importance between various features according to the specific requirements and environmental changes of the current scene, and thus outputs the corresponding weight coefficients.

[0075] Step S23: Calculate the target score of each frame of the image data according to the weight coefficients corresponding to the stability score, the environment score, the quality score, and the time difference respectively.

[0076] In this embodiment, the calculation formula of the target score Score of each frame of image is as follows: Score = α·ΔT + β·S_imu + γ·S_visual + δ·S_ambient, where α, β, γ, and δ are weight coefficients, ΔT is the time difference, S_imu is the stability score, S_visual is the quality score, and S_ambient is the environment score.

[0077] Step S24: Determine the target image frame according to the target score.

[0078] The smaller the target score value, the closer the frame of the image is to the expected image corresponding to the shooting instruction. Therefore, usually, the frame of the image with the lowest target score is used as the target image frame.

[0079] It should be noted that since the image score includes multiple sub-score items, during the caching process of the image data, inertial sensing data and environmental data also need to be cached simultaneously, and when clearing the cached data, these data need to be cleared simultaneously.

[0080] This embodiment provides an image acquisition method, which calculates the target score of each frame of the image according to the time difference and multiple sub-score items of the image score, and selects the target image frame that meets the requirements based on the target score value, so as to intelligently match the best image frame through a multi-dimensional scoring mechanism, so that when the target image frame is used as the target image, the image meets the acquisition requirements of the user and at the same time avoids shooting delay.

[0081] Exemplarily, in order to help understand the implementation process of the image acquisition method obtained by combining the above various embodiments, please refer to Figure 4 , Figure 4A brief flow schematic diagram of an image acquisition method is provided. Specifically, after the TWS earphones enter the pre-buffer mode, they continuously acquire and store image frames, timestamps of each image frame, IMU data (Inertial Measurement Unit Data), environmental data, etc. Subsequently, feature extraction is performed to calculate the score values and time differences corresponding to each data, including the time difference ΔT, stability score S_imu, quality score S_visual, and environmental score S_ambient. Then, the data obtained from feature extraction is passed into the dynamic attention module, and the weight coefficients α, β, γ, and δ corresponding to each score value and time difference are output. Subsequently, the comprehensive score (target score) of each image frame is calculated, that is, Score = α·ΔT + β·S_imu + γ·S_visual + δ·S_ambient. Then, the candidate frames are sorted and filtered based on the calculated comprehensive score to define the target image frames.

[0082] Then, in the single-sided mode, select the image frame with the lowest score as the candidate frame, that is, the single-sided best matching frame, and output the single-sided best matching frame, which is used as the captured image corresponding to the shooting instruction. Finally, the weights and model parameters are updated in an online feedback and adaptive learning manner, and then the process ends or continues with the content of synchronous shutter photography.

[0083] In the double-sided mode, it is necessary to determine the candidate image frames of the left and right cameras respectively. Then, after selecting the frames with the lowest scores of the left and right cameras, calculate the time difference between them, ΔT_LR = T_left + T_right. Immediately afterwards, calculate the joint score of the image data through the two-stream fusion network: Joint score = Score_left + Score_right + λ·ΔT_LR. Finally, after using DTW to adjust the time alignment of the left and right frames, select the left and right frames with the lowest joint score as the best matching frames and output them. Finally, the weights and model parameters are updated in an online feedback and adaptive learning manner, and then the process ends or continues with the content of synchronous shutter photography.

[0084] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the image acquisition method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.

[0085] This application provides an image acquisition device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the image acquisition method in the above first embodiment.

[0086] Next, refer to Figure 5, which shows a schematic structural diagram of an image acquisition device suitable for implementing the embodiments of the present application. Figure 5 The shown image acquisition device is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0087] As Figure 5 shown, the image acquisition device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM, Read Only Memory) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM, Random Access Memory) 1004. In the random access memory 1004, various programs and data required for the operation of the image acquisition device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD, Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the image acquisition device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an image acquisition device having various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be alternatively implemented or had.

[0088] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiments disclosed in the present application are executed.

[0089] The image acquisition device provided by the present application adopts the image acquisition method in the above embodiment, and can solve the technical problem of high image shooting delay. Compared with the prior art, the beneficial effects of the image acquisition device provided by the present application are the same as those of the image acquisition method provided by the above embodiment, and other technical features in the image acquisition device are the same as those disclosed in the method of the previous embodiment, which will not be elaborated here.

[0090] It should be understood that the various parts disclosed in the present application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0091] The above is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0092] The present application provides a computer-readable storage medium, on which computer-readable program instructions (i.e., computer programs) are stored, and the computer-readable program instructions are used to execute the image acquisition method in the above embodiment.

[0093] The computer-readable storage medium provided by the present application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memories, read-only memories, erasable programmable read-only memories (EPROMs), optical fibers, portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination of the above.

[0094] The above computer-readable storage medium can be included in the image acquisition device; it can also exist separately without being assembled into the image acquisition device.

[0095] The above computer-readable storage medium carries one or more programs, which, when executed by an image acquisition device, cause the image acquisition device to: continuously acquire and store image data if the pre-buffering mode is triggered; respond to a shooting instruction, and determine the time difference between the instruction response time and the timestamp of each frame of the image data; determine a target image frame according to the time difference and the image score of the image data; use the target image frame as the target image corresponding to the shooting instruction.

[0096] Computer program code for performing the operations of the present application may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0097] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0098] The modules involved in the embodiments of the present application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.

[0099] The readable storage medium provided by the present application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above image acquisition method, and can solve the technical problem of high image capture latency. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as those of the image acquisition method provided by the above embodiments, and will not be elaborated here.

[0100] The above are only some embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structural transformation made by using the content of the specification and drawings of the present application under the technical concept of the present application, or direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. An image acquisition method, characterized in that, The described image acquisition method includes: If the pre-buffering mode is triggered, continuously acquire and store image data; In response to a shooting instruction, determine the time difference between the instruction response time and the timestamp of each frame of the image data; Based on the time difference and the image score of the image data, determine the target image frame; Use the target image frame as the target image corresponding to the shooting instruction.

2. The image acquisition method according to claim 1, characterized in that After the step of determining the target image frame based on the time difference and the image score of the image data, it further includes: If the shooting mode is single-sided shooting, execute using the target image frame as the target image corresponding to the shooting instruction; If the shooting mode is double-sided shooting, determine the combined matching frame based on the target image frame; Use the combined matching frame as the target image corresponding to the shooting instruction.

3. The image acquisition method according to claim 2, wherein The step of determining the combined matching frame based on the target image frame when the shooting mode is double-sided shooting includes: Determine the first image frame and the second image frame respectively corresponding to the target image frame, and the target time difference between the first image frame and the second image frame; Based on the target time difference, determine the combined score value between the image data corresponding to the first image frame and the second image frame, and determine the combined score between the image data corresponding to the second image frame and the first image frame; Set the image frame corresponding to the minimum value of the combined score as the combined image frame.

4. The image acquisition method according to claim 2, characterized in that Before the step of determining the target image frame based on the time difference and the image score of the image data, it further includes: Obtain the device status information of the image acquisition device; Based on the shooting instruction and / or the device status information, determine the shooting mode of the image acquisition device.

5. The image acquisition method according to claim 1, wherein The step of determining the target image frame based on the time difference and the image score of the image data includes: Obtain the stability score, environment score, and quality score corresponding to the image score; Determine the weight coefficients corresponding to the stability score, the environment score, the quality score, and the time difference; Based on the weight coefficients corresponding to the stability score, the environment score, the quality score, and the time difference respectively, calculate the target score of each frame of the image data; Determine the target image frame based on the target score.

6. The image acquisition method according to claim 5, wherein Before the step of determining the target image frame based on the time difference and the image score of the image data, it further includes: Obtain the inertial sensing data and environmental sensing data corresponding to the image data; Determine the stability score based on the inertial sensing data, and determine the environment score based on the environmental sensing data; Determine the quality score based on the visual information of the image data, where the visual information includes at least one of contrast, clarity, and exposure.

7. The image acquisition method according to claim 1, characterized in that After the step of continuously acquiring and storing image data if the pre-buffering mode is triggered, it further includes: If the shooting instruction is not detected within the caching period, clear the image data cached in the cache address; Control the video module of the image acquisition device to enter the standby state.

8. The image acquisition method according to any one of claims 1 to 7, characterized in that The pre-buffering mode is triggered when at least one of the following conditions is met: A pre-shooting instruction is received; A device operation instruction is detected; The preset cache time point is reached.

9. An image acquisition device, characterized in that, The image acquisition device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program being configured to implement the steps of the image acquisition method according to any one of claims 1 to 8.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the image acquisition method according to any one of claims 1 to 8 are implemented.