Image data processing method, image acquisition equipment and storage medium

Through the combination of inertial sensing data and visual information, the noise covariance of the process is dynamically updated, and the Kalman filtering algorithm is used to solve the problem of poor image correction effect caused by sensor error amplification, and the image correction effect and user experience of smart headphones are improved.

CN120343415APending Publication Date: 2025-07-18GEER TECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510571184.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, the image correction method based on an inertial sensor is continuously amplified during the integration process, resulting in poor image calibration effect. Especially in dual wireless camera applications of smart headphones, the difference in user wearing habits and postures leads to the camera viewing angle offset affecting the shooting effect.

Method used

The inertial state vector and target state vector are determined through inertial sensing data and image visual information, the process noise covariance is dynamically updated, and the Kalman filtering algorithm is combined to reflect the sensor error changes in real time and improve the image correction effect.

Benefits of technology

It enhances the adaptability and robustness of image correction, significantly improves the image correction effect, reduces sensor error accumulation, and improves the shooting quality of smart headphones in remote meetings and live broadcasts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343415A_ABST
    Figure CN120343415A_ABST
Patent Text Reader

Abstract

The invention discloses an image data processing method, image acquisition equipment and a storage medium, and relates to the technical field of image correction, and the image data processing method comprises the following steps: determining an inertial state vector of the image acquisition equipment at an image acquisition moment according to inertial sensing data, and determining an inertial state vector of the image acquisition equipment at an image acquisition moment according to image visual information and the inertial state vector; determining a target state vector; updating a current process noise covariance according to the error variation of the inertial state vector and the target state vector to obtain a target process noise covariance; and based on the target process noise covariance and a preset Kalman filtering algorithm, updating the target state vector at the next moment corresponding to the image acquisition moment. Through dynamic updating of the process noise covariance during Kalman filtering processing, the dynamic change of the sensor error is reflected in real time based on the dynamically updated noise covariance, and the image correction effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image correction technology, and particularly to an image data processing method, an image acquisition device, and a storage medium. Background Art

[0002] With the continuous progress of intelligent earphones and camera technology, dual wireless camera intelligent earphones are widely used in the fields of remote conferencing and live streaming. Differences in users' wearing habits and postures easily cause the camera view angle of the earphones to shift, affecting the shooting effect.

[0003] In related technologies, image correction is usually performed based on data collected by inertial sensors, and speed and displacement are derived through integral operations. However, the tiny errors of the sensors will be continuously amplified during the integration process, accumulating sensing drift errors and resulting in poor image calibration effects.

[0004] The above content is only used to assist in understanding the technical solution of this application, and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main purpose of this application is to provide an image data processing method, an image acquisition device, and a storage medium, aiming to solve the technical problem of poor image correction effect.

[0006] To achieve the above purpose, this application proposes an image data processing method, which includes: Determine the inertial state vector of the image acquisition device at the image acquisition moment according to inertial sensing data, and determine the target state vector according to the image visual information and the inertial state vector; Update the current process noise covariance according to the error change amount between the inertial state vector and the target state vector to obtain the target process noise covariance; Update the target state vector at the next moment corresponding to the image acquisition moment based on the target process noise covariance and a preset Kalman filtering algorithm.

[0007] In one embodiment, the step of updating the current process noise covariance according to the error change amount between the inertial state vector and the target state vector to obtain the target process noise covariance includes: Determine the forgetting factor of the current process covariance and the transpose matrix of the error change amount, where the forgetting factor is the weight ratio of the current process covariance; Determine the change weight of the first product of the error change amount and the transpose matrix according to the forgetting factor; Set the sum of the change weight and the first product, and the second product of the forgetting factor and the current process noise covariance as the target process noise covariance.

[0008] In one embodiment, the step of updating the target state vector at the next moment corresponding to the image acquisition moment based on the target process noise covariance and a preset Kalman filtering algorithm includes: Updating the state covariance according to the target process noise covariance to obtain a target covariance; Updating the Kalman gain coefficient at the next moment based on the target covariance, a preset observation noise covariance, and the Jacobian matrix of the visual measurement model and its transpose; Updating the target state vector at the next moment corresponding to the image acquisition moment based on the Kalman gain coefficient at the next moment and the preset Kalman filtering algorithm.

[0009] In one embodiment, the step of determining the inertial state vector of the image acquisition device at the image acquisition moment according to the inertial sensing data, and determining the target state vector according to the image visual information and the inertial state vector includes: Determining a preset state transition matrix, the current inertial state vector calculated at the previous moment corresponding to the image acquisition moment, and a control input matrix; Calculating the inertial state vector according to the state transition matrix, the current inertial state vector, the control input matrix, the inertial sensing data, and the current process noise covariance; Determining a visual observation value according to the image visual information, and calculating the target state vector based on the preset Kalman filtering algorithm and the inertial state vector.

[0010] In one embodiment, the step of calculating the inertial state vector according to the state transition matrix, the current inertial state vector, the control input matrix, the inertial sensing data, and the current process noise covariance includes: Determining a third product of the state transition matrix and the current inertial state vector, and a fourth product of the control input matrix and the inertial sensing data; Setting the sum of the third product, the fourth product, and the current process noise covariance as the inertial state vector.

[0011] In one embodiment, before the step of determining the target state vector according to the visual pose information and the inertial state vector, the image data processing method further includes: Obtaining the Jacobian matrix of the visual measurement model and its transpose matrix, a preset observation noise covariance, and determining the state covariance at the image acquisition moment according to the current process noise covariance; Determine the Kalman gain coefficient at the image acquisition moment according to the Jacobian matrix and its transposed matrix, the preset observation noise covariance, and the state covariance; According to the Kalman gain coefficient at the image acquisition moment and the preset Kalman filtering algorithm, calculate the Kalman filtering fusion result of the visual pose information, the Kalman gain coefficient, and the inertial state vector to obtain the target state vector.

[0012] In one embodiment, before the steps of determining the inertial state vector of the image acquisition device at the image acquisition moment according to the inertial sensing data, and determining the target state vector according to the image visual information and the inertial state vector, further include: Obtain the image information collected at the image acquisition moment and determine the target area of the image information; Obtain the image coordinate information of the visual feature points in the target area and convert the image coordinate information into three-dimensional space coordinate information; Generate the image visual information according to the observation result of the three-dimensional space coordinate information under the visual measurement model and the measurement noise of the image coordinate information.

[0013] In one embodiment, after the step of updating the target state vector at the next moment corresponding to the image acquisition moment based on the target process noise covariance and the preset Kalman filtering algorithm, the image data processing method further includes: Output a calibration prompt message according to the target state vector and obtain the correction pose corresponding to the calibration prompt message; If the correction pose meets the pose result corresponding to the target state vector, output a corrected prompt.

[0014] In addition, to achieve the above object, the present application also proposes an image acquisition device, where the image acquisition device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the image data processing method as described above.

[0015] In addition, to achieve the above object, the present application also proposes a storage medium, where the storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the image data processing method as described above.

[0016] In addition, to achieve the above object, the present application also provides a computer program product, where the computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps of the image data processing method as described above.

[0017] One or more technical solutions proposed in this application have at least the following technical effects: According to the inertial sensing data and image visual information collected by the image acquisition device at the current image acquisition moment, determine the inertial state vector and target state vector that need to be corrected. Then, through the error change amount between the inertial state vector and the target state vector, update the process noise covariance to obtain the target process noise covariance. Finally, update the target state vector at the next acquisition moment through the target process noise covariance and the preset Kalman filtering algorithm. By fusing Kalman filtering and inertial sensing data, on the basis of image correction processing, realize the dynamic update of the process noise covariance during Kalman filtering processing, so as to reflect the dynamic change of the sensor error in real time based on the dynamically updated target process noise covariance, improve the adaptive ability and robustness of the algorithm when correcting the image through the Kalman filtering algorithm and inertial sensing data, and thus improve the image correction effect. Brief Description of the Drawings

[0018] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.

[0019] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 It is a schematic flowchart provided for the first embodiment of the image data processing method of this application; Figure 2 It is a schematic flowchart provided for the second embodiment of the image data processing method of this application; Figure 3 It is a schematic flowchart provided for the third embodiment of the image data processing method of this application; Figure 4 It is a schematic flowchart of the brief image data processing method provided by combining the embodiments of this application; Figure 5 It is a schematic diagram of the device structure of the hardware operating environment involved in the image data processing method in the embodiments of this application.

[0021] The implementation, functional features, and advantages of the purpose of this application will be further described with reference to the embodiments and the drawings. Detailed Embodiments

[0022] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0023] With the continuous progress of smart headphones and camera technology, dual wireless camera smart headphones are widely used in the fields of remote conferencing and live streaming. Differences in users' wearing habits and postures easily cause the viewing angle of the headphone camera to shift, affecting the shooting effect.

[0024] In the related art, image correction is usually performed based on the data collected by inertial sensors, and the speed and displacement are deduced through integral operations. However, the tiny errors of the sensors will be continuously amplified during the integration process, accumulating sensing drift errors, resulting in poor image calibration effects.

[0025] The main solution of the embodiments of the present application is: determining the inertial state vector of the image acquisition device at the image acquisition moment according to the inertial sensing data, and determining the target state vector according to the image visual information and the inertial state vector; Updating the current process noise covariance according to the error change amount between the inertial state vector and the target state vector to obtain the target process noise covariance; Updating the target state vector at the next moment corresponding to the image acquisition moment based on the target process noise covariance and a preset Kalman filtering algorithm.

[0026] Specifically, on the basis of correcting the image by means of Kalman filtering and inertial sensing data fusion, the dynamic update of the process noise covariance during Kalman filtering processing is realized, so as to reflect the dynamic change of the sensor error in real time based on the dynamically updated target process noise covariance, improve the adaptive ability and robustness of the algorithm when correcting the image by the Kalman filtering algorithm and inertial sensing data, and further improve the image correction effect.

[0027] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, or an image acquisition device capable of implementing the above functions. Among them, the image acquisition device includes headphones with a camera, such as TWS headphones, and wearable devices such as smart home systems, online education systems, and security monitoring systems equipped with video modules that need to take images through the video module. This can improve the user experience and market competitiveness of these devices in actual use and promote the development of intelligent wearable technology. Hereinafter, taking headphones with a camera as an example, this embodiment and the following embodiments will be described. Among them, the headphones are equipped with two independent cameras (left and right ears), and can effectively capture high-precision depth (distance) information based on their inherent stereo vision structure advantages, such as sufficient baseline length. At the same time, the spatial positioning of feature points is more accurate and stable, significantly reducing the measurement error of monocular vision and improving the accuracy of attitude angle estimation.

[0028] To better understand the technical solution of the present application, the following will be described in detail in conjunction with the accompanying drawings of the specification and specific implementation manners.

[0029] The embodiment of the present application provides an image data processing method. Refer to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the image data processing method of the present application.

[0030] In this embodiment, the image data processing method includes steps S10 to S30: Step S10, determining the inertial state vector of the image acquisition device at the image acquisition moment according to the inertial sensing data, and determining the target state vector according to the image visual information and the inertial state vector.

[0031] In addition to setting a camera, the image acquisition device is also provided with an inertial sensor (Inertial Measurement Unit), and the inertial sensor usually includes a gyroscope, an accelerometer, and a magnetometer, and can collect triaxial angular velocity, linear acceleration data, and geomagnetic data. Therefore, the inertial sensing data collected by the image acquisition device includes at least one of triaxial angular velocity, linear acceleration, and geomagnetic data.

[0032] The inertial state vector refers to the predicted value of the attitude obtained by predicting the attitude by fusing angular velocity and acceleration data at the current image acquisition moment. The image visual information includes the two-dimensional and three-dimensional spatial coordinate information of the coordinates of the visual feature points in the image data. The target state vector is the corrected result obtained by correcting the predicted value of the inertial sensing data through visual observation values during the attitude fusion process using the EKF algorithm (Extended Kalman Filter, extended Kalman filter fusion algorithm), that is, the target state vector is the corrected result of the inertial state vector. It should be noted that in addition to fusing inertial and visual data through the extended Kalman filter algorithm, unscented Kalman filter, particle filter, and Madgwick filter / Mahony filter (lightweight attitude resolver) can also be used for fusing image visual information and inertial state vector.

[0033] In this embodiment, when the earphone executes a shooting instruction or the earphone is in a moving state, inertial sensing data is acquired based on an inertial sensor, and then the inertial state vector corresponding to the inertial sensing data is calculated through the prediction state equation in the state space model. At the same time, the current image data is collected, and the image visual information is calculated according to the image data, that is, the coordinate information of the visual feature points in the image is calculated, so as to calculate the visual measurement value of the attitude through this coordinate information, and then based on the correction formula in the EKF algorithm, the predicted value of the inertial sensing data is corrected according to the visual observation value, thereby obtaining the target state vector.

[0034] Exemplarily, after fusing the image visual information and the inertial state vector based on the EKF algorithm, the obtained target state vector is: , where, is the target state vector, is the horizontal attitude angle (such as the left-right deflection angle of the head), with the unit of radian (rad) or degree (°); is the vertical attitude angle (such as the up-down pitching angle of the head), with the unit of radian (rad) or degree (°); is the horizontal angle change rate (angular velocity), with the unit of rad / s; is the vertical angle change rate (angular velocity), with the unit of rad / s.

[0035] In this embodiment, the extended Kalman filter algorithm is used to fuse visual and inertial sensing data, thereby reducing the error accumulation in the integration process of inertial sensing data, and thus improving the image calibration effect in subsequent image processing.

[0036] It should be noted that the image acquisition time is consistent with the EKF filtering frequency. For example, if the filtering frequency is f, the time interval T between each image acquisition time is: T = 1 / f. Therefore, within each time interval, data needs to be continuously acquired and the corresponding inertial state vector and target state vector for this time interval need to be calculated to ensure that the device has high precision, real-time performance, and adaptability.

[0037] Step S20: Update the current process noise covariance according to the error change amount between the inertial state vector and the target state vector to obtain the target process noise covariance.

[0038] Since the traditional EKF algorithm usually takes the process noise covariance Q as a fixed matrix and cannot reflect the dynamic changes of sensor errors in real time, and in a normal image acquisition environment, the noise characteristics of sensors and vision change over time. For example, in a real scenario, the inertial sensor may have noise fluctuations due to temperature changes, vibrations, etc.; visual data may introduce instantaneous noise due to lighting changes, occlusions, or dynamic objects. If the process noise covariance and the observation noise covariance are static values, they cannot reflect the real-time changes of the noise at this time, resulting in a mismatch between the model and the real noise distribution. At this time, when the quality of sensor data fluctuates, the fusion accuracy of the data decreases, and the subsequent calculated target state vector cannot effectively reflect the attitude change of the earphone, making it difficult to achieve accurate perspective calibration and resulting in poor image calibration effects.

[0039] Therefore, in this embodiment, it is necessary to use the vector matrix after fusing the current visual and inertial sensing data to update the process noise covariance at the image acquisition time in real time, so as to process the visual and inertial sensing data at the next time based on the dynamically updated process noise covariance. In this way, by dynamically adjusting the process noise covariance matrix , respond to the error changes of the visual and IMU sensors in real time, and adaptively optimize the fusion parameters through the amplitude of the attitude deviation measured by the real-time visual and IMU, significantly enhancing the adaptive ability and robustness of the algorithm.

[0040] Specifically, when determining the target process covariance, it is necessary to first calculate the error change amount (matrix) in the process of fusing visual and IMU data, that is, the vector matrix difference between the target state vector and the inertial state vector, then calculate the product of the error change amount and the transposed matrix of the error change amount, and then use the sum of this product and the current process covariance at the image acquisition time as the target process noise covariance, so as to update the process noise covariance through the error change amount and its transpose of the predicted turntable vector and the corrected state vector, so as to achieve intelligent adaptive adjustment of visual and IMU data fusion in different environments.

[0041] Optionally, a weight ratio can also be set. When updating the process noise covariance of the noise process, set the weight ratio of the current process noise covariance, as well as the weight ratio of the product of the error change amount and its transpose, and control the balance between the historical information of the covariance matrix and the current error information based on this weight ratio, so as to improve the accuracy of subsequent image correction processing.

[0042] Step S30: Update the target state vector at the next moment corresponding to the image acquisition moment based on the target process noise covariance and a preset Kalman filtering algorithm.

[0043] In this embodiment, at each image acquisition moment, inertial sensing data and visual information of the image are fused through an extended Kalman filter. In the EKF framework, the purpose of calculating the target state vector is to use visual observation information to correct the predicted value of inertial sensing data. The fused attitude data, that is, the target state vector, is the result of image correction. During the data fusion and correction process, when predicting the inertial state (calculating the inertial state vector), the prediction state equation is directly affected by the process noise covariance, so that the result of the inertial state vector is affected by this covariance. At the same time, the dynamically updated process noise covariance affects the Kalman filter gain, so that the result of the target state vector is indirectly affected by the process noise covariance.

[0044] Therefore, after updating the current process noise covariance through the error change amount between the inertial state vector and the target state vector, in subsequent image acquisition moments, the subsequent image correction result can be adjusted through the updated target process noise covariance, and so on in a cycle. That is, at each image acquisition moment, the inertial state vector and the target state vector are calculated through the current process noise covariance (the covariance updated at the previous acquisition moment), and then the current process noise covariance is updated based on the inertial state vector and the target state vector at the current image acquisition moment to obtain the target process noise covariance used to calculate the inertial state vector and the target state vector at the next image acquisition moment, and calculate in this cycle. By dynamically adjusting the process noise covariance matrix, it can respond in real time to the error changes of the vision and inertial sensors, so as to utilize the attitude deviation amplitude of the real-time vision and inertial sensing data, adaptively optimize the fusion parameters, and enhance the adaptive ability and robustness of the EKF algorithm.

[0045] It should be noted that after calculating the fused attitude data, the earphone can perform image stitching and / or viewing angle correction processing based on the attitude data. The specific stitching and correction process and the algorithms used are not limited in this application.

[0046] This embodiment provides an image data processing method, which performs fusion correction processing on image visual data and inertial data through an extended Kalman filter algorithm, thereby correcting the predicted attitude value calculated based on inertial sensing data. Then, to reduce the interference of inertial sensing data fluctuations on the fusion accuracy during the data processing, the process noise covariance at the current moment is dynamically updated according to the error change amount between the inertial state vector and the target state vector, so that at the subsequent image acquisition moment, when performing data fusion correction processing based on the extended Kalman filter algorithm, the attitude fusion result is dynamically updated through the dynamically updated process noise covariance, thereby improving the image fusion accuracy when the sensor data quality fluctuates, and thus improving the image correction effect.

[0047] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as that in the above first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, the noise process covariance is updated by setting the weight ratio. Specifically, please refer to Figure 2 , step S20 includes steps S21 to S23: Step S21, determine the forgetting factor of the current process covariance and the transpose matrix of the error change amount, and the forgetting factor is the weight ratio of the current process covariance.

[0048] Step S22, determine the change weight of the first product of the error change amount and the transpose matrix according to the forgetting factor.

[0049] Step S23, set the sum of the change weight and the first product, and the second product of the forgetting factor and the current process noise covariance as the target process noise covariance.

[0050] In this embodiment, the update formula of the process noise covariance is as follows: , where, represents the covariance matrix dynamically adjusted at the current moment k, that is, the target process noise covariance, represents the covariance matrix at the previous moment, that is, the current process noise covariance, , is the corrected state after fusing the visual measurement data, that is, the target state vector, is the inertial state vector, is the error change amount during the visual and IMU fusion process, is the transpose matrix of the error change amount, is the forgetting factor (0 < λ < 1), which is equivalent to the weight ratio.

[0051] Therefore, when updating the current process state covariance, by introducing a forgetting factor, the balance between the historical information of the covariance matrix and the current error information is controlled, and the accuracy of the updated process noise covariance is improved. Herein, the forgetting factor can be a fixed value or a parameter that dynamically changes based on actual requirements, and the present application does not limit this.

[0052] This embodiment provides an image data processing method. By introducing a forgetting factor, the proportion of the current process noise covariance and the error variation when updating the process noise covariance is balanced, so that when the environment is stable and the error variation is small, it slowly changes and depends on the historical covariance information. While if the environment changes drastically and the sensor error increases, that is, the error variation is large, it updates quickly and depends more on the error information measured in real time, thereby realizing the intelligent adaptive adjustment of the process noise covariance in different environments when fusing visual and inertial sensing data.

[0053] Based on the first embodiment of the present application, in the third embodiment of the present application, the same or similar content as the above first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 3 , step S30 includes steps S31 to S33: Step S31, update the state covariance according to the target process noise covariance to obtain the target covariance.

[0054] In this embodiment, the target process noise covariance affects the state covariance, and the state covariance is used to calculate the Kalman gain coefficient, and the Kalman gain coefficient is used for the EKF correction value. Therefore, the target process noise covariance indirectly affects the result of the target state vector. Therefore, when it is necessary to update the target state vector at the next moment, it is necessary to first update the state covariance through the target process noise covariance.

[0055] Specifically, the calculation formula of the state covariance is as follows: , wherein, is the updated state covariance, that is, the target covariance, is the state covariance to be updated, A is the state transition matrix (pre-set), is the transpose of the state transition matrix, is the covariance matrix dynamically adjusted at the current moment k, that is, the target process noise covariance.

[0056] Step S32, update the Kalman gain coefficient at the next moment based on the target covariance, the preset observation noise covariance, and the Jacobian matrix of the visual measurement model and its transpose.

[0057] In this embodiment, the calculation formula of the Kalman gain coefficient is as follows: , wherein, is the Kalman gain coefficient, is the Jacobian matrix of the visual measurement model, is the transpose matrix of the Jacobian matrix, is the preset observation noise covariance.

[0058] Update the state covariance through the target process noise covariance, and update the Kalman filter coefficient based on the updated state covariance, so as to indirectly adjust the target state vector during subsequent EKF correction.

[0059] Step S33: Update the target state vector corresponding to the next moment of the image acquisition moment based on the Kalman gain coefficient at the next moment and the preset Kalman filter algorithm.

[0060] In this embodiment, the preset Kalman filter algorithm is the correction equation of the EKF algorithm, and the calculation formula is as follows: , wherein, is the corrected state after fusing visual measurement data, that is, the target state vector, is the inertial state vector, is the Kalman gain coefficient, is the visual observation value, that is, the image visual information, is the visual measurement model.

[0061] Furthermore, the calculation formula of the inertial state vector is: , A is the state transition matrix, B is the control input matrix, is the inertial sensing data, also known as the inertial input data, is the process noise, and its covariance is (target process noise covariance), is the best estimated state after fusing visual and inertial sensing data at the previous moment k - 1, that is, the state obtained after the state update (correction) of the filter at the previous moment, that is, the current inertial state vector calculated at the previous moment.

[0062] The calculation formula of the visual observation value is: , is the measurement noise, and its covariance is .

[0063] Therefore, when updating the target state vector at the next moment, based on the correction equation of the EKF, through the target process noise covariance the inertial state vector is directly updated, and at the same time, the Kalman gain coefficient is indirectly updated through the target process noise covariance, so as to indirectly update the target state vector. When performing data fusion and correction processing based on the extended Kalman filter algorithm at subsequent image acquisition moments, the attitude fusion result is updated in real time through the dynamically updated process noise covariance, thereby improving the image fusion accuracy when the sensor data quality fluctuates, and thus improving the image correction effect.

[0064] This embodiment provides an image data processing method. Based on the correction equation of the extended Kalman filter, the inertial state vector is directly updated through the dynamically updated process noise covariance matrix, and the Kalman filter gain is indirectly updated, so as to realize the dynamic update of the target state vector, so as to dynamically update the data at each acquisition moment during the image acquisition moment, reduce the sensor error, and improve the image correction effect.

[0065] Based on the first embodiment or the third embodiment of the present application, in the fourth embodiment of the present application, the same or similar content as the above embodiments can be referred to the above introduction and will not be repeated hereinafter. On this basis, the calculation formula of the inertial state vector is as shown in the calculation formula in the third embodiment. Therefore, when determining the inertial state vector, the preset state transition matrix, the current inertial state vector calculated at the previous moment corresponding to the image acquisition moment, and the control input matrix can be determined first, and then the final inertial state vector can be calculated according to the state transition matrix, the current inertial state vector, the control input matrix, the inertial sensing data, and the current process noise covariance.

[0066] Specifically, the product of the state transition matrix and the current inertial state vector, and the product of the control input matrix and the inertial sensing data can be determined first, and then the sum of these two products and the current process noise covariance is set as the inertial state vector.

[0067] The calculation formula of the inertial state vector is as follows: , where the current image acquisition moment is K, is the inertial state vector, A is the preset state transition matrix, B is the control input matrix, is the inertial sensing data, also known as the inertial input data, is the process noise, and its covariance is , is the best estimated state after fusing vision and inertial sensing data at the previous moment k - 1, that is, the state obtained after the state update (correction) of the filter at the previous moment, that is, the current inertial state vector calculated at the previous moment.

[0068] Further, when determining the target state vector, the Jacobian matrix of the visual measurement model and its transpose matrix, a preset observation noise covariance, and a state covariance determined by the current process noise covariance can be obtained first. Then, the Kalman gain coefficient is determined by the Jacobian matrix and its transpose matrix, the preset observation noise covariance, and the state covariance. Subsequently, the Kalman filter fusion result of the visual attitude information, the Kalman gain coefficient, and the inertial sensing vector is calculated according to the Kalman gain coefficient and the preset Kalman filter algorithm to obtain the target state vector.

[0069] The preset Kalman filter algorithm is the EKF correction equation, and its calculation formula is as follows: , where, is the target state vector, is the inertial state vector, is the Kalman filter gain, is the visual observation value, that is, the image visual information, is the visual measurement model.

[0070] Exemplarily, in the fused target state vector, , where the horizontal attitude angle is: , and the vertical attitude angle is . Therefore, the attitude angle output after fusion is the first two elements in the target state vector .

[0071] It should be noted that in the above calculation formula, when calculating the target state vector at the current image acquisition moment, the current process noise covariance is used, while when calculating the target state vector at the next moment, the target process noise covariance is used. If the previous moment is empty, the current process noise covariance is the initial covariance. Therefore, different process noise covariances are used when calculating the target state vectors at different image acquisition moments.

[0072] Based on the first embodiment of the present application, in the fifth embodiment of the present application, the same or similar content as that in the above embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, before step S10, steps S01 to S03 are further included: Step S01, obtaining the image information collected at the image acquisition moment and determining the target area of the image information.

[0073] The target area is the area corresponding to the visual key feature points. In this embodiment, when collecting images based on the headset, the cheek area and the temple position area on the side view of the headset can be used as the target area. Among them, these areas have stable structures and are not easily blocked. At the same time, the side view is easier to capture in real time and is more suitable for the perspective correction task of the head-mounted device, which can significantly improve the stability and reliability of visual measurement. Therefore, the cheek area and the temple position area on the side view of the headset are used as the target area.

[0074] Step S02: Obtain the image coordinate information of the visual feature points in the target area, and convert the image coordinate information into three-dimensional space coordinate information.

[0075] In this embodiment, the extraction of image coordinates and the conversion of coordinates can be performed through a neural network model.

[0076] Exemplarily, the target area is the cheek area and the temple position area. As an alternative implementation, a lightweight face key point detection network (MediaPipe Face Mesh) can be used to detect the two-dimensional coordinates of the side face cheek and the temple in real time. The two-dimensional coordinates are expressed as: , represents the two-dimensional coordinates of the cheek feature points in the image, : The two-dimensional coordinates of the temple feature points in the image.

[0077] Then, three-dimensional space positioning can be performed based on stereo vision, including using the parallax principle of the binocular cameras of the headset for feature point matching and three-dimensional reconstruction to obtain the real space coordinates. Among them, the three-dimensional coordinates are expressed as:

[0078] represents the coordinates of the feature points in the real three-dimensional space.

[0079] Step S03: Generate the image visual information according to the observation result of the three-dimensional space coordinate information under the visual measurement model and the measurement noise of the image coordinate information.

[0080] In this embodiment, the image visual information is , that is, the visual observation value, and its calculation formula is: , is the three-dimensional space coordinate information, is the measurement noise, and its covariance is , is the visual observation model.

[0081] This embodiment provides an image data processing method, which uses the cheek area and the temple position from the perspective of the earphone side as visual key feature points, and calculates the actual image visual information based on the two-dimensional coordinate information of these key feature points. In this way, during the process of wearing and shooting the earphone, analysis and calculation are carried out based on the stable and non-blocking feature area, reducing the calculation complexity and improving the stability and reliability of visual measurement.

[0082] Based on the first embodiment of this application, in the sixth embodiment of this application, the same or similar content as the above first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, after step S30, steps S40 to S50 are further included: Step S40, output a calibration prompt message according to the target state vector, and obtain the corrected posture corresponding to the calibration prompt message.

[0083] In this embodiment, the earphone can provide clear voice instructions according to the fused posture data to guide the user to precisely adjust the posture. For example, the earphone generates calibration prompt messages such as "Please adjust slightly to the left until the temple position is calibrated" and "Slightly down, align the cheek position" according to the target state vector, and then obtains the corrected posture feedback by the user based on this calibration prompt message, that is, performs real-time posture detection to automatically trigger the calibration prompt and reduce the posture correction error.

[0084] Step S50, if the corrected posture meets the posture result corresponding to the target state vector, output a corrected prompt.

[0085] In this embodiment, when the user adjusts the shooting posture and the adjusted posture enters the allowable error range, it means that the current posture meets the shooting conditions. At this time, a corrected prompt is output so that the user can learn the prompt for completing the posture correction.

[0086] Exemplarily, in order to help understand the implementation process of the image data processing method obtained by combining the above various embodiments of this application, please refer to Figure 4 , Figure 4 A brief flow schematic diagram of an image data processing method is provided. Specifically, at the moment of image acquisition, the dual cameras of the earphone collect image data, and at the same time, inertial sensing data is collected in real time through the inertial sensor. Then, feature point detection is performed on the positions of the side face cheek and the temple in the image data, and the visual posture value is calculated. At the same time, the posture prediction value of the inertial sensing data, that is, the inertial state vector, is calculated. Then, the correction value for correcting this posture prediction value is calculated, that is, the posture prediction value is corrected by the visual posture value to obtain the target state vector. Subsequently, the posture error estimation between the two state vectors is calculated, and the process noise covariance is dynamically updated through the posture error estimation, so as to update the posture data of the next time period through the process noise covariance.

[0087] Then, real-time voice feedback is performed through the fused accurate attitude data so that the user can correct the attitude. Subsequently, the earphone automatically synchronizes image stitching and corrected perspective, thereby completing image correction. Finally, the attitude deviation change of the earphone is continuously detected. If the deviation is too large, the attitude correction is triggered, and at this time, it jumps to the process of collecting data by the camera and inertial sensor. Otherwise, the current attitude fusion output is maintained, and data collection at the next moment is performed.

[0088] This embodiment provides an image data processing method. After detecting the change in the fused attitude of vision and IMU, it can adaptively trigger a calibration program to prompt the user to correct the attitude, thereby improving the image calibration effect and enhancing the user experience at the same time.

[0089] This application provides an image acquisition device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the image data processing method in the above first embodiment.

[0090] Next, refer to Figure 5 , which shows a schematic structural diagram of an image acquisition device suitable for implementing the embodiments of the present application. Figure 5 The shown image acquisition device is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0091] As Figure 5As shown, the image acquisition device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM, Read Only Memory) 1002 or the program loaded from the storage device 1003 into the random access memory (RAM, Random Access Memory) 1004. In the random access memory 1004, various programs and data required for the operation of the image acquisition device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD, Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the image acquisition device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an image acquisition device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be alternatively implemented or had.

[0092] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the read-only memory 1002. When the computer program is executed by the processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.

[0093] The image acquisition device provided by the present application adopts the image data processing method in the above embodiments and can solve the technical problem of poor image correction effect. Compared with the prior art, the beneficial effects of the image acquisition device provided by the present application are the same as those of the image data processing method provided by the above embodiments, and the other technical features in the image acquisition device are the same as those disclosed in the method of the previous embodiment, which will not be elaborated here.

[0094] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0095] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0096] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the image data processing method in the above embodiments.

[0097] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination of the above.

[0098] The above computer-readable storage medium can be included in the image acquisition device; or it can exist separately without being assembled into the image acquisition device.

[0099] The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by the image acquisition device, the image acquisition device is caused to: determine the inertial state vector of the image acquisition device at the image acquisition moment according to the inertial sensing data, and determine the target state vector according to the image visual information and the inertial state vector; Update the current process noise covariance according to the error variation between the inertial state vector and the target state vector to obtain the target process noise covariance; Update the target state vector at the next moment corresponding to the image acquisition moment based on the target process noise covariance and a preset Kalman filtering algorithm.

[0100] Computer program code for performing the operations of the present application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0101] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0102] The modules involved in the embodiments described in the present application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.

[0103] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above image data processing method, which can solve the technical problem of poor image correction effect. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the image data processing method provided by the above embodiments, and will not be elaborated here.

[0104] The above are only partial embodiments of this application, and do not limit the patent scope of this application accordingly. Any equivalent structural transformation made under the technical concept of this application by using the content of the specification and drawings of this application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of this application.

Claims

1. An image data processing method, characterized in that, The described image data processing method includes: Determining an inertial state vector of an image acquisition device at an image acquisition moment according to inertial sensing data, and determining a target state vector according to image visual information and the inertial state vector; Updating a current process noise covariance according to an error change amount between the inertial state vector and the target state vector to obtain a target process noise covariance; Updating the target state vector at the next moment corresponding to the image acquisition moment based on the target process noise covariance and a preset Kalman filtering algorithm.

2. The image data processing method according to claim 1, wherein The step of updating a current process noise covariance according to an error change amount between the inertial state vector and the target state vector to obtain a target process noise covariance includes: Determining a forgetting factor of the current process covariance and a transpose matrix of the error change amount, where the forgetting factor is a weight ratio of the current process covariance; Determining a change weight of a first product of the error change amount and the transpose matrix according to the forgetting factor; Setting a sum of the change weight and the first product, and a second product of the forgetting factor and the current process noise covariance as the target process noise covariance.

3. The image data processing method according to claim 1, wherein The step of updating the target state vector at the next moment corresponding to the image acquisition moment based on the target process noise covariance and a preset Kalman filtering algorithm includes: Updating a state covariance according to the target process noise covariance to obtain a target covariance; Updating a Kalman gain coefficient at the next moment based on the target covariance, a preset observation noise covariance, and a Jacobian matrix of a visual measurement model and its transpose; Updating the target state vector at the next moment corresponding to the image acquisition moment based on the Kalman gain coefficient at the next moment and the preset Kalman filtering algorithm.

4. The image data processing method according to claim 1, wherein, The step of determining an inertial state vector of an image acquisition device at an image acquisition moment according to inertial sensing data, and determining a target state vector according to image visual information and the inertial state vector includes: Determining a preset state transition matrix, a current inertial state vector calculated at the previous moment corresponding to the image acquisition moment, and a control input matrix; Calculating the inertial state vector according to the state transition matrix, the current inertial state vector, the control input matrix, the inertial sensing data, and the current process noise covariance; Determining a visual observation value according to the image visual information, and calculating the target state vector based on the preset Kalman filtering algorithm and the inertial state vector.

5. The image data processing method according to claim 4, wherein The step of calculating the inertial state vector according to the state transition matrix, the current inertial state vector, the control input matrix, the inertial sensing data, and the current process noise covariance includes: Determining a third product of the state transition matrix and the current inertial state vector, and a fourth product of the control input matrix and the inertial sensing data; Setting a sum of the third product, the fourth product, and the current process noise covariance as the inertial state vector.

6. The image data processing method according to claim 4, wherein Before the step of determining the target state vector according to the visual pose information and the inertial state vector, the image data processing method further includes: Obtain the Jacobian matrix of the visual measurement model and its transpose matrix, a preset observation noise covariance, and determine the state covariance at the image acquisition moment according to the current process noise covariance; Determine the Kalman gain coefficient at the image acquisition moment according to the Jacobian matrix and its transpose matrix, the preset observation noise covariance, and the state covariance; Calculate the Kalman filter fusion result of the visual pose information, the Kalman gain coefficient, and the inertial state vector according to the Kalman gain coefficient at the image acquisition moment and the preset Kalman filter algorithm to obtain the target state vector.

7. The image data processing method according to claim 1, wherein Before the steps of determining the inertial state vector of the image acquisition device at the image acquisition moment according to the inertial sensing data, and determining the target state vector according to the image visual information and the inertial state vector, the image data processing method further includes: Obtain the image information acquired at the image acquisition moment and determine the target area of the image information; Obtain the image coordinate information of the visual feature points in the target area and convert the image coordinate information into three-dimensional space coordinate information; Generate the image visual information according to the observation result of the three-dimensional space coordinate information under the visual measurement model and the measurement noise of the image coordinate information.

8. The image data processing method according to claim 1, wherein After the step of updating the target state vector at the next moment corresponding to the image acquisition moment based on the target process noise covariance and the preset Kalman filter algorithm, the image data processing method further includes: Output a calibration prompt message according to the target state vector and obtain the correction pose corresponding to the calibration prompt message; If the correction pose meets the pose result corresponding to the target state vector, output a corrected prompt.

9. An image acquisition device, characterized in that, The image acquisition device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the image data processing method according to any one of claims 1 to 8.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the image data processing method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Eye movement tracking method, intelligent glasses, mobile terminal and eye movement tracking system

    CN120891931A

  • Moving object three-dimensional model reconstruction system and method based on cooperation of multiple unmanned aerial vehicles

    CN121414977A

  • Multi-modal visual positioning method and device and electronic equipment

    CN121767453A

  • Multimodal visual positioning method, apparatus, and electronic device

    CN121767453B