An image processing method, apparatus and device

CN122820852APending Publication Date: 2026-09-25DEXFORCE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611317073.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-28
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]本发明提供了一种图像处理方法、装置及设备,以解决现有姿态估计方法高度依赖视差图的质量,易在视差图出现偏差时导致视差误差传导累积,使得姿态估计精度与鲁棒性降低的问题

Benefits of technology

[0020]本发明实施例提供的图像处理方法,通过获取通过双目视觉传感器针对目标对象同步采集的左目图像与右目图像,并确定与左目图像关联的语义特征和左目视图匹配特征、以及与右目图像关联的右目视图匹配特征;根据左目视图匹配特征和右目视图匹配特征构建双目代价体,双目代价体用于表征左目图像像素与右目图像像素在各预设候选视差下的匹配关联关系;根据双目代价体和语义特征,获得目标对象对应的二维任务预测结果,并对双目代价体进行三维正则化处理,获得初始双目视差估计结果;其中,根据双目代价体和语义特征,获得目标对象对应的二维任务预测结果,包括:对双目代价体进行特征映射得到对应的双目几何特征,根据双目几何特征与语义特征的融合特征进行二维任务推理得到二维任务预测结果;根据语义特征、双目代价体、二维任务预测结果以及初始双目视差估计结果,确定目标对象的目标姿态估计结果。利用该方法,通过在统一网络结构中联合建模二维视觉任务与三维空间任务,实现对目标六维姿态的直接预测,并且通过双目特征级信息融合策略,在不依赖高精度双目视差显式预测的情况下,充分利用双目输入图像所蕴含的几何与语义信息,从而有效削弱传统方法对双目视差质量的依赖,并整体提升了六维姿态估计的准确性、鲁棒性及工程适用性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820852A_ABST
    Figure CN122820852A_ABST
Patent Text Reader

Abstract

The application discloses an image processing method, device and equipment, and relates to the technical field of image processing. The method comprises the following steps: acquiring a left-eye image and a right-eye image which are synchronously collected by a binocular vision sensor for a target object, and determining semantic features associated with the left-eye image and left-eye view matching features, and right-eye view matching features associated with the right-eye image; constructing a binocular cost volume according to the left-eye view matching features and the right-eye view matching features; obtaining a two-dimensional task prediction result corresponding to the target object according to the binocular cost volume and the semantic features, and performing three-dimensional regularization processing on the binocular cost volume to obtain an initial binocular disparity estimation result; and determining a target pose estimation result of the target object according to the semantic features, the binocular cost volume, the two-dimensional task prediction result and the initial binocular disparity estimation result. By using the method, the dependence on the binocular disparity quality is reduced, and the accuracy, robustness and engineering applicability of six-dimensional pose estimation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to an image processing method, apparatus, and device. Background Technology

[0002] Six-dimensional pose estimation aims to simultaneously recover the three-dimensional position and spatial pose of a target in three-dimensional space from visual input. It is usually represented in the form of six degrees of freedom and is widely used in fields such as robot grasping, autonomous driving, augmented reality and industrial inspection.

[0003] In binocular vision systems, the parallax relationship formed by the left and right cameras can alleviate the uncertainty of scale and depth estimation in monocular vision to a certain extent. Therefore, binocular vision is widely used in high-precision 6D pose estimation tasks.

[0004] However, most existing binocular six-dimensional attitude estimation methods rely heavily on disparity maps corresponding to binocular images to complete attitude calculation. The accuracy and stability of attitude estimation are entirely dependent on the quality of disparity map solving, failing to fully explore and utilize the effective feature information of binocular images. In actual complex working conditions, once the disparity estimation result deviates, the error will be directly transmitted and accumulated to the subsequent attitude calculation stage, resulting in poor stability and low accuracy of attitude estimation results, which cannot meet the industrial requirements for high-precision and high-robust attitude estimation. Summary of the Invention

[0005] This invention provides an image processing method, apparatus, and device to address the problem that existing attitude estimation methods are highly dependent on the quality of disparity maps, and that disparity error propagation and accumulation are easily caused by deviations in the disparity map, resulting in reduced attitude estimation accuracy and robustness.

[0006] In a first aspect, embodiments of the present invention provide an image processing method, the method comprising:

[0007] Acquire the left and right eye images of the target object simultaneously captured by the binocular vision sensor, and determine the semantic features and left eye view matching features associated with the left eye image, as well as the right eye view matching features associated with the right eye image;

[0008] A binocular cost body is constructed based on the left-eye view matching features and the right-eye view matching features. The binocular cost body is used to characterize the matching association relationship between left-eye image pixels and right-eye image pixels under each preset candidate disparity.

[0009] Based on the binocular cost body and the semantic features, a two-dimensional task prediction result corresponding to the target object is obtained, and the binocular cost body is subjected to three-dimensional regularization processing to obtain an initial binocular disparity estimation result; wherein, obtaining the two-dimensional task prediction result corresponding to the target object based on the binocular cost body and the semantic features includes: performing feature mapping on the binocular cost body to obtain corresponding binocular geometric features, and performing two-dimensional task inference based on the fusion features of the binocular geometric features and the semantic features to obtain the two-dimensional task prediction result;

[0010] Based on the semantic features, the binocular cost volume, the two-dimensional task prediction results, and the initial binocular disparity estimation results, the target pose estimation result of the target object is determined.

[0011] In a second aspect, embodiments of the present invention provide an image processing apparatus, the apparatus comprising:

[0012] The feature extraction module is used to acquire the left and right eye images of the target object simultaneously collected by the binocular vision sensor, and to determine the semantic features and left eye view matching features associated with the left eye image, and the right eye view matching features associated with the right eye image.

[0013] The cost body construction module is used to construct a binocular cost body based on the left-eye view matching features and the right-eye view matching features. The binocular cost body is used to characterize the matching association relationship between left-eye image pixels and right-eye image pixels under each preset candidate disparity.

[0014] The first estimation module is used to obtain a two-dimensional task prediction result corresponding to the target object based on the stereo cost body and the semantic features, and to perform three-dimensional regularization processing on the stereo cost body to obtain an initial stereo disparity estimation result; wherein, obtaining the two-dimensional task prediction result corresponding to the target object based on the stereo cost body and the semantic features includes: performing feature mapping on the stereo cost body to obtain corresponding stereo geometric features, and performing two-dimensional task inference based on the fusion features of the stereo geometric features and the semantic features to obtain the two-dimensional task prediction result;

[0015] The second estimation module is used to determine the target pose estimation result of the target object based on the semantic features, the binocular cost volume, the two-dimensional task prediction result, and the initial binocular disparity estimation result.

[0016] Thirdly, embodiments of the present invention provide an electronic device, the electronic device comprising:

[0017] At least one processor;

[0018] and a memory communicatively connected to the at least one processor;

[0019] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the image processing method according to any embodiment of the present invention.

[0020] The image processing method provided in this invention acquires left and right eye images of a target object simultaneously captured by a binocular vision sensor, and determines semantic features and left-eye view matching features associated with the left eye image, as well as right-eye view matching features associated with the right eye image. A binocular cost body is constructed based on the left-eye view matching features and the right-eye view matching features. The binocular cost body characterizes the matching relationship between pixels in the left and right eye images under various preset candidate disparities. Based on the binocular cost body and semantic features, a two-dimensional task prediction result corresponding to the target object is obtained, and the binocular cost body is subjected to three-dimensional regularization processing to obtain an initial binocular disparity estimation result. The process of obtaining the two-dimensional task prediction result corresponding to the target object based on the binocular cost body and semantic features includes: performing feature mapping on the binocular cost body to obtain corresponding binocular geometric features; performing two-dimensional task inference based on the fusion features of the binocular geometric features and semantic features to obtain the two-dimensional task prediction result; and determining the target pose estimation result of the target object based on the semantic features, the binocular cost body, the two-dimensional task prediction result, and the initial binocular disparity estimation result. This method enables direct prediction of target six-dimensional pose by jointly modeling two-dimensional vision tasks and three-dimensional spatial tasks in a unified network structure. Furthermore, through a binocular feature-level information fusion strategy, it fully utilizes the geometric and semantic information contained in the binocular input images without relying on high-precision binocular disparity explicit prediction. This effectively reduces the dependence of traditional methods on binocular disparity quality and improves the accuracy, robustness, and engineering applicability of six-dimensional pose estimation as a whole.

[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 A flowchart of an image processing method provided in an embodiment of the present invention;

[0024] Figure 2 This is a schematic diagram illustrating the principle of an image processing method provided in an embodiment of the present invention;

[0025] Figure 3 This is a schematic diagram of the structure of an image processing device provided in an embodiment of the present invention;

[0026] Figure 4 A schematic diagram of an electronic device that can be used to implement embodiments of the present invention is shown. Detailed Implementation

[0027] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0029] It is understood that before using the technical methods disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0030] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as the electronic device, application, server, or storage medium performing the operations of this disclosed technology.

[0031] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0032] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0033] Specifically, embodiments of the present invention provide an image processing method. Figure 1 This is a flowchart of an image processing method provided by an embodiment of the present invention. The embodiment of the present invention is applicable to scenarios involving pose estimation of a target object. The method can be executed by an image processing device, which can be implemented in the form of software and / or hardware. Optionally, it can be implemented by an electronic device, preferably a mobile terminal, desktop computer, laptop computer, or server.

[0034] like Figure 1 As shown, the image processing method provided in this embodiment of the invention may specifically include:

[0035] S101. Acquire the left and right eye images synchronously acquired by the binocular vision sensor for the target object, and determine the semantic features and left eye view matching features associated with the left eye image, as well as the right eye view matching features associated with the right eye image.

[0036] In this context, a binocular vision sensor can be understood as an imaging acquisition device capable of simultaneously acquiring images from two perspectives within the same scene. Examples include binocular cameras, stereo cameras, or a combination of two calibrated, time-synchronized ordinary monocular cameras. The device undergoes pre-calibration of its intrinsic and extrinsic parameters and can output paired left and right eye images. The left and right eye images can be understood as a pair of images simultaneously acquired by the binocular vision sensor. For example, the left eye image is the image output by the imaging unit on the left side of the sensor, and the right eye image is the image output by the imaging unit on the right side of the sensor. Both images capture a target object located in the same scene, and there is parallax between the images. The target object can be considered the object to be measured for 6D pose estimation, which can be an industrial workpiece, an object grasped by a robot, or a real-world entity. One or more target objects can be deployed in the shooting scene as needed.

[0037] In this embodiment, the method of acquiring the left and right eye images synchronously acquired by the binocular vision sensor for the target object may include: receiving the paired left and right eye images acquired and output in real time by the binocular vision sensor for the target object, or reading the left and right eye images synchronously acquired and stored by the binocular vision sensor from a preset storage location, etc., without limitation. Next, semantic feature extraction and matching feature extraction can be performed on the left eye image to obtain corresponding multi-scale semantic features and left eye view matching features, and matching feature extraction can be performed on the right eye image to obtain corresponding right eye view matching features. The semantic features can be used to carry semantic information such as the category, region, outline, meaning, concept, and context of the target object; the left eye view matching features and right eye view matching features can be considered as matching features with geometrical alignment in the left and right eye views, respectively.

[0038] S102. Construct a binocular cost body based on the matching features of the left and right views. The binocular cost body is used to characterize the matching relationship between the pixels of the left and right images under each preset candidate disparity.

[0039] The binocular cost volume can be understood as a matching relationship tensor formed by the left-eye view matching features and the right-eye view matching features under preset candidate disparities. It is used to characterize the degree of matching association between each pixel position in the left-eye image and the corresponding offset pixel in the right-eye image under each preset candidate disparity. The candidate disparity can be considered as a disparity hypothesis value determined in advance within a set range, representing the horizontal offset of the pixel position of the left-eye image pixel in the right-eye image. Each candidate disparity corresponds to one horizontal offset of the right-eye matching feature.

[0040] In this embodiment, based on the left and right sets of view matching features, the feature matching degree of each pixel under each candidate disparity can be calculated by traversing each candidate disparity, and integrated to form a stereo cost body, establishing the correspondence between pixels, candidate disparities, and feature matching strength. As one possible approach, a set of candidate disparities can be pre-defined; for each candidate disparity, the right-eye view matching features are offset according to the horizontal offset corresponding to that candidate disparity; the view matching features at each pixel position in the left-eye image are matched with the view matching features at the same coordinate position in the offset right-eye image to obtain the feature matching strength; all pixels and all candidate disparities are traversed, and all feature matching strengths are integrated according to a set dimension (such as pixel dimension, candidate disparity dimension) to form a stereo cost body.

[0041] S103. Based on the binocular cost body and semantic features, obtain the two-dimensional task prediction result corresponding to the target object, and perform three-dimensional regularization processing on the binocular cost body to obtain the initial binocular disparity estimation result; wherein, obtaining the two-dimensional task prediction result corresponding to the target object based on the binocular cost body and semantic features includes: performing feature mapping on the binocular cost body to obtain the corresponding binocular geometric features, and performing two-dimensional task inference based on the fusion features of the binocular geometric features and semantic features to obtain the two-dimensional task prediction result.

[0042] The 2D task prediction result can be understood as the 2D output result corresponding to the preset 2D task. For example, the 2D task may include object detection, object segmentation, and 2D keypoint prediction. Correspondingly, the 2D task prediction result may include the 2D bounding box of the target object, pixel-level segmentation results, and 2D keypoint prediction results (such as the predicted 2D keypoint positions). The initial binocular disparity estimation result can be considered a coarse-grained binocular disparity estimation result, which has problems such as blurred edges and insufficient accuracy for small objects, and requires refinement.

[0043] In this embodiment, this step includes two parallel processing branches. The first branch fuses the binocular geometric information carried by the binocular cost body with the semantic features of the left-eye image to form a unified feature representation containing binocular geometric constraints. Based on this feature representation, two-dimensional task inference is performed to obtain the two-dimensional task prediction result corresponding to the target object. The second branch performs three-dimensional regularization processing on the binocular cost body to suppress noise within the binocular cost body, correct abnormal matching strength caused by mismatches, and output the initial binocular disparity estimation result.

[0044] S104. Based on semantic features, binocular cost volume, two-dimensional task prediction results, and initial binocular disparity estimation results, determine the target pose estimation result of the target object.

[0045] Among them, the target pose estimation result can be understood as the final predicted six-dimensional pose result, which represents the three-dimensional position and spatial rotational attitude of the target object in the three-dimensional coordinate system of the binocular vision sensor.

[0046] In this embodiment, the initial pose estimation result is obtained based on the two-dimensional task prediction result; then, the initial binocular disparity estimation result is refined by combining semantic features and binocular cost volume to obtain the target binocular disparity estimation result; the target binocular disparity estimation result is combined with the intrinsic parameters of the binocular vision sensor to construct a three-dimensional point cloud representation; finally, the three-dimensional point cloud representation is used as a geometric constraint to perform registration optimization on the initial pose estimation result to obtain the final target pose estimation result.

[0047] The image processing method provided in this invention acquires left and right eye images of a target object simultaneously captured by a binocular vision sensor, and determines semantic features and left-eye view matching features associated with the left eye image, as well as right-eye view matching features associated with the right eye image. A binocular cost body is constructed based on the left-eye view matching features and the right-eye view matching features. The binocular cost body characterizes the matching relationship between pixels in the left and right eye images under various preset candidate disparities. Based on the binocular cost body and semantic features, a two-dimensional task prediction result corresponding to the target object is obtained, and the binocular cost body is subjected to three-dimensional regularization processing to obtain an initial binocular disparity estimation result. The process of obtaining the two-dimensional task prediction result corresponding to the target object based on the binocular cost body and semantic features includes: performing feature mapping on the binocular cost body to obtain corresponding binocular geometric features; performing two-dimensional task inference based on the fusion features of the binocular geometric features and semantic features to obtain the two-dimensional task prediction result; and determining the target pose estimation result of the target object based on the semantic features, the binocular cost body, the two-dimensional task prediction result, and the initial binocular disparity estimation result. This method enables direct prediction of target six-dimensional pose by jointly modeling two-dimensional vision tasks and three-dimensional spatial tasks in a unified network structure. Furthermore, through a binocular feature-level information fusion strategy, it fully utilizes the geometric and semantic information contained in the binocular input images without relying on high-precision binocular disparity explicit prediction. This effectively reduces the dependence of traditional methods on binocular disparity quality and improves the accuracy, robustness, and engineering applicability of six-dimensional pose estimation as a whole.

[0048] Specifically, this method no longer relies solely on a pre-estimated single high-precision disparity feature. Instead, it constructs a binocular cost body based on the feature information of the binocular input image, fully preserving the feature matching and association information of each pixel under all candidate disparity assumptions. Pose estimation is then performed based on the binocular cost body, combined with semantic features, 2D task prediction results, and initial binocular disparity estimation results. This effectively suppresses error accumulation caused by disparity mismatches and takes into account both 2D semantic priors and 3D geometric constraints. This improves the accuracy and robustness of target pose estimation results in complex scenarios such as low binocular disparity quality, difficult stereo matching, or the presence of occlusion and weak textures, thus meeting the needs of practical applications. Furthermore, it can simultaneously complete 2D tasks (including but not limited to target detection, target segmentation, and 2D keypoint prediction) and 3D tasks (including binocular geometric information modeling and 6D pose estimation) during a single forward inference process, improving inference efficiency and resource conservation. This makes it suitable for scenarios with higher requirements for multi-task collaboration and computational efficiency, such as humanoid robots and augmented reality.

[0049] As a first optional embodiment of the present invention, based on the above embodiments, the steps of determining the semantic features and left-eye view matching features associated with the left-eye image, and the right-eye view matching features associated with the right-eye image, can be specifically optimized as follows:

[0050] The left eye image is input into the semantic feature extraction network of the feature extraction module in the image processing model to extract semantic features from the left eye image and obtain the semantic features corresponding to the left eye image.

[0051] The left and right images are respectively input into the matching feature extraction network of the feature extraction module. The left image is used to extract matching features to obtain left view matching features, and the right image is used to extract matching features to obtain right view matching features.

[0052] In this embodiment, the pose estimation of the target object can be considered to be mainly achieved through an image processing model. The image processing model may include modules for performing different estimation processing logics. The feature extraction module can be viewed as a pre-trained module for feature extraction built into the image processing model, which may contain a semantic feature extraction network and a matching feature extraction network.

[0053] In this embodiment, the left and right eye images are used as inputs to the image processing model. The left eye image is input to a semantic feature extraction network, which extracts the semantic features required for object detection, object segmentation, and 2D keypoint prediction for 2D tasks. The left and right eye images are simultaneously input to a matching feature extraction network to extract matching features with geometrical alignment between the left and right views, resulting in left eye view matching features and right eye view matching features to support subsequent binocular geometric modeling.

[0054] The above technical solution uses two independent networks in the feature extraction module to complete semantic feature extraction and matching feature extraction respectively, providing data support for subsequent two-dimensional task fusion reasoning and binocular cost volume construction.

[0055] As a second optional embodiment of the present invention, based on the above embodiments, one implementation of constructing a binocular cost body according to the left eye view matching features and the right eye view matching features can be described as follows: through the binocular cost body construction module in the image processing model, the feature matching relationship between the left eye view matching features and the right eye view matching features is determined according to each preset candidate disparity, and the binocular cost body is formed based on the feature matching relationship.

[0056] As one implementation, the method of determining the feature matching relationship between the left-eye view matching feature and the right-eye view matching feature based on each preset candidate disparity, and forming the binocular cost body based on the feature matching relationship, can be described as follows: Obtain each preset candidate disparity; for each candidate disparity, perform a horizontal offset on the right-eye view matching feature according to the offset corresponding to the candidate disparity to obtain the offset right-eye view matching feature under the candidate disparity; for each pixel position in the left-eye view matching feature, perform a matching operation between the first feature vector and the second feature vector corresponding to the pixel position in the offset right-eye view matching feature to obtain the feature matching intensity of the pixel position under the candidate disparity; after traversing all pixel positions and all candidate disparities, integrate all feature matching intensities according to a preset dimension to form a multi-dimensional binocular cost body.

[0057] The first feature vector can be understood as the feature vector corresponding to a certain pixel position on the left eye view matching feature; the second feature vector can be understood as the feature vector corresponding to the same spatial pixel position on the right eye view matching feature after candidate disparity horizontal offset processing.

[0058] In this embodiment, each preset candidate disparity can be determined through discrete search or other methods. The binocular cost volume can be considered as the matching relationship volume of the left-eye view matching features and the right-eye view matching features under each candidate disparity, that is, the matching relationship tensor on "pixel position × candidate disparity". It can be in matrix form, used to record the feature matching strength / feature relationship of each pixel position under each candidate disparity, and to represent the geometric correspondence between the left-eye image and the right-eye image. Optionally, it can be integrated according to the height dimension, width dimension, channel dimension and / or candidate disparity dimension.

[0059] The above technical solution provides a specific implementation method for constructing a stereo cost volume. By storing the feature matching intensity of each pixel position under all candidate disparities in the stereo cost volume, it provides data support for subsequent pose estimation, reduces the dependence of pose estimation on high-precision disparity features or disparity prediction results (such as disparity maps), and avoids error accumulation.

[0060] As a third optional embodiment of the present invention, based on the above embodiments, the steps of performing feature mapping on the stereo cost body to obtain corresponding stereo geometric features, and performing two-dimensional task inference based on the fusion features of the stereo geometric features and the semantic features to obtain two-dimensional task prediction results are specified as follows:

[0061] The stereo cost volume is input into the feature mapping network of the geometric feature injection module in the image processing model, and the stereo cost volume is subjected to feature mapping to obtain the corresponding stereo geometric features. The stereo geometric features are matched with the spatial resolution and feature dimension of the semantic features.

[0062] The binocular geometric features and the semantic features are input into the feature fusion network of the geometric feature injection module, and the binocular geometric features and the semantic features are fused to obtain fused features.

[0063] The fused features are input into the two-dimensional task inference module of the image processing model to obtain the two-dimensional task prediction result output after performing two-dimensional task inference.

[0064] In this embodiment, the geometric feature injection module can be understood as a pre-trained functional module within the image processing model, used to inject binocular geometric features into the two-dimensional task feature space. The geometric feature injection module includes a feature mapping network and a feature fusion network. To improve the prediction performance of the two-dimensional task using binocular geometric information, the feature mapping network can project the binocular cost volume onto a feature space that matches the two-dimensional task features in terms of spatial resolution and semantic level, obtaining the projected binocular geometric features. The feature mapping network can be a two-dimensional convolutional network. Next, the feature fusion network can fuse the binocular geometric features with the two-dimensional task features (i.e., semantic features) to form a unified feature representation containing binocular geometric constraints, i.e., fused features, which are then used as input to the two-dimensional task inference module.

[0065] For example, one implementation of feature fusion can be described as follows: concatenating binocular geometric features and semantic features by channel to obtain concatenated features, then performing a convolution operation on the concatenated features using a convolution kernel of size 1, and reprojecting the number of channels of the concatenated features back to the original number of channels to obtain fused features.

[0066] In this embodiment, the fused features are input into a two-dimensional task inference module. This module may include multiple parallel-executing two-dimensional task inference networks, each responsible for outputting the target's two-dimensional bounding box, pixel-level segmentation results, and two-dimensional keypoint locations. It is understood that each two-dimensional task inference network in the module shares a unified semantic feature extraction network, which is optimized through multi-task joint training.

[0067] Through the above technical solution, the binocular cost volume enables binocular geometric information to participate in the two-dimensional task inference process in the form of features. By using the binocular cost volume instead of monocular disparity matching features, the binocular cost volume integrates binocular key information. Integrating the binocular cost volume into the two-dimensional task inference module can provide a larger field of view for two-dimensional target detection, segmentation and key point prediction. Thus, without relying on high-precision explicit disparity output, the inference accuracy and robustness of two-dimensional tasks can be effectively improved by using binocular input. Furthermore, it blocks the transmission of disparity estimation error to the two-dimensional perception link, eliminates multi-level error accumulation, and thus improves the accuracy of target pose estimation results.

[0068] As a fourth optional embodiment of the present invention, based on the above embodiments, one way to perform three-dimensional regularization processing on the binocular cost volume to obtain the initial binocular disparity estimation result can be described as follows:

[0069] The stereo cost volume is input into the three-dimensional convolutional network in the three-dimensional regularization module of the image processing model. The stereo cost volume is subjected to convolutional filtering in a set dimension to obtain the first processing result. The set dimension includes the channel dimension, the disparity dimension and the size dimension.

[0070] The first processing result is input into the feature-guided attention network in the three-dimensional regularization module. The corresponding positions in the first processing result are weighted according to the spatial attention weights to obtain the second processing result. The spatial attention weights are generated based on the first feature, which is obtained by filtering the left eye view matching features based on preset feature filtering parameters.

[0071] The second processing result is input into the hourglass aggregation network in the three-dimensional regularization module, and the second processing result is aggregated across scale contexts in the disparity dimension and spatial dimension to obtain the initial binocular disparity estimation result.

[0072] In this embodiment, the 3D regularization module can be understood as a module within the image processing model used for regularization processing such as noise reduction and cross-scale information aggregation on the binocular cost volume to obtain the regularized geometric code volume (i.e., the initial binocular disparity estimation result). The 3D regularization module sequentially includes a 3D convolutional network, a feature-guided attention network, and an hourglass aggregation network.

[0073] One implementation approach is to input the constructed binocular cost volume into a 3D convolutional network. 3×3×3 convolutions can be performed across the channel dimension, disparity dimension, and spatial size dimension (height and width) to perform preliminary filtering on the binocular cost volume, yielding a first processing result. Next, this first processing result can be input into a feature-guided attention network. This network uses the first features obtained from the left-view matching features to generate spatial attention weights. These weights are then used to weight the first processing result spatially, assigning different confidence weights to different image locations to suppress interference from the background and low-confidence regions, outputting a second processing result. The second processing result is then input into an hourglass aggregation network. This network uses an hourglass-shaped encoding / decoding structure to first downsample the spatial and disparity dimensions to capture global disparity context information, and then upsamples to restore the original resolution. This achieves cross-disparity and cross-spatial context aggregation, fusing local detail information with global disparity constraints, eliminating conflicts caused by local mismatches, and finally outputting the initial binocular disparity estimation result through disparity regression calculation.

[0074] In this embodiment, the feature-guided attention network can pre-select the left-eye view matching features based on the acquired preset feature selection parameters to obtain the first feature. One implementation of feature selection based on the preset feature selection parameters for the left-eye view matching features is to downsample the resolution of the left-eye view matching features. For example, when the preset feature selection parameters are... In this case, it can be done according to The magnification is used to downsample the matching features of the left eye view to obtain the first feature.

[0075] The technical solution described in this embodiment effectively eliminates mismatch noise within the binocular cost body through a 3D convolutional network. It guides the generation of spatial attention weights using left-eye view matching features and adaptively weights the first processing result by position based on confidence, reducing the impact of background and low-confidence regions on subsequent disparity calculations. This is equivalent to using image context to suppress unreliable matching and strengthen reliable matching. Combined with an hourglass aggregation network, it achieves cross-disparity and cross-spatial context aggregation, balancing local details and global constraints. This improves the accuracy of the initial binocular disparity estimation results and reduces disparity estimation bias caused by weak textures and lighting variations, providing higher-quality intermediate disparity data for subsequent disparity refinement and pose determination. The data processing procedures of the feature-guided attention network and the hourglass aggregation network can be found in descriptions of existing open-source spatial attention networks and hourglass encoding / decoding networks (such as the 3D U-Net network), and are not limited here.

[0076] As a fifth optional embodiment of the present invention, based on the above embodiments, the step of determining the target pose estimation result of the target object according to the semantic features, the binocular cost volume, the two-dimensional task prediction result, and the initial binocular disparity estimation result can be further optimized to the following steps:

[0077] The pose estimation module in the image processing model inputs the two-dimensional key point prediction results from the two-dimensional task prediction results into the geometric solution network in the pose estimation module. The geometric solution network obtains the initial pose estimation results based on the two-dimensional key point prediction results, the three-dimensional structural information of the target object, and the intrinsic parameters of the binocular vision sensor combined with a robust solver.

[0078] The initial binocular disparity estimation result is refined by the disparity refinement network in the pose estimation module based on the semantic features and the binocular cost body to obtain the target binocular disparity estimation result.

[0079] The point cloud representation in the attitude estimation module is used to obtain a 3D point cloud representation based on the target binocular disparity estimation result and the intrinsic parameters of the binocular vision sensor.

[0080] The target pose estimation result is obtained by combining the pose estimation network in the pose estimation module with the point cloud registration algorithm based on the 3D point cloud representation and the initial pose estimation result.

[0081] In this embodiment, as one implementation method, the 2D keypoint prediction results can first be input into a geometric solving network. The geometric solving network combines the known 3D structural information of the target object with the intrinsic parameters of the binocular vision sensor, and calls a robust solver to suppress interference from points other than mismatched points, thus calculating the initial pose estimation result of the target object. The initial pose estimation result can be understood as a coarse-grained pose estimation result corresponding to the target object, meeting the requirements for finer-grained estimation. For example, the robust solver can be a Perspective-n-Point (PnP) solver based on the Random Sample Consensus (RANSAC) algorithm.

[0082] Next, the initial binocular disparity estimation results are input into the disparity refinement network. It is known that stereo disparity estimation usually requires the participation of contextual features, while two-dimensional tasks require the participation of semantic features. Through experiments, it was found that the semantic features of two-dimensional tasks and the contextual features used in the binocular disparity estimation process are basically the same in terms of network structure and feature representation. Therefore, it is proposed that part of the feature extraction process be shared, that is, the semantic features are transformed into the contextual features (both are high-level information features). By sharing, redundant calculations are reduced and the overall computational cost of the system is reduced.

[0083] Following the above description, after sharing, semantic features are used as the required contextual features to participate in the estimation of the target pose. Semantic features and the stereo cost volume are used as auxiliary constraints. By fusing semantic features, the stereo cost volume, and the initial stereo disparity estimation results through convolution, the noisy initial stereo disparity estimation results can be refined through edge repair, detail completion, and sub-pixel optimization, outputting a more accurate target stereo disparity estimation result. Then, the point cloud representation network uses the intrinsic parameters of the stereo vision sensor and the stereo imaging triangulation principle to convert each effective pixel of the target object in the target stereo disparity estimation result into a 3D spatial point in the stereo vision sensor coordinate system. These spatial points are then aggregated to form a 3D point cloud representation, filtering out invalid background points. Finally, the pose estimation network uses the initial pose estimation result as the initial value for registration iteration. Combined with the observation constraints of the 3D point cloud representation, the point cloud registration algorithm is executed to iteratively optimize and correct the pose, continuously updating the rotation and translation pose parameters until the iterative convergence condition is met, outputting the target pose estimation result corresponding to the target object. For example, the point cloud registration algorithm can employ the Iterative Closest Point (ICP) registration algorithm.

[0084] The above technical solution obtains the initial pose estimation result by geometrically solving based on two-dimensional key points and three-dimensional structural information, without relying solely on the displayed high-precision binocular disparity result. This allows for the output of usable initial pose values ​​even when the quality of the binocular disparity result is poor. Then, by using semantic features and binocular cost volume to jointly constrain and complete disparity refinement, higher quality disparity and three-dimensional point cloud are obtained. Point cloud registration and fine-tuning are performed starting from the reliable initial pose estimation result, alleviating the problem of disparity error accumulating and propagating to the final pose layer by layer. This improves the accuracy and robustness of the overall pose estimation in complex scenes such as weak texture and low binocular disparity quality.

[0085] As one implementation, based on the above embodiments, the training steps of the image processing model include:

[0086] Obtain a sample training set and an initial image processing model. The sample training set includes sample groups, which include binary sample image pairs as input and supervision labels corresponding to the binary sample image pairs. The supervision labels include actual disparity results, actual two-dimensional task prediction results, or actual pose estimation results. The binary sample image pairs include a sample left eye image and a sample right eye image.

[0087] The binary sample image pairs are input into the initial image processing model to obtain the processing results of the initial image processing model. The processing results include target binocular disparity estimation results, two-dimensional task prediction results, and target pose estimation results.

[0088] Based on the supervision label and the processing result corresponding to the supervision label, determine the first loss function value of the complete intersection-union loss function, the second loss function value of the binary cross-entropy loss function, the third loss function value of the target key point similarity loss function, and the fourth loss function value of the smooth L1 loss function. Based on the first loss function value, the second loss function value, the third loss function value, and the fourth loss function value, determine the target loss function value.

[0089] Update the network parameters of each module included in the initial image processing model according to the target loss function value, return and re-execute the step of inputting the binary sample image pair into the initial image processing model until the training termination condition is met, and determine the initial image processing model corresponding to the training termination as the image processing model.

[0090] In this embodiment, the sample training set can be considered as a dataset used to train the initial image processing model to obtain the final image processing model that can be used for inference, and it consists of several sample groups. The initial image processing model can be understood as a model to be trained that has the same functional modules as the image processing model, but whose network parameters have not yet converged. Multiple tasks share the same semantic feature extraction network and intermediate representation, which not only reduces the model parameter size and redundant computation overhead, but also forms a more compact and efficient computation flow, effectively reducing the overall inference latency and resource consumption of the system. The processing process of the left and right eye images of the sample in the binary sample image pair by the initial image processing model is similar to the processing process of the left and right eye images in the above embodiment, and will not be repeated here.

[0091] A sample set can be considered the smallest data unit for training, containing a pair of binary sample images and their corresponding supervision labels. Supervision labels can be understood as the ground truth labeled data corresponding to the binary sample image pairs, used to measure the deviation between the model output and the ground truth. The training termination condition can be considered the criterion for stopping iterative training, which could be conditions such as reaching a preset upper limit for the number of iterations, the target loss function value decreasing to a corresponding threshold, and / or the target loss function value no longer decreasing and converging.

[0092] In this embodiment, as one implementation method, the initial image processing model can be pre-trained using a sample training set with actual disparity results as supervision labels (such as public datasets like SceneFlow and KITTI) to optimize the stereo disparity estimation capability of the initial image processing model. Then, the initial image processing model can be fine-tuned based on a sample training set with actual two-dimensional task prediction results as supervision labels to optimize the two-dimensional task inference capability of the initial image processing model. Finally, the initial image processing model can be further fine-tuned using a sample training set with actual pose estimation results as supervision labels to optimize the six-dimensional pose estimation capability of the initial image processing model.

[0093] For example, when performing further fine-tuning, the input image resolution during training can be the same as that during inference, which can be 512x512. The AdamW optimizer can be used with a base learning rate of 0.0002, a batch size of 4, a weight decay of 0.00001, and a maximum number of iterations of 20,000.

[0094] For example, one way to determine the target loss function value can be described as follows: The target loss function value is obtained by weighted summing of the first loss function value, the second loss function value, the third loss function value, and the fourth loss function value. For example, all weights can be 1. One possible way to express this is:

[0095] ;

[0096] in, The first loss function value of the complete intersection-union loss function; This represents the second loss function value of the binary cross-entropy loss function; The value of the third loss function in the target keypoint similarity loss function; This is the fourth loss function value of the smoothed L1 loss function.

[0097] The above technical solution constructs a unified multi-task model framework (i.e., the initial image processing model) that simultaneously supports 2D and 3D tasks. Within a single model, it jointly completes 2D tasks (including object detection, object segmentation, and 2D keypoint prediction) and 3D tasks (including binocular disparity modeling). Furthermore, each module is jointly optimized through a unified end-to-end training framework, leveraging 2D task constraints, binocular geometric constraints, and pose constraints to improve overall 6D pose estimation performance. Simultaneously, a joint training mechanism for 2D and 3D tasks is introduced within the unified framework, fully utilizing the complementary information provided by binocular input to determine the spatial location and structure of the target. Multi-dimensional, strongly constrained joint supervision, through collaborative optimization among multiple tasks, allows 2D tasks to benefit from the auxiliary constraints of 3D geometric information during training, thereby significantly improving the accuracy and stability of 2D task inference (such as 2D keypoint prediction). The improvement in 2D perception results also promotes further improvement in the performance of 6D pose estimation. In addition, by realizing the end-to-end modeling and inference process from binocular input to instance-level 6D pose output, the error accumulation problem caused by multi-stage pipelined processing is avoided, so that the trained image processing model can be applied to scenarios with high requirements for real-time performance, stability, multi-task collaboration and computational efficiency.

[0098] To better understand the image processing method provided in the embodiments of the present invention, a specific example is given here. Figure 2 This is a schematic diagram illustrating the principle of an image processing method provided in an embodiment of the present invention.

[0099] like Figure 2As shown, image processing is performed using an image processing model, which includes a feature extraction module, a binocular cost body construction module, a geometric feature injection module, a 2D task inference module, a 3D regularization module, and a pose estimation module. The image processing model receives the left and right eye images simultaneously acquired by the binocular cameras as input. These images are simultaneously input into the matching feature extraction network of the feature extraction module, resulting in multi-scale left-eye view matching features associated with the left eye image and multi-scale right-eye view matching features associated with the right eye image. These left-eye and right-eye view matching features are collectively referred to as left and right-eye view matching features. Additionally, the left eye image is input into the semantic feature extraction network set up for the 2D vision task in the feature extraction module to obtain corresponding multi-scale semantic features. Then, the right-eye view matching features are downsampled, and the left-eye view matching features and the downsampled right-eye view matching features are input into the binocular cost body construction module to obtain the output binocular cost body. Finally, the binocular cost body is input into the geometric feature injection module and the 3D regularization module, respectively. The stereo cost volume is mapped to obtain corresponding stereo geometric features through the feature mapping network of the geometric feature injection module, and the stereo geometric features and semantic features are fused through the feature fusion network of the geometric feature injection module to obtain fused features. The stereo cost volume is then subjected to three-dimensional regularization processing by the three-dimensional regularization module to obtain the initial stereo disparity estimation result.

[0100] The fused features are then input into the 2D task inference module for 2D task inference, yielding a 2D task prediction result. This module can include multiple parallel 2D task inference networks, such as an object detection network, an object segmentation network, and a 2D keypoint prediction network. The 2D task prediction result includes the 2D bounding box of the target object, pixel-level segmentation results, and 2D keypoint prediction results. The 2D keypoint prediction results are then input into the geometry solving network of the pose estimation module to obtain the initial pose estimation result. The initial binocular disparity estimation result, semantic features, and binocular cost volume are then input into the disparity refinement network of the pose estimation module to obtain the target binocular disparity estimation result. Finally, the target binocular disparity estimation result and the intrinsic parameters of the binocular vision sensor are input into the point cloud representation network of the pose estimation module to obtain a 3D point cloud representation. The 3D point cloud representation and the initial pose estimation result are then input into the pose estimation network of the pose estimation module to obtain the target pose estimation result.

[0101] Figure 3 This is a schematic diagram of the structure of an image processing apparatus provided in an embodiment of the present invention. Figure 3 As shown, the device includes: a feature extraction module 31, a cost body construction module 32, a first estimation module 33, and a second estimation module 34.

[0102] The feature extraction module 31 is used to acquire the left and right eye images synchronously collected by the binocular vision sensor for the target object, and to determine the semantic features and left eye view matching features associated with the left eye image, and the right eye view matching features associated with the right eye image.

[0103] The cost body construction module 32 is used to construct a binocular cost body based on the left eye view matching features and the right eye view matching features. The binocular cost body is used to characterize the matching association relationship between left eye image pixels and right eye image pixels under each preset candidate disparity.

[0104] The first estimation module 33 is used to obtain a two-dimensional task prediction result corresponding to the target object based on the stereo cost body and the semantic features, and to perform three-dimensional regularization processing on the stereo cost body to obtain an initial stereo disparity estimation result; wherein, obtaining the two-dimensional task prediction result corresponding to the target object based on the stereo cost body and the semantic features includes: performing feature mapping on the stereo cost body to obtain corresponding stereo geometric features, and performing two-dimensional task inference based on the fusion features of the stereo geometric features and the semantic features to obtain the two-dimensional task prediction result;

[0105] The second estimation module 34 is used to determine the target pose estimation result of the target object based on the semantic features, the binocular cost volume, the two-dimensional task prediction result, and the initial binocular disparity estimation result.

[0106] The image processing apparatus provided in this embodiment of the invention acquires left and right eye images synchronously collected by a binocular vision sensor for a target object, and determines semantic features and left eye view matching features associated with the left eye image, and right eye view matching features associated with the right eye image; constructs a binocular cost body based on the left eye view matching features and right eye view matching features, the binocular cost body being used to characterize the matching association relationship between left eye image pixels and right eye image pixels under each preset candidate disparity; obtains a two-dimensional task prediction result corresponding to the target object based on the binocular cost body and semantic features, and performs three-dimensional regularization processing on the binocular cost body to obtain an initial binocular disparity estimation result; wherein, obtaining a two-dimensional task prediction result corresponding to the target object based on the binocular cost body and semantic features includes: performing feature mapping on the binocular cost body to obtain corresponding binocular geometric features, performing two-dimensional task inference based on the fusion features of binocular geometric features and semantic features to obtain a two-dimensional task prediction result; and determining a target pose estimation result for the target object based on the semantic features, binocular cost body, two-dimensional task prediction result, and initial binocular disparity estimation result. This device eliminates the reliance on a single, pre-estimated disparity feature. By constructing a binocular cost body based on feature information from the binocular input image, it fully preserves the feature matching and association information of each pixel under all candidate disparity assumptions. Pose estimation is then performed based on the binocular cost body, combined with semantic features, 2D task prediction results, and initial binocular disparity estimation results. This effectively suppresses error accumulation caused by disparity mismatches and balances 2D semantic priors with 3D geometric constraints. It improves the accuracy and robustness of target pose estimation in complex scenarios with low binocular disparity quality, difficult stereo matching, or occlusion and weak textures, thus meeting the needs of practical applications. Furthermore, it can simultaneously complete 2D tasks (including but not limited to target detection, target segmentation, and 2D keypoint prediction) and 3D tasks (including binocular geometric information modeling and 6D pose estimation) during a single forward inference process, improving inference efficiency and resource conservation. This makes it suitable for scenarios with higher requirements for multi-task collaboration and computational efficiency, such as humanoid robots and augmented reality.

[0107] Furthermore, the feature extraction module 31 can specifically be used for:

[0108] The left eye image is input into the semantic feature extraction network of the feature extraction module in the image processing model to extract semantic features from the left eye image and obtain the semantic features corresponding to the left eye image.

[0109] The left and right images are respectively input into the matching feature extraction network of the feature extraction module. The left image is used to extract matching features to obtain left view matching features, and the right image is used to extract matching features to obtain right view matching features.

[0110] Furthermore, the cost body construction module 32 may specifically include a cost body construction unit.

[0111] The cost volume construction unit is used to determine the feature matching relationship between the left eye view matching feature and the right eye view matching feature according to each preset candidate disparity through the binocular cost volume construction module in the image processing model, and form the binocular cost volume based on the feature matching relationship.

[0112] Furthermore, the cost body construction unit can specifically be used for:

[0113] Obtain the candidate disparity for each preset value;

[0114] For each candidate disparity, the right eye view matching feature is horizontally offset according to the offset corresponding to the candidate disparity to obtain the right eye view matching feature after the candidate disparity is offset.

[0115] For each pixel position in the left eye view matching feature, the first feature vector is matched with the second feature vector in the offset right eye view matching feature corresponding to the pixel position to obtain the feature matching strength of the pixel position under the candidate disparity.

[0116] After traversing all pixel positions and all candidate disparities, all feature matching intensities are integrated according to a preset dimension to form a multidimensional binocular cost volume.

[0117] Furthermore, the first estimation module 33 can specifically be used for:

[0118] The stereo cost volume is input into the feature mapping network of the geometric feature injection module in the image processing model, and the stereo cost volume is subjected to feature mapping to obtain the corresponding stereo geometric features. The stereo geometric features are matched with the spatial resolution and feature dimension of the semantic features.

[0119] The binocular geometric features and the semantic features are input into the feature fusion network of the geometric feature injection module, and the binocular geometric features and the semantic features are fused to obtain fused features.

[0120] The fused features are input into the two-dimensional task inference module of the image processing model to obtain the two-dimensional task prediction result output after performing two-dimensional task inference.

[0121] Furthermore, the first estimation module 33 can specifically be used for:

[0122] The stereo cost volume is input into the three-dimensional convolutional network in the three-dimensional regularization module of the image processing model. The stereo cost volume is subjected to convolutional filtering in a set dimension to obtain the first processing result. The set dimension includes the channel dimension, the disparity dimension and the size dimension.

[0123] The first processing result is input into the feature-guided attention network in the three-dimensional regularization module. The corresponding positions in the first processing result are weighted according to the spatial attention weights to obtain the second processing result. The spatial attention weights are generated based on the first feature, which is obtained by filtering the left eye view matching features based on preset feature filtering parameters.

[0124] The second processing result is input into the hourglass aggregation network in the three-dimensional regularization module, and the second processing result is aggregated across scale contexts in the disparity dimension and spatial dimension to obtain the initial binocular disparity estimation result.

[0125] Furthermore, the second estimation module 34 can specifically be used for:

[0126] The pose estimation module in the image processing model inputs the two-dimensional key point prediction results from the two-dimensional task prediction results into the geometric solution network in the pose estimation module. The geometric solution network obtains the initial pose estimation results based on the two-dimensional key point prediction results, the three-dimensional structural information of the target object, and the intrinsic parameters of the binocular vision sensor combined with a robust solver.

[0127] The initial binocular disparity estimation result is refined by the disparity refinement network in the pose estimation module based on the semantic features and the binocular cost body to obtain the target binocular disparity estimation result.

[0128] The point cloud representation in the attitude estimation module is used to obtain a 3D point cloud representation based on the target binocular disparity estimation result and the intrinsic parameters of the binocular vision sensor.

[0129] The target pose estimation result is obtained by combining the pose estimation network in the pose estimation module with the point cloud registration algorithm based on the 3D point cloud representation and the initial pose estimation result.

[0130] Furthermore, the device also includes a training module, which can be specifically used for:

[0131] Obtain a sample training set and an initial image processing model. The sample training set includes sample groups, which include binary sample image pairs as input and supervision labels corresponding to the binary sample image pairs. The supervision labels include actual initial disparity results, actual target disparity results, actual two-dimensional task prediction results, or actual pose estimation results. The binary sample image pairs include a sample left eye image and a sample right eye image.

[0132] The binary sample image pairs are input into the initial image processing model to obtain the processing results of the initial image processing model. The processing results include initial binocular disparity estimation results, target binocular disparity estimation results, two-dimensional task prediction results, and target pose estimation results.

[0133] Based on the supervision label and the processing result corresponding to the supervision label, determine the first loss function value of the complete intersection-union loss function, the second loss function value of the binary cross-entropy loss function, the third loss function value of the target key point similarity loss function, and the fourth loss function value of the smooth L1 loss function. Based on the first loss function value, the second loss function value, the third loss function value, and the fourth loss function value, determine the target loss function value.

[0134] Update the network parameters of each module included in the initial image processing model according to the target loss function value, return and re-execute the step of inputting the binary sample image pair into the initial image processing model until the training termination condition is met, and determine the initial image processing model corresponding to the training termination as the image processing model.

[0135] The image processing apparatus provided in the embodiments of the present invention can execute the image processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0136] Figure 4 A schematic diagram of an electronic device 40 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0137] like Figure 4 As shown, the electronic device 40 includes at least one processor 41 and a memory, such as a read-only memory (ROM) 42 or a random access memory (RAM) 43, communicatively connected to the at least one processor 41. The memory stores computer programs executable by the at least one processor. The processor 41 can perform various appropriate actions and processes based on the computer program stored in the ROM 42 or loaded from storage unit 48 into the RAM 43. The RAM 43 may also store various programs and data required for the operation of the electronic device 40. The processor 41, ROM 42, and RAM 43 are interconnected via a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.

[0138] Multiple components in electronic device 40 are connected to I / O interface 45, including: input unit 46, such as keyboard, mouse, etc.; output unit 47, such as various types of monitors, speakers, etc.; storage unit 48, such as disk, optical disk, etc.; and communication unit 49, such as network card, modem, wireless transceiver, etc. Communication unit 49 allows electronic device 40 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0139] Processor 41 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 41 performs the various methods and processes described above, such as image processing methods.

[0140] In some embodiments, the image processing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 48. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 40 via ROM 42 and / or communication unit 49. When the computer program is loaded into RAM 43 and executed by processor 41, one or more steps of the image processing method described above may be performed. Alternatively, in other embodiments, processor 41 may be configured to perform the image processing method by any other suitable means (e.g., by means of firmware).

[0141] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include: implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0142] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0143] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0144] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0145] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0146] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0147] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0148] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. An image processing method, characterized in that, include: Acquire the left and right eye images of the target object simultaneously captured by the binocular vision sensor, and determine the semantic features and left eye view matching features associated with the left eye image, and the right eye view matching features associated with the right eye image; A binocular cost body is constructed based on the left-eye view matching features and the right-eye view matching features. The binocular cost body is used to characterize the matching association relationship between left-eye image pixels and right-eye image pixels under each preset candidate disparity. Based on the binocular cost body and the semantic features, a two-dimensional task prediction result corresponding to the target object is obtained, and the binocular cost body is subjected to three-dimensional regularization processing to obtain an initial binocular disparity estimation result; wherein, obtaining the two-dimensional task prediction result corresponding to the target object based on the binocular cost body and the semantic features includes: performing feature mapping on the binocular cost body to obtain corresponding binocular geometric features, and performing two-dimensional task inference based on the fusion features of the binocular geometric features and the semantic features to obtain the two-dimensional task prediction result; Based on the semantic features, the binocular cost volume, the two-dimensional task prediction results, and the initial binocular disparity estimation results, the target pose estimation result of the target object is determined.

2. The method according to claim 1, characterized in that, The determination of semantic features and left-eye view matching features associated with the left-eye image, and right-eye view matching features associated with the right-eye image, includes: The left eye image is input into the semantic feature extraction network of the feature extraction module in the image processing model to extract semantic features from the left eye image and obtain the semantic features corresponding to the left eye image. The left and right images are respectively input into the matching feature extraction network of the feature extraction module. The left image is used to extract matching features to obtain left view matching features, and the right image is used to extract matching features to obtain right view matching features.

3. The method according to claim 1, characterized in that, The step of constructing a stereo cost body based on the left-eye view matching features and the right-eye view matching features includes: The binocular cost body construction module in the image processing model determines the feature matching relationship between the left-eye view matching feature and the right-eye view matching feature based on each preset candidate disparity, and forms the binocular cost body based on the feature matching relationship.

4. The method according to claim 3, characterized in that, The step of determining the feature matching relationship between the left-eye view matching feature and the right-eye view matching feature based on each preset candidate disparity, and forming the binocular cost body based on the feature matching relationship, includes: Obtain the candidate disparity for each preset value; For each candidate disparity, the right eye view matching feature is horizontally offset according to the offset corresponding to the candidate disparity to obtain the right eye view matching feature after the candidate disparity is offset. For each pixel position in the left eye view matching feature, the first feature vector is matched with the second feature vector in the offset right eye view matching feature corresponding to the pixel position to obtain the feature matching strength of the pixel position under the candidate disparity. After traversing all pixel positions and all candidate disparities, all feature matching intensities are integrated according to a preset dimension to form a multi-dimensional binocular cost volume.

5. The method according to claim 1, characterized in that, The step of performing feature mapping on the stereo cost body to obtain corresponding stereo geometric features, and performing two-dimensional task inference based on the fusion features of the stereo geometric features and the semantic features to obtain a two-dimensional task prediction result includes: The stereo cost volume is input into the feature mapping network of the geometric feature injection module in the image processing model, and the stereo cost volume is subjected to feature mapping to obtain the corresponding stereo geometric features. The stereo geometric features are matched with the spatial resolution and feature dimension of the semantic features. The binocular geometric features and the semantic features are input into the feature fusion network of the geometric feature injection module, and the binocular geometric features and the semantic features are fused to obtain fused features. The fused features are input into the two-dimensional task inference module of the image processing model to obtain the two-dimensional task prediction result output after two-dimensional task inference.

6. The method according to claim 1, characterized in that, The step of performing three-dimensional regularization on the binocular cost volume to obtain the initial binocular disparity estimation result includes: The stereo cost volume is input into the three-dimensional convolutional network in the three-dimensional regularization module of the image processing model. The stereo cost volume is subjected to convolutional filtering in a set dimension to obtain the first processing result. The set dimension includes the channel dimension, the disparity dimension and the size dimension. The first processing result is input into the feature-guided attention network in the three-dimensional regularization module. The corresponding positions in the first processing result are weighted according to the spatial attention weights to obtain the second processing result. The spatial attention weights are generated based on the first feature, which is obtained by filtering the left eye view matching features based on preset feature filtering parameters. The second processing result is input into the hourglass aggregation network in the three-dimensional regularization module, and the second processing result is aggregated across scale contexts in the disparity dimension and spatial dimension to obtain the initial binocular disparity estimation result.

7. The method according to claim 1, characterized in that, The step of determining the target pose estimation result of the target object based on the semantic features, the binocular cost volume, the two-dimensional task prediction result, and the initial binocular disparity estimation result includes: The pose estimation module in the image processing model inputs the two-dimensional key point prediction results from the two-dimensional task prediction results into the geometric solution network in the pose estimation module. The geometric solution network obtains the initial pose estimation results based on the two-dimensional key point prediction results, the three-dimensional structural information of the target object, and the intrinsic parameters of the binocular vision sensor combined with a robust solver. The initial binocular disparity estimation result is refined by the disparity refinement network in the pose estimation module based on the semantic features and the binocular cost body to obtain the target binocular disparity estimation result. The point cloud representation in the attitude estimation module is used to obtain a 3D point cloud representation based on the target binocular disparity estimation result and the intrinsic parameters of the binocular vision sensor. The target pose estimation result is obtained by combining the pose estimation network in the pose estimation module with the point cloud registration algorithm based on the 3D point cloud representation and the initial pose estimation result.

8. The method according to any one of claims 2-5, characterized in that, The training steps of the image processing model include: Obtain a sample training set and an initial image processing model. The sample training set includes sample groups, which include binary sample image pairs as input and supervision labels corresponding to the binary sample image pairs. The supervision labels include actual initial disparity results, actual target disparity results, actual two-dimensional task prediction results, or actual pose estimation results. The binary sample image pairs include a sample left eye image and a sample right eye image. The binary sample image pairs are input into the initial image processing model to obtain the processing results of the initial image processing model. The processing results include initial binocular disparity estimation results, target binocular disparity estimation results, two-dimensional task prediction results, and target pose estimation results. Based on the supervision label and the processing result corresponding to the supervision label, determine the first loss function value of the complete intersection-union loss function, the second loss function value of the binary cross-entropy loss function, the third loss function value of the target key point similarity loss function, and the fourth loss function value of the smooth L1 loss function. Based on the first loss function value, the second loss function value, the third loss function value, and the fourth loss function value, determine the target loss function value. Update the network parameters of each module included in the initial image processing model according to the target loss function value, return and re-execute the step of inputting the binary sample image pair into the initial image processing model until the training termination condition is met, and determine the initial image processing model corresponding to the training termination as the image processing model.

9. An image processing apparatus, characterized in that, include: The feature extraction module is used to acquire the left and right eye images of the target object simultaneously collected by the binocular vision sensor, and to determine the semantic features and left eye view matching features associated with the left eye image, and the right eye view matching features associated with the right eye image. The cost body construction module is used to construct a binocular cost body based on the left-eye view matching features and the right-eye view matching features. The binocular cost body is used to characterize the matching association relationship between left-eye image pixels and right-eye image pixels under each preset candidate disparity. The first estimation module is used to obtain a two-dimensional task prediction result corresponding to the target object based on the stereo cost body and the semantic features, and to perform three-dimensional regularization processing on the stereo cost body to obtain an initial stereo disparity estimation result; wherein, obtaining the two-dimensional task prediction result corresponding to the target object based on the stereo cost body and the semantic features includes: performing feature mapping on the stereo cost body to obtain corresponding stereo geometric features, and performing two-dimensional task inference based on the fusion features of the stereo geometric features and the semantic features to obtain the two-dimensional task prediction result; The second estimation module is used to determine the target pose estimation result of the target object based on the semantic features, the binocular cost volume, the two-dimensional task prediction result, and the initial binocular disparity estimation result.

10. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the image processing method according to any one of claims 1-8.