Charging equipment, charging control method and device thereof and medium

By determining the initial pose in the charging robot system and combining binocular stereo matching and multi-task model optimization, the problem of insufficient positioning accuracy of the charging port was solved, and sub-millimeter-level precise docking was achieved in complex environments.

CN120863391AActive Publication Date: 2025-10-31CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202511406234.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2025-10-31
Estimated Expiration
2045-09-29

AI Technical Summary

Technical Problem

Existing charging robot systems have significant technical bottlenecks in millimeter-level positioning accuracy, especially in complex environments where the positioning accuracy of the charging port is insufficient. Traditional methods are easily affected by environmental interference and lack closed-loop optimization mechanisms, resulting in large positioning errors and high system latency.

Method used

By determining the initial pose of the charging port, and combining binocular stereo matching with the three-dimensional spatial coordinates of the preset charging port model, the depth map and pose are alternately optimized and updated. The multi-task model is used to improve image quality and perform environmental perception and image enhancement, thus establishing a closed-loop optimization mechanism.

Benefits of technology

It achieves sub-millimeter-level positioning of the charging port in complex environments, improving positioning accuracy and environmental robustness, ensuring precise alignment of the charging plug, and meeting the needs of automatic charging of electric vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120863391A_ABST
    Figure CN120863391A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of charging control, and discloses a charging device and a charging control method and device thereof, and a medium, and the method comprises the steps: determining the initial pose of a charging port, obtaining an initial depth map through binocular stereo matching, carrying out the alternative optimization and updating of the depth map and the pose, achieving the precise alignment of a charging plug and the charging port, and improving the charging precision. The problems of insufficient depth map precision and large pose calculation error caused by environmental interference in traditional charging positioning are effectively solved, depth deviation and pose errors can be dynamically corrected through alternate optimization of the depth map and the pose, limitation of single-mode positioning is avoided, the target pose obtained through final optimization can achieve charging port submillimeter positioning, and the positioning precision of the charging port is improved. According to the invention, accurate alignment and insertion of the charging plug are ensured, a high docking success rate can still be maintained in complex environments such as rain and fog, low illumination and the like, the environment robustness and the positioning accuracy of automatic charging equipment are greatly improved, and the millimeter-level plugging requirement of automatic charging of an electric vehicle is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of charging control technology, specifically to charging equipment and its charging control methods, devices, and media. Background Technology

[0002] With the rapid breakthroughs in artificial intelligence, robot vision, and force sensing fusion technologies, and the continuous increase in the number of new energy vehicles, the closed-loop scenario of Automated Valet Parking (AVP) + Automatic Charging has become a key node in the integration of smart cities and new energy strategies in recent years due to its ability to reconstruct user experience and energy efficiency. However, existing charging robot systems face significant technical bottlenecks in areas such as millimeter-level positioning accuracy. Traditional single-geometric algorithms, such as Perspective-n-Point (PNP), are susceptible to image noise interference, and pose calculation errors increase dramatically when the charging port is partially occluded or lacks texture. Monocular depth estimation algorithms exhibit greater depth estimation errors as the target becomes smaller and farther away, making it difficult to meet millimeter-level insertion requirements for distant, small targets (such as car charging ports). Therefore, achieving millimeter-level charging port positioning has become an urgent problem to be solved for automated charging robots. Summary of the Invention

[0003] In view of this, the present invention provides a charging device and its charging control method, apparatus and medium to solve the problem of low positioning accuracy when charging robots locate the charging port in related technologies.

[0004] In a first aspect, the present invention provides a charging control method for a charging device, the charging device comprising: a charging plug, the method comprising: After controlling the charging plug to move to the preset position of the target charging port, the initial pose of the target charging port is determined, and a binocular image containing the target charging port is acquired. The initial pose is the pose of the target charging port relative to the camera coordinate system corresponding to the binocular image. Perform stereo matching on the binocular images to obtain the initial depth map corresponding to the binocular images; Based on the three-dimensional spatial coordinates of key points in the preset charging port model, the initial depth map and the initial pose are alternately optimized and updated until the preset optimization conditions are met, and the optimized initial pose is determined as the target pose of the target charging port; the preset charging port model is of the same type as the target charging port. Based on the target pose, the charging plug is aligned with the target charging port and then inserted.

[0005] This invention determines the initial pose of the charging port and uses an initial depth map obtained through binocular stereo matching. It then alternately optimizes and updates the depth map and pose, finally using the optimized target pose to achieve precise alignment between the charging plug and the charging port. This effectively solves the problems of insufficient depth map accuracy and large pose calculation errors caused by environmental interference in traditional charging positioning. By first determining the initial pose and obtaining the initial depth map based on binocular images, and then combining the 3D coordinates of key points of the preset charging port model for alternating optimization of the depth map and pose, it can dynamically correct depth deviation and pose error, avoiding the limitations of single-modal positioning. This allows the final optimized target pose to achieve sub-millimeter-level positioning of the charging port, ensuring precise alignment and insertion of the charging plug. It maintains a high docking success rate even in complex environments such as rain, fog, and low light, significantly improving the environmental robustness and positioning accuracy of automatic charging equipment and meeting the millimeter-level insertion requirements of electric vehicle automatic charging.

[0006] In one optional implementation, the preset optimization conditions include: the change in the initial pose before and after optimization satisfies a preset pose change condition or the number of iterations reaches a preset iteration threshold.

[0007] This invention solves the problem of independent depth map and pose calculation, which easily leads to error accumulation, in traditional positioning by first correcting the depth map, then using the corrected depth map for pose optimization, and finally performing a closed-loop operation with severe repetition. Furthermore, each alternating optimization cycle is based on the 3D coordinates of key points of a preset charging port model. The initial depth map is first corrected to eliminate depth deviations caused by environmental interference, and then the initial pose is optimized based on the corrected depth map. The optimization termination is controlled by preset pose change conditions and iteration number thresholds, ensuring optimization accuracy while avoiding system delays caused by excessive iteration. Compared to traditional single-step positioning schemes, this invention can progressively compress positioning errors, significantly improving the final pose accuracy, while also considering system real-time performance to meet the rapid response requirements of charging devices.

[0008] In one optional implementation, the step of alternately optimizing and updating the initial depth map and the initial pose based on the three-dimensional spatial coordinates of key points in the preset charging port model includes: Based on the initial pose, the three-dimensional spatial coordinates of key points in the preset charging port model are mapped to the camera coordinate system and projected onto the image plane of the camera coordinate system to obtain the projection point coordinates and theoretical depth corresponding to the key points; Based on the theoretical depth corresponding to the key point, the first depth corresponding to the key point in the initial depth map is corrected to obtain the corrected depth map. Based on the depth error between the key points in the initial depth map and the corrected depth map, and the reprojection error between the image coordinates of the key points in the monocular image and the projection point coordinates of the key points, the pose change is calculated. The monocular image is one of the images in the binocular image. The process involves obtaining an optimized initial pose based on the initial pose and the pose change, and then mapping the three-dimensional spatial coordinates of key points in the preset charging port model to the camera coordinate system based on the initial pose, and projecting them onto the image plane of the camera coordinate system to obtain the projection point coordinates and theoretical depth of the key points.

[0009] This invention maps the 3D spatial coordinates of key points in a pre-defined charging port model to the camera coordinate system to obtain the theoretical depth, thereby correcting the initial depth map and resolving depth errors caused by texture loss and occlusion in binocular matching. Furthermore, by combining depth error and reprojection error to calculate pose change, it comprehensively considers deviations in both depth and image coordinates, avoiding pose deviations caused by a single error metric. This process can accurately locate the charging port, reducing docking failures caused by inaccurate depth and coordinate offsets. It maintains high positioning accuracy even in scenarios with partial occlusion of the charging port and weak texture, improving the reliability of charging docking.

[0010] In one optional implementation, the step of correcting the first depth corresponding to the key point in the initial depth map based on the theoretical depth corresponding to the key point includes: The corrected first depth value is obtained by weighted summation of the theoretical depth and the first depth.

[0011] This invention corrects depth values ​​by weighting the theoretical depth and the first depth in the initial depth map. The weights are dynamically allocated based on the reliability of both. For example, when the reliability of the first depth is low due to environmental interference (such as fog), the weight of the theoretical depth is increased; conversely, the weight of the first depth is increased. This flexible weighting strategy fully utilizes effective depth information, reduces the impact of single-depth data deviations on positioning, and makes the corrected depth value more closely match the actual scene. This provides a precise depth foundation for subsequent pose optimization, further improving the positioning accuracy of the charging port and ensuring the accurate docking of the charging plug.

[0012] In an optional implementation, the method further includes: Obtain the confidence level of the first depth corresponding to the key point in the initial depth map; A first weight is determined based on the confidence level of the first depth, and the first weight is positively correlated with the confidence level. The second weight corresponding to the theoretical depth is determined based on the first weight; And / or, determine the computational weight corresponding to the depth error based on the confidence level of the depth, wherein the computational weight is positively correlated with the confidence level.

[0013] This invention addresses the problems of traditional weight allocation, such as lack of reliable data and strong subjective assumptions, by introducing depth confidence scores corresponding to key points in the initial depth map. By obtaining the confidence score of the first depth, the first weight is positively correlated with the confidence score, ensuring that the first depth with high confidence plays a greater role in depth correction. Simultaneously, the second weight of the theoretical depth and the calculation weight of the depth error are reasonably determined. For example, in the reflective metal area of ​​the charging port, the first depth confidence score is low, corresponding to a small first weight and a small depth error calculation weight, reducing the interference of unreliable depth data in this area on optimization. This dynamic weight adjustment strategy based on confidence scores allows the optimization process to better reflect the actual data quality, improving the accuracy of depth correction and pose optimization, and enhancing the system's adaptability to complex environments.

[0014] In one alternative implementation, the pose change is calculated using the following formula:

[0015] in, Indicates the change in pose. Indicates the first The image coordinates of each key point in a monocular image. Indicates the first The coordinates of the projection points corresponding to each key point Indicates the first The three-dimensional spatial coordinates of the key points R This indicates that the rotation matrix is ​​used to describe the orientation of the target charging port. t This indicates that the translation vector is used to describe the position of the target charging port. Represents the projection function. This represents the camera intrinsic parameter matrix. Indicates the first The calculation weights corresponding to the depth errors of each key point Indicates the first The depth values ​​corresponding to the key points in the initial depth map. Indicates the first The depth values ​​corresponding to the key points in the corrected depth map.

[0016] This invention addresses the problems of singular error considerations and one-sided optimization directions in traditional pose optimization by incorporating reprojection error and depth error into a unified pose change optimization objective. In the formula, reprojection error considers the deviation between the keypoint image coordinates and the projected point coordinates, while depth error, combined with dynamically calculated weights, considers the deviation in depth before and after correction. The synergistic optimization of both comprehensively corrects pose deviations. Furthermore, the introduction of calculation weights adjusts the influence of depth error in optimization based on depth confidence, preventing low-confidence depth data from interfering with the optimization results. This formula provides a clear mathematical basis for pose optimization, enabling precise calculation of pose adjustments, ensuring the optimized pose is more realistic, significantly improving charging port positioning accuracy, and reducing charging docking failure rates.

[0017] In an optional implementation, the method further includes: Obtain a first image containing the target charging port; The first image is input into a pre-trained multi-task model based on image quality classification and image quality enhancement to obtain the image quality classification result and the second image after image quality enhancement. The image quality classification result is related to the environmental state of the target charging port. Based on the second image, the target location of the target charging port is determined; Based on the target position, the charging plug is controlled to move to a preset position of the target charging port, where the preset position is a position at a set distance from the target position.

[0018] This invention addresses the problems of fragmented environmental perception and image enhancement in traditional positioning, as well as positioning failures caused by poor image quality in complex environments, by introducing a multi-task model to process the first image during the initial charging positioning phase. The multi-task model simultaneously outputs image quality classification results (reflecting the environmental state) and an enhanced second image, clearly defining the current environment (e.g., rain, fog) while eliminating environmental interference, thus providing a high-quality image foundation for subsequent target location determination of the charging port. By determining the target location based on the second image and controlling the charging plug to move to a preset position, misjudgments of the target location due to blurry or low-contrast original images are avoided, significantly improving the accuracy of target positioning in complex environments and laying the foundation for subsequent precise docking.

[0019] In one optional implementation, the multi-task model based on image quality classification and image quality enhancement includes: an encoder, a decoder, and an image quality classification branch, wherein the encoder includes a deep feature extraction network, a channel-space joint attention module, a first convolutional layer, a stacking operation layer, and a first feature dimension transformation layer connected in sequence, and the input end of the deep feature extraction network is used to receive the first image; The decoder includes a second feature dimension transformation layer, a splicing layer, a second convolutional layer, and an upsampling layer connected in sequence. The input of the second feature dimension transformation layer is connected to the output of the deep feature extraction network, and the splicing layer is also connected to the output of the first feature dimension transformation layer. The decoder is used to perform image quality enhancement tasks based on the input image features. The image quality classification branch includes a C3T module, a third convolutional layer, and a fully connected layer connected in sequence. The input of the C3T module is connected to the output of the feature dimension transformation layer. The image quality classification branch is used to perform an image quality classification task based on the input image features.

[0020] This invention addresses the limitations of traditional single-task models and the high latency caused by multiple models being cascaded, through a collaborative design of the encoder, decoder, and image quality classification branch. The encoder sequentially extracts deep image features through a deep feature extraction network and a channel-spatial joint attention module, providing high-quality feature support for subsequent tasks. The decoder performs image quality enhancement based on these features, outputting a clear image. The image quality classification branch classifies environmental states through modules such as C3T. All three components share encoder features, avoiding redundant computation and significantly reducing inference latency. Compared to traditional independent image enhancement and classification models, this structure can perform both tasks simultaneously, improving processing efficiency. Furthermore, the collaborative optimization of each module results in more accurate image quality enhancement and classification results, providing efficient and reliable image preprocessing support for charging positioning.

[0021] In an optional implementation, the multi-task model based on image quality classification and image quality enhancement further includes: A first local attention module is provided between the channel-space joint attention module and the first convolutional layer; A second local attention module is provided between the second convolutional layer and the upsampling layer.

[0022] This invention addresses the problems of insufficient attention to key features of the charging port area and low task accuracy caused by complex background interference in traditional models by adding first and second local attention modules to a multi-task model. The first local attention module, located between the channel-space joint attention module and the first convolutional layer, enhances the feature representation of key areas of the charging port (such as terminals) extracted by the encoder and suppresses background noise. The second local attention module, located between the second convolutional layer and the upsampling layer, focuses on optimizing the details of the charging port area during the decoder's image quality enhancement process, making the charging port features clearer in the output image. Through this local attention mechanism, the model can focus on key areas, improving the accuracy of image quality classification and the targeted nature of image quality enhancement, providing more accurate image information for subsequent localization, and further improving charging port positioning accuracy.

[0023] In an optional implementation, before performing stereo matching on the binocular images to obtain the initial depth map corresponding to the binocular images, the method further includes: The stereo images are input into a pre-trained multi-task model based on image quality classification and image quality enhancement to obtain the image quality classification result and the stereo images after image quality enhancement.

[0024] This invention addresses the problems of large matching errors and low initial depth map accuracy caused by poor original image quality (such as blurring due to rain or fog) in traditional binocular matching by performing multi-task model processing on the binocular images before stereo matching. The multi-task model improves the quality of the binocular images, eliminates the influence of environmental interference on the left and right eye images, and ensures consistency and clarity between the binocular images, providing high-quality input for subsequent stereo matching. Simultaneously, the image quality classification results can help determine the current environment, providing a basis for adjusting subsequent optimization strategies. Compared to directly matching the original binocular images, this processing significantly reduces matching noise, improves the accuracy of the initial depth map, provides a more reliable depth foundation for subsequent alternating optimization updates, and further enhances the accuracy and stability of charging port positioning.

[0025] In a second aspect, the present invention provides a charging control device for a charging device, the charging device comprising: a charging plug, and the device comprising: The first processing module is used to determine the initial pose of the target charging port after controlling the charging plug to move to the preset position of the target charging port, and to acquire a binocular image containing the target charging port. The initial pose is the pose of the target charging port relative to the camera coordinate system corresponding to the binocular image. The second processing module is used to perform stereo matching on the stereo image to obtain the initial depth map corresponding to the stereo image. The third processing module is used to alternately optimize and update the initial depth map and the initial pose based on the three-dimensional spatial coordinates of key points in the preset charging port model, until the preset optimization conditions are met, and then determine the optimized initial pose as the target pose of the target charging port; the preset charging port model and the target charging port are of the same type. The fourth processing module is used to control the charging plug to be aligned with the target charging port and then inserted based on the target pose. The third processing module includes: a first processing subunit, used to map the three-dimensional spatial coordinates of key points in the preset charging port model to the camera coordinate system based on the initial pose, and project them onto the image plane of the camera coordinate system to obtain the projection point coordinates and theoretical depth corresponding to the key points; a second processing subunit, used to correct the first depth corresponding to the key points in the initial depth map based on the theoretical depth corresponding to the key points, to obtain a corrected depth map; a third processing subunit, used to calculate the pose change based on the depth error between the key points in the initial depth map and the corrected depth map, and the reprojection error between the image coordinates corresponding to the key points in the monocular image and the projection point coordinates corresponding to the key points, wherein the monocular image is one image in the binocular image; and a fourth processing subunit, used to obtain the optimized initial pose based on the initial pose and the pose change, and to call the first processing subunit to run.

[0026] Thirdly, the present invention provides a charging device comprising: a charging plug and a controller, the controller comprising: A memory and a processor are communicatively connected, the memory storing computer instructions, and the processor executing the computer instructions to perform the method described in the first aspect and any of its alternative embodiments.

[0027] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions that cause a computer to perform the method provided in the first aspect or any corresponding embodiment thereof.

[0028] The beneficial effects of this invention are: This invention determines the initial pose of the charging port and uses an initial depth map obtained through binocular stereo matching. It then alternately optimizes and updates the depth map and pose, finally using the optimized target pose to achieve precise alignment between the charging plug and the charging port. This effectively solves the problems of insufficient depth map accuracy and large pose calculation errors caused by environmental interference in traditional charging positioning. By first determining the initial pose and obtaining the initial depth map based on binocular images, and then combining the 3D coordinates of key points of the preset charging port model for alternating optimization of the depth map and pose, it can dynamically correct depth deviation and pose error, avoiding the limitations of single-modal positioning. This allows the final optimized target pose to achieve sub-millimeter-level positioning of the charging port, ensuring precise alignment and insertion of the charging plug. It maintains a high docking success rate even in complex environments such as rain, fog, and low light, significantly improving the environmental robustness and positioning accuracy of automatic charging equipment and meeting the millimeter-level insertion requirements of electric vehicle automatic charging. Attached Figure Description

[0029] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0030] Figure 1 This is a flowchart of a charging control method for a charging device according to an embodiment of the present invention; Figure 2 This is a flowchart of a charging control method for another charging device according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the network structure of a multi-task model based on image quality classification and image quality improvement according to an embodiment of the present invention. Figure 4 This is a schematic diagram of the structure of a local attention module according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the C3T module according to an embodiment of the present invention; Figure 6 These are comparison images of image denoising effects according to embodiments of the present invention; Figure 7 This is a diagram illustrating the binocular depth estimation effect according to an embodiment of the present invention; Figure 8 This is a schematic diagram of the main process of automatic charging of the charging device according to an embodiment of the present invention; Figure 9 This is a schematic diagram of the structure of the charging control device of the charging equipment according to an embodiment of the present invention; Figure 10 This is a schematic diagram of the structure of the controller of the charging device according to an embodiment of the present invention. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0032] Existing charging robot systems suffer from significant technical bottlenecks in areas such as adaptability to complex environments and millimeter-level positioning accuracy. Lack of environmental robustness: In severe weather conditions such as rain, fog, and snow, or under low light conditions, images acquired by traditional vision systems are severely degraded (blurred, low contrast), leading to a sharp decline in the performance of deep learning-based recognition models (such as semantic segmentation and object detection). Insufficient positioning accuracy: Single-reliability geometry algorithms (such as PNP) are susceptible to image noise interference, and pose calculation errors increase dramatically when the charging port is partially occluded or texture is missing. Monocular depth estimation algorithms are insufficient to meet millimeter-level insertion requirements for small targets at long distances (such as car charging ports); System-level coordination defects: Existing solutions lack a closed-loop optimization mechanism from environmental perception to image enhancement to localization decision-making, and cannot dynamically adjust processing strategies according to real-time environmental conditions.

[0033] While current mainstream solutions attempt to improve localization robustness by introducing image enhancement or fusing multiple sensors, two major limitations still exist: Enhancement and recognition are separated: Traditional image enhancement methods (such as histogram equalization and physical model-based dehazing algorithms) are designed independently from the subsequent recognition and localization modules, and the enhanced images may not be suitable for the task requirements; Multi-source data are not deeply coupled: visual and depth information are often simply cascaded without establishing a joint optimization mechanism between geometric priors and deep learning estimation, resulting in insufficient positioning accuracy in complex scenarios.

[0034] Overcoming environmental interference and accuracy bottlenecks to achieve "all-weather, all-scenario, millimeter-level" charging port positioning has become a key technological challenge restricting automatic charging robots.

[0035] In real-world electric vehicle automatic charging scenarios, millimeter-level accuracy of the charging connection relies on high-precision 6D pose estimation (3D position + 3D rotation) of the charging port. However, complex environmental degradation and cross-scale positioning error accumulation present existing solutions with a dual challenge: 1. Perception failure under environmental disturbances: Scenes such as rain, fog, and snow cause image blurring and reduced contrast. Traditional vision algorithms cannot simultaneously achieve environmental state discrimination and task-oriented image enhancement, resulting in subsequent segmentation / detection failure. 2. Bottleneck of single-modal localization: Pure geometric methods (such as PNP) rely on the texture features of the charging port, and the pose calculation error exceeds 10cm in occluded or weakly textured scenes. Monocular depth estimation has a ranging error of >5% for small targets (charging ports). While binocular depth algorithms improve accuracy, they lack geometric constraints and are susceptible to matching noise interference. 3. System latency due to module fragmentation: Environmental perception, image enhancement, and localization decision-making operate independently without a closed-loop optimization mechanism, resulting in processing latency exceeding 200ms.

[0036] To address the shortcomings of existing technologies, this invention aims to build a highly robust visual positioning and cross-scene adaptive capability by integrating a multi-task model that combines cross-domain adaptation and image quality evaluation. This will overcome the challenge of precise positioning in the "last millimeter" of complex environments and provide key infrastructure support for the autonomous driving ecosystem.

[0037] According to an embodiment of the present invention, a charging control method for a charging device is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0038] This embodiment provides a charging control method for a charging device, which is an automatic charging device such as an automatic charging robot. The charging device includes a charging plug for insertion into a vehicle's charging port to charge the vehicle. This method is applied to the controller of the charging device, such as a microcontroller (MCU) or other control chip. Figure 1 This is a flowchart of a charging control method for a charging device according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps: Step S101: After controlling the charging plug to move to the preset position of the target charging port, determine the initial pose of the target charging port and acquire a binocular image containing the target charging port.

[0039] The preset position is a pre-calibrated proximity of approximately 200mm to the target charging port, facilitating subsequent precise positioning. The initial pose is the pose of the target charging port relative to the camera coordinate system corresponding to the stereo image. Specifically, this initial pose is the 6D pose (including 3D position and 3D rotation information) of the target charging port relative to the camera coordinate system corresponding to the stereo image, used to describe the position and orientation of the charging port in three-dimensional space, and is the basis for subsequent positioning optimization. This camera coordinate system is a three-dimensional coordinate system established with the optical center of the left camera of the stereo camera as the origin, the X-axis horizontally to the right along the camera, the Y-axis vertically downward, and the Z-axis forward along the camera's optical axis, serving as the reference coordinate system for pose calculation. Furthermore, in practical applications, the left camera can be replaced with the right camera; this invention is not limited to this. The stereo image is two images containing the target charging port, simultaneously captured by the stereo cameras (including two synchronously acquiring RGB cameras, left and right) mounted on the charging device.

[0040] Specifically, the initial pose of the target charging port can be determined using existing charging port pose determination algorithms, which will not be elaborated upon here. For example, the initial pose can be determined through the following steps: A binocular image containing the target charging port is acquired using a binocular image sensor mounted on the vehicle. Then, the charging port is segmented using one of the monocular images from the binocular sensor. The location of the charging port is detected in the segmented image, and interference from obstacles (such as fallen leaves, dust, etc.) is eliminated. Using the PNP algorithm, combined with the relative geometric constraints of the seven fixed terminals in the preset charging port model (the charging port terminals are standardized structures, and the three-dimensional relative positions of the seven terminals are known), the initial pose of the target charging port relative to the camera coordinate system is calculated (denoted as R0 and t0, where R0 is a 3×3 rotation matrix describing the charging port orientation; t0 is a 3×1 translation vector describing the charging port position).

[0041] Step S102: Perform stereo matching on the stereo image to obtain the initial depth map corresponding to the stereo image.

[0042] Specifically, stereo matching obtains depth information by finding matching relationships between corresponding pixels in stereo images and calculating the disparity between pixels (the horizontal distance between corresponding pixels in the left and right images). Existing stereo matching algorithms can be used to convert stereo images into disparity maps, and the depth maps corresponding to the stereo images can be obtained based on the disparity maps. For example, in this embodiment of the invention, the Unimatch depth estimation algorithm is used as an example for explanation. It is a Transformer-based stereo depth estimation algorithm that improves the matching accuracy in low-texture and occluded areas through multi-scale feature enhancement and global matching strategies.

[0043] For example, the specific implementation process of step S102 above includes: Preprocessing of binocular images: First, distortion correction is performed on the images based on the camera distortion coefficients (obtained in advance through calibration) to eliminate image distortion caused by lens optical distortion; second, adaptive histogram equalization (CLAHE) is used to process the images to reduce the impact of uneven lighting in rain, fog, and low-light environments and improve image contrast.

[0044] Subsequently, the Unimatch depth estimation algorithm was used to perform stereo matching: First, multi-scale features (left image feature FL and right image feature FR) of the stereo images were extracted using a Transformer encoder, and position encoding (generated by a trigonometric function, preserving pixel spatial position information) was introduced to enhance the feature representation of low-texture areas (such as the pure black surface of the charging port); Second, the feature similarity matrix S (dimensions H×W×H×W, where H and W are the image height and width, respectively, and S(i,j,k,l) ​​represents the cosine similarity between the left image (i,j) pixel and the right image (k,l) pixel) was calculated, and a matching probability distribution was generated within a local search window (horizontal ±90 pixels, based on the preset charging port size) using the Softmax function (temperature coefficient τ=0.01, controlling the sharpness of the probability distribution), and the point with the highest probability was selected as the corresponding matching point; Third, the disparity map D (D(i,j)=|ik|, k is the horizontal coordinate of the matching point in the right image) was calculated based on the matching point, and then the geometric formula Z= (B×f) / D (where B f is the baseline length of the binocular camera, pre-calibrated to 120mm; f is the camera focal length, pre-calibrated to 800 pixels) converting the disparity map into an initial depth map.

[0045] Step S103: Based on the three-dimensional spatial coordinates of key points in the preset charging port model, the initial depth map and initial pose are alternately optimized and updated until the preset optimization conditions are met, and the optimized initial pose is determined as the target pose of the target charging port.

[0046] The preset charging port model is the same type as the target charging port. This refers to a 3D model that is completely identical to the target charging port type, containing the 3D coordinates of all structures within the charging port. The 3D spatial coordinates of the seven terminals are known key parameters and serve as the benchmark reference for optimization. By iteratively executing the "depth optimization" and "pose optimization" steps, depth deviation and pose error are gradually corrected, achieving coordinated optimization of both.

[0047] For example, the above-mentioned preset optimization conditions include two termination criteria: one is the threshold of the change in pose before and after optimization (such as a change in rotation angle < 0.1°, a change in translation distance < 0.05mm), and the other is the maximum number of iterations (preset to 5 times). Optimization stops when either criterion is met.

[0048] Step S104: Based on the target pose, the charging plug is aligned with the target charging port and then inserted.

[0049] Specifically, based on the target pose (R, t) obtained in step S103, the robot motion control system adjusts the pose of the charging plug to precisely align it with the target charging port. This ensures that the central axis of the charging plug coincides with the central axis of the charging port (position alignment) and the orientation of the charging plug matches that of the charging port (posture alignment), thus completing the insertion action. The robot motion control system is used to control the movement of charging devices, such as the robotic arm of a charging robot. It calculates the motion parameters (such as angle and velocity) of each joint of the robotic arm based on the target pose, achieving precise displacement and posture adjustment of the charging plug.

[0050] This invention determines the initial pose of the charging port and uses an initial depth map obtained through binocular stereo matching. It then performs alternating optimization and updates of the depth map and pose, finally using the optimized target pose to achieve precise alignment between the charging plug and the charging port. This effectively solves the problems of insufficient depth map accuracy and large pose calculation errors caused by environmental interference in traditional charging positioning. By first determining the initial pose and obtaining the initial depth map based on binocular images, and then combining the three-dimensional coordinates of key points of the preset charging port model for alternating optimization of the depth map and pose, it can dynamically correct depth deviation and pose error, avoiding the limitations of single-modal positioning. This allows the final optimized target pose to achieve sub-millimeter-level positioning of the charging port, ensuring precise alignment and insertion of the charging plug. It maintains a high docking success rate even in complex environments such as rain, fog, and low light, significantly improving the environmental robustness and positioning accuracy of automatic charging equipment and meeting the millimeter-level insertion requirements of electric vehicle automatic charging.

[0051] This embodiment also provides a charging control method for a charging device, which is an automatic charging device such as an automatic charging robot. The charging device includes a charging plug for insertion into a charging port on a vehicle to charge the vehicle. This method is applied to the controller of the charging device, such as a microcontroller (MCU) or other control chip. Figure 2 This is a flowchart of a charging control method for a charging device according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps: Step S201: Obtain a first image containing the target charging port.

[0052] Specifically, the image quality of a monocular image containing the target charging port can be directly affected by factors such as ambient light and weather conditions. This image can be captured by the image acquisition device carried by the charging robot (hereinafter referred to as the charging robot) such as the aforementioned binocular camera.

[0053] Step S202: Input the first image into a pre-trained multi-task model based on image quality classification and image quality enhancement to obtain the image quality classification result and the second image after image quality enhancement.

[0054] The image quality classification result is related to the environmental state of the target charging port. For example, the environmental state includes: sunny, rainy, foggy, snowy, low light, etc. Correspondingly, the image quality classification result can also be one of these five categories: sunny, rainy, foggy, snowy, low light. This is just an example and is not a limitation of the present invention.

[0055] Specifically, a pre-trained multi-task model for image quality classification and enhancement is used to simultaneously perform environmental state classification and image quality enhancement on the first image. This addresses the issues of fragmented functionality and low processing efficiency inherent in traditional single-model approaches. Simultaneously, it outputs both the image quality classification result (reflecting the environmental state) and the enhanced second image. Furthermore, the second image, enhanced by the multi-task model, exhibits significant improvements in sharpness, contrast, and noise suppression compared to the first image, and can be directly used for subsequent determination of the charging port target location.

[0056] For example, the model training dataset includes charging port images under five different environmental conditions: sunny, rainy, foggy, snowy, and low light, to ensure the model's generalization ability. The specific training process of the model can be implemented by referring to the existing model training process, and will not be elaborated here. Furthermore, the image quality classification result, i.e., the environmental state label output by the above model, can include eight categories: "sunny," "light rain," "heavy rain," "light fog," "dense fog," "light snow," "heavy snow," and "low light (<10 lux)," to reflect the environmental state of the target charging port and provide a basis for subsequent adjustment of positioning parameters.

[0057] Furthermore, such as Figure 3 As shown, the above multi-task model based on image quality classification and image quality enhancement includes: encoder 301, decoder 302, and image quality classification branch 303. The encoder 301 includes a deep feature extraction network, a channel-space joint attention module, a first convolutional layer, a stacking operation layer, and a first feature dimension transformation layer connected in sequence. The input end of the deep feature extraction network is used to receive the first image. Decoder 302 includes a second feature dimension transformation layer, a concatenation layer, a second convolutional layer and an upsampling layer connected in sequence. The input of the second feature dimension transformation layer is connected to the output of the deep feature extraction network. The concatenation layer is also connected to the output of the first feature dimension transformation layer. The decoder is used to perform image quality enhancement tasks based on the input image features. Image quality classification branch 303 includes a C3T module, a third convolutional layer, and a fully connected layer connected in sequence. The input of the C3T module is connected to the output of the feature dimension transformation layer. The image quality classification branch is used to perform image quality classification tasks based on the input image features.

[0058] This invention addresses the limitations of traditional single-task models and the high latency caused by multiple models being cascaded, through a collaborative design of the encoder, decoder, and image quality classification branch. The encoder sequentially extracts deep image features through a deep feature extraction network and a channel-spatial joint attention module, providing high-quality feature support for subsequent tasks. The decoder performs image quality enhancement based on these features, outputting a clear image. The image quality classification branch classifies environmental states through modules such as C3T. All three components share encoder features, avoiding redundant computation and significantly reducing inference latency. Compared to traditional independent image enhancement and classification models, this structure can perform both tasks simultaneously, improving processing efficiency. Furthermore, the collaborative optimization of each module results in more accurate image quality enhancement and classification results, providing efficient and reliable image preprocessing support for charging positioning.

[0059] In some optional implementations, the above-mentioned multi-task model based on image quality classification and image quality enhancement also includes: A first local attention module is set between the channel-space joint attention module and the first convolutional layer.

[0060] A second local attention module is placed between the second convolutional layer and the upsampling layer.

[0061] This invention addresses the problems of insufficient attention to key features of the charging port area and low task accuracy caused by complex background interference in traditional models by adding first and second local attention modules to the multi-task model. The first local attention module, located between the channel-space joint attention module and the first convolutional layer, enhances the feature representation of key charging port areas (such as terminals) extracted by the encoder and suppresses background noise. The second local attention module, located between the second convolutional layer and the upsampling layer, focuses on optimizing the details of the charging port area during the decoder's image quality improvement process, making the charging port features clearer in the output image. Through the local attention mechanism, the model can focus on key areas, improving the accuracy of image quality classification and the targeted nature of image quality improvement, providing more accurate image information for subsequent positioning, and further improving charging port positioning accuracy.

[0062] Specifically, existing image quality processing models for intelligent charging robots suffer from three major bottlenecks: First, traditional single-task models (either image enhancement or environment classification) are fragmented, requiring multiple models to run in series, resulting in processing delays exceeding 200ms, which cannot meet real-time requirements. Second, the feature extraction process lacks targeted enhancement for key areas of the charging port (such as terminal edges), making it susceptible to interference from background noise and resulting in weak feature discrimination capabilities in rain, fog, and low-light environments. Third, the feature interaction between the encoder and decoder is insufficient, leading to inadequate fusion of high-level semantic features and low-level detail features, resulting in problems such as blurred edges and texture distortion in the enhanced image. Furthermore, existing models do not incorporate local attention mechanisms, resulting in insufficient attention to features of local targets such as the charging port, further limiting the image quality improvement effect and environment classification accuracy, making it difficult to adapt to the stringent requirements of automatic charging scenarios.

[0063] To address the aforementioned issues, the multi-task model provided in this embodiment of the invention adopts an end-to-end architecture of "encoder-decoder-classification branch." The encoder is responsible for extracting multi-scale features of the first image, the decoder performs image quality enhancement based on the encoder features to output the second image, and the image quality classification branch performs environmental state classification based on the encoder features. Simultaneously, first and second local attention modules (i.e., CASS local attention modules) are set at key nodes of the encoder and decoder respectively to strengthen the key features of the charging port. The model input is a 1280×720×3 (RGB) first image, and the output is a classification result of 8 environmental states (such as sunny, heavy rain, dense fog, etc.) and a 1280×720×3 second image. The model inference speed is ≥30fps, meeting real-time processing requirements. For example, the specific structure of the local attention module is as follows... Figure 4 As shown.

[0064] Specifically, in Figure 3The depth feature extraction network of the encoder 301 mentioned above is a convolutional neural network based on ResNet-50. By introducing dilated convolution (dilation rate=2) to replace part of the ordinary convolution, the receptive field is expanded without reducing spatial resolution. It can extract multi-scale features from low-level (edges, texture) to high-level (semantics) features of the image, and uses dilated convolution to extract the depth information features of the image. Dilated convolution can effectively capture multi-scale information in the image while maintaining spatial resolution and avoiding information loss caused by pooling operations. The channel-space joint attention module, abbreviated as CA_ODS module, is used to enhance the discriminative ability of features. By jointly optimizing features through channel attention and spatial attention, important regions and features are highlighted, noise and redundant information are suppressed, and the quality and robustness of features are further improved. The feature channel weights are calculated through channel attention (SENet mechanism) to highlight effective channel features; the spatial attention map is generated through spatial attention (CBAM mechanism) to suppress background noise and achieve accurate feature selection. The first local attention module (CASS) includes an LS Block (Local Spatial Feature Extraction Block) and a TSM (Temporal Shifting Module), which can focus on the local features of the charging port and enhance the expression of key features such as terminal edges and contours. The first convolutional layer uses two 3×3 convolutions with a stride of 1 and padding of 1, further extracting deep semantic information from the output features of the first local attention module. The activation function used is SiLU. It should be noted that... Figure 3 The first local attention module is not shown. Stacked operation layer ( Figure 3 (Abbreviated as stacking): The multi-channel features output from the first convolutional layer are stacked according to channel dimension to increase the number of feature channels and retain richer feature information. For example, 256-channel features are stacked into 512 channels. First feature dimension transformation layer: Uses 1×1 convolution, with one layer, stride = 1, and no padding. It is used to adjust the channel dimension of the stacked features to adapt to the input requirements of subsequent decoder concatenation layers and classification branches. For example, 512-channel features are reduced to 256 channels to further extract deeper semantic features. Dimensionality reduction or expansion of features reduces computational complexity while retaining key information.

[0065] Furthermore, the second feature dimension transformation layer of the decoder 302 above employs a 1×1 convolution, consistent with the structure of the first feature dimension transformation layer. It is used to adjust the low-level feature channel dimension of the deep feature extraction network output, matching it with the number of high-level feature channels in the encoder output, facilitating subsequent splicing and fusion. (Splicing layer) Figure 3(Short for splicing): This involves splicing the high-level features output by the encoder with the low-level features adjusted by the decoder along the channel dimension. The low-level features contain rich detail information, while the high-level features contain semantic information. The combination of the two can achieve a balance between semantics and detail, realizing the fusion of semantic and detail features and improving the texture and edge quality of the enhanced image. Second convolutional layer: Uses 3×3 convolutions, with 3 layers, stride = 1, padding = 1, and the activation function is SiLU. It is used to refine the spliced ​​fused features and eliminate redundant information generated during feature fusion. Second local attention module: Has the same structure as the first local attention module, located between the second convolutional layer and the upsampling layer. It is used to enhance the detail features of the charging port in the fused features (such as terminal texture and edge contours) and avoid detail loss during upsampling. It should be noted that... Figure 3 The second local attention module is not shown. Upsampling layer ( Figure 3 (Abbreviated as upsampling four times): The feature map size is enlarged by using transposed convolution. The transposed convolution kernel size is 4×4, the stride is 2, and the padding is 1. Through multi-stage upsampling, the feature map is restored to the original image resolution, such as (1280×720).

[0066] Furthermore, the C3T module of the image quality classification branch 303 above combines the C3 module (residual bottleneck structure) with the Transformer attention mechanism. It preserves local features (such as fog droplet texture and raindrop trajectory) through residual connections and captures the correlation of global features (such as the overall fog density distribution) through Transformer multi-head self-attention, thereby enhancing the discriminative ability of environmental state features. For example, the specific network structure of the C3T module is as follows: Figure 5 As shown, the activation function of the activation block is the SiLU function.

[0067] The third convolutional layer: uses a 1×1 kernel, has one layer, stride = 1, and no padding. It is used to compress the channel dimension of the C3T module's output features, reducing the computational cost of subsequent fully connected layers. Fully connected layers: contain two fully connected network layers. The first layer is a hidden layer (dimension = 128), using the SiLU activation function; the second layer is the output layer (dimension = 8), corresponding to eight environmental states. It outputs probability values ​​for each class using the Softmax activation function. The class with the highest probability is the image quality classification result (e.g., sunny: 0.01, heavy rain: 0.02, dense fog: 0.95, low light: 0.02). The class with the highest probability, "dense fog," is selected as the image quality classification result, completing the environmental state classification task.

[0068] Furthermore, in practical applications, the aforementioned deep feature extraction network can be replaced with: ordinary convolution + multi-scale pooling (such as the Pyramid Pooling Module, PPM). The main purpose of dilated convolution is to capture multi-scale information, while the Pyramid Pooling Module can also achieve similar functionality and may be more flexible in certain scenarios. The first and second feature dimension transformation layers mentioned above can both be replaced with other dimensionality reduction or expansion operations, such as GroupConvolution or Depthwise Separable Convolution.

[0069] Furthermore, both the first and second local attention modules mentioned above adopt the CASS Block structure, whose core includes the LS Block and TSM modules. The LS Block structure uses a 3×3 depthwise separable convolution (number of kernels = number of input channels, groups = number of input channels), stride = 1, padding = 1, and is coupled with the SiLU activation function to extract local spatial features of the input features, focusing on the charging port region (receptive field size = 31×31, adapted to the size of the charging port in the feature map), and suppressing background noise.

[0070] The TSM module divides the input feature map into three parts along the channel dimension (channel ratio 1:1:1). Two parts are shifted by one pixel horizontally and vertically, respectively, while the third part remains unchanged. These three feature maps are then concatenated along the channel dimension and fused using a 1×1 convolution to enhance the spatial correlation of features and avoid isolated local features. The LS Block effectively extracts local spatial information from the input feature map to compensate for insufficient local modeling capabilities. The CASS Block is the core module of the model. It undergoes a series of processing steps during the input stage, enabling the network to learn deeper and richer feature representations. Batch normalization ensures efficient and stable training and inference processes. The input features first undergo conventional convolution and batch normalization, then enter the LS and CASS modules. Finally, feature fusion and the TSM channel transfer module enhance spatial and temporal features.

[0071] This invention employs an integrated "encoder-decoder-classification branch" architecture to simultaneously improve image quality and classify the environment, reducing processing latency while accelerating inference speed. It also improves environment classification accuracy, addresses the fragmentation issues of traditional single-task models, and meets both real-time and accuracy requirements. Furthermore, by setting up first and second local attention modules (CASS modules) to focus on the charging port area, and through local spatial feature extraction and spatial correlation enhancement, it improves the response values ​​of key features such as the charging port terminal edges and textures by 30%-40%, while reducing background noise response values ​​by 20%-30%. It demonstrates excellent performance in rain, fog, and low-light scenarios.

[0072] For example, as Figure 3 Taking the multi-task model shown as an example, the overall operation process of the model is described as follows: 1. First, each batch of 1280x720 images is entered into the model; 2. Extract depth information features (feature1) from the image using a dilated convolutional network; 3. Feature1 is used to extract features to obtain feature2; 4. Feature 2 is used to obtain feature 3 through 3x3 convolution, stacking, and 1x1 convolution; 5. Feature 3 is obtained by passing through the C3T module, convolution, and fully connected layers to obtain Feature 1, which is used for image quality classification; 6. Feature 1 is converted into feature 4 through a 1x1 convolution operation; 7. Feature 4 and feature 3, which has been upsampled by 4 times, are concatenated to obtain feature 5; 8. Feature5 is generated into a high-resolution image by performing a 3x3 convolution and upsampling by 4 times.

[0073] Step S203: Based on the second image, determine the target location of the target charging port.

[0074] The target location refers to the pixel coordinates (u,v) of the charging port center in the image coordinate system, and the actual distance (d) from the charging port center to the camera. The combination of these two can completely describe the position of the charging port in three-dimensional space, providing a basis for subsequent movement of the charging plug.

[0075] Specifically, the target location of the charging port can be determined by a fusion algorithm of "semantic segmentation + target detection", which solves the problem of large positioning error of traditional single detection algorithm in complex environment and ensures the accuracy of target location.

[0076] For example, the charging port region in an image can be accurately segmented using the Deeplabv3 + semantic segmentation model, and a mask of the charging port region can be output (only the charging port pixels are retained, and the background pixels are set to 0). The pixel center coordinates of the charging port region can be calculated using the mask. Then, the bounding box coordinates of the charging port can be obtained using the YOLOv8 object detection model, and the center pixel coordinates of the bounding box can be calculated. Finally, the pixel center coordinates of the charging port region obtained by semantic segmentation and the center pixel coordinates of the bounding box obtained by object detection are weighted and fused to obtain the final pixel coordinates of the charging port center. The actual distance from the charging port center to the camera can be calculated using the pixel size of the charging port in the second image and the camera intrinsic parameters. The final pixel coordinates of the charging port center, i.e., its actual distance to the camera, are determined as the target position of the target charging port.

[0077] Step S203: Based on the target position, control the charging plug to move to the preset position of the target charging port.

[0078] The preset position is a location at a set distance from the target position. This distance is designed to avoid collisions between the plug and the vehicle while providing a sufficiently close observation range for subsequent fine positioning; for example, this set distance is 200mm. By converting the target position from the image coordinate system to the charging robot's base coordinate system, and then moving the charging plug towards the preset position until it reaches a distance of 200mm from the target charging port, the coarse positioning of the target charging port is completed.

[0079] This invention addresses the problems of fragmented environmental perception and image enhancement in traditional positioning, as well as positioning failures caused by poor image quality in complex environments, by introducing a multi-task model to process the first image during the initial charging positioning phase. The multi-task model simultaneously outputs image quality classification results (reflecting the environmental state) and an enhanced second image, clearly defining the current environment (e.g., rain, fog) while eliminating environmental interference, thus providing a high-quality image foundation for subsequent target location determination of the charging port. By determining the target location based on the second image and controlling the charging plug to move to a preset position, misjudgments of the target location due to blurry or low-contrast original images are avoided, significantly improving the accuracy of target positioning in complex environments and laying the foundation for subsequent precise docking.

[0080] Step S204: After controlling the charging plug to move to the preset position of the target charging port, determine the initial pose of the target charging port and acquire a stereo image containing the target charging port. The initial pose is the pose of the target charging port relative to the camera coordinate system corresponding to the stereo image. See the above for details. Figure 1 The relevant descriptions of step S101 shown will not be repeated here.

[0081] Step S205: Perform stereo matching on the binocular images to obtain the initial depth map corresponding to the binocular images. See the above for details. Figure 1 The relevant description of step S102 shown will not be repeated here.

[0082] Step S206: Based on the three-dimensional spatial coordinates of key points in the preset charging port model, the initial depth map and initial pose are alternately optimized and updated until the preset optimization conditions are met, and the optimized initial pose is determined as the target pose of the target charging port; the preset charging port model and the target charging port are of the same type.

[0083] Specifically, the aforementioned preset optimization conditions include: the change in the initial pose before and after optimization meets the preset pose change condition or the number of iterations reaches the preset iteration threshold.

[0084] This invention addresses the problem of independent depth map and pose calculation, which easily leads to error accumulation, in traditional positioning by first correcting the depth map, then using the corrected depth map for pose optimization, and finally performing a closed-loop operation with severe repetition. Each alternating optimization cycle is based on the preset 3D coordinates of key points on the charging port model. The initial depth map is first corrected to eliminate depth deviations caused by environmental interference, and then the initial pose is optimized based on the corrected depth map. Furthermore, the optimization termination is controlled by preset pose change conditions and iteration number thresholds, ensuring optimization accuracy while avoiding system delays caused by excessive iteration. Compared to traditional single-step positioning schemes, this approach can progressively compress positioning errors, significantly improving the final pose accuracy while maintaining system real-time performance to meet the rapid response requirements of charging devices.

[0085] Furthermore, in step S206 above, based on the three-dimensional spatial coordinates of key points in the preset charging port model, the initial depth map and initial pose are alternately optimized and updated, including: Step a1: Based on the initial pose, map the three-dimensional spatial coordinates of key points in the preset charging port model to the camera coordinate system, and project them onto the image plane of the camera coordinate system to obtain the projection point coordinates and theoretical depth corresponding to the key points.

[0086] The preset charging port model is a 3D digital model that is completely identical in structure to the target charging port (such as a Type-C charging port), containing the 3D spatial coordinates of 7 fixed terminals (key points) of the charging port. By obtaining the 3D spatial coordinates of the 7 key points in the preset charging port model, the key point coordinates are mapped to the camera coordinate system, and their Z-axis coordinates in the camera coordinate system represent the theoretical depth of the key point. Then, the 3D coordinates of the key points in the camera coordinate system are converted into 2D pixel coordinates using a projection function and the camera intrinsic parameter matrix to obtain the aforementioned projection point coordinates. The process of transforming the 3D spatial coordinates from the model coordinate system to the camera coordinate system and further projecting them onto the image plane of the camera coordinate system is existing technology and will not be described in detail here.

[0087] Step a2: Based on the theoretical depth corresponding to the key point, correct the first depth corresponding to the key point in the initial depth map to obtain the corrected depth map.

[0088] The initial depth map consists of two images taken by a stereo camera, showing the distances from each pixel to the camera. This distance is calculated using stereo matching. The distances to key points in the initial depth map represent the first depth, but this first depth may have errors (e.g., rain or fog can cause inaccurate stereo matching, leading to an overestimation or underestimation of the first depth). The theoretical depth calculated in step a1 is obtained by transforming the 3D spatial coordinates of key points in a preset charging port model. Therefore, the accuracy of this theoretical depth is highly reliable. Taking the seven fixed terminals of the charging port as key points as an example, the depth value corresponding to each fixed terminal in the initial depth map can be extracted, which is the first depth. This high-precision theoretical depth is then used to correct the first depth, which may have depth errors in the initial depth map. The corrected depth values ​​of all key points are then updated in the initial depth map to obtain the corrected depth map, ensuring that the depth value of each key point in the corrected depth map is accurate and reducing depth estimation errors.

[0089] Specifically, by using the theoretical depth as a benchmark, the first depth corresponding to key points in the initial depth map is corrected, eliminating depth deviations caused by matching noise and environmental interference in the initial depth map, generating a more accurate corrected depth map, and providing reliable depth data support for subsequent pose optimization. Specifically, the measured depth (first depth) affected by environmental interference is corrected to a depth closer to the true value. For example, in rain and fog scenarios, the first depth of key points in the initial depth map may deviate by 5mm; after correction, the deviation can be reduced to within 0.5mm, eliminating depth measurement errors. Furthermore, since the key points are the core area of ​​the charging port, correcting the depth of the key points can improve the accuracy of the entire depth map (for example, by using interpolation to spread the correction logic of the key points to surrounding pixels), providing reliable depth data for more accurate pose calculations and improving the overall accuracy of the depth map. Even in complex environments such as low light, rain, and fog, the theoretical depth serves as a benchmark, avoiding the problem of inaccurate depth maps, solving the pain point of traditional binocular depth estimation where accuracy collapses in poor environments, and enhancing environmental robustness.

[0090] Furthermore, step a2 above obtains the corrected first depth value by weighted summation of the theoretical depth and the first depth.

[0091] This invention corrects depth values ​​by weighting the theoretical depth and the first depth in the initial depth map. The weights are dynamically allocated based on the reliability of both. For example, when the reliability of the first depth is low due to environmental interference (such as fog), the weight of the theoretical depth is increased; conversely, the weight of the first depth is increased. This flexible weighting strategy fully utilizes effective depth information, reduces the impact of single-depth data deviations on positioning, and makes the corrected depth value more closely match the actual scene. This provides a precise depth foundation for subsequent pose optimization, further improving the positioning accuracy of the charging port and ensuring the accurate docking of the charging plug.

[0092] Step a3: Based on the depth error between the key points in the initial depth map and the corrected depth map, and the reprojection error between the image coordinates of the key points in the monocular image and the coordinates of the projection points of the key points, calculate the pose change. The monocular image is one of the images in the binocular image.

[0093] Among them, depth error refers to the difference between the first depth of the same key point in the initial depth map and the depth in the corrected depth map (reflecting the change before and after depth correction, and also reflecting the magnitude of the initial depth error); the image coordinates corresponding to the key point in the monocular image are the 2D positions of the key point actually found in this monocular image; the reprojection error is the distance between this actual image coordinate and the projection point coordinates calculated in step a1 above (the theoretical position it should be in) (for example, the actual position is at (100,200) pixels in the image, and the projection point is at (102,201) pixels, and the reprojection error is the distance between these two points).

[0094] Then, by combining the "depth error" and "reprojection error," the pose change is calculated. This pose change reflects how much the initial pose needs adjustment (e.g., rotating the angle by 0.1° or translating the distance by 0.2mm) to achieve more accurate depth and make the key points in the image more closely match their theoretical positions. By simultaneously considering both "image position error" and "depth error," the problem of "looking correct but actually having the wrong distance" caused by traditional positioning methods that only consider a single error (e.g., only looking at image position and ignoring depth) is avoided, allowing for more comprehensive pose adjustment. A precise pose change is calculated by quantifying the pose adjustment magnitude, rather than relying on experience. This ensures that each pose optimization has a precise direction and magnitude; for example, it can accurately calculate the need to rotate the charging port's pose by 0.3° and translate it by 0.5mm, rather than making approximate adjustments based on experience.

[0095] Specifically, by constructing an objective function for pose change that integrates depth error and reprojection error, and using a nonlinear optimization algorithm to solve for pose change, we can achieve accurate correction of the initial pose and solve the problem of one-sided pose optimization caused by traditional single error index.

[0096] Step a4: Obtain the optimized initial pose based on the initial pose and pose change, and return to execute step a1 above.

[0097] Specifically, the optimized initial pose can be obtained by adjusting the pose based on the pose change. By adjusting the initial pose according to the pose change, a more accurate charging port pose is obtained; this new pose is the optimized initial pose.

[0098] Furthermore, since the pose is composed of a rotation matrix R and a translation vector t, the specific optimization process involves updating these two parameters. Specifically, the rotation matrix is ​​updated as follows: the initial rotation matrix is... R 0, the rotation adjustment in the pose change corresponds to a small rotation matrix Δ R Optimized rotation matrix Rnew = R 0×Δ R (Rotation matrix multiplication represents rotation superposition). Translation vector update: The initial translation vector is... t 0, the translation adjustment in the pose change is Δ t Optimized translation vector tnew = t 0+Δ t (The addition of translation vectors represents the superposition of positions).

[0099] After the previous deep correction and multi-error optimization, the optimized pose can significantly reduce the error of the initial pose. For example, the initial pose may have a translation error of ±5mm and a rotation error of ±1°, which can be reduced to a translation error of ±0.5mm and a rotation error of ±0.3° after optimization, achieving "sub-millimeter" positioning accuracy. This is exactly what electric vehicle automatic charging requires, otherwise it will not be able to be plugged in or the interface will be damaged.

[0100] The optimized pose is the final instruction for adjusting the charging plug's position, providing a precise basis for final docking. Based on this accurate positioning, the charging device can precisely control the plug's movement, ensuring perfect alignment between the plug and the charging port, thus solving the docking failure problem caused by insufficient positioning accuracy in traditional methods. Furthermore, if the error after one optimization is still insufficient, the optimized initial pose can be used as the new initial pose, repeating the aforementioned depth correction, error calculation, and pose adjustment steps until the error is small enough to meet charging requirements (for example, sub-millimeter accuracy can be achieved after 3-5 iterations), further ensuring positioning reliability in complex environments.

[0101] This invention maps the 3D spatial coordinates of key points in a preset charging port model to the camera coordinate system to obtain the theoretical depth, thereby correcting the initial depth map and resolving depth errors caused by texture loss and occlusion in binocular matching. Furthermore, by combining depth error and reprojection error to calculate pose change, it comprehensively considers deviations in both depth and image coordinates, avoiding pose deviations caused by a single error metric. This process can accurately locate the charging port, reducing docking failures caused by inaccurate depth and coordinate offsets. It maintains high positioning accuracy even in scenarios with partial occlusion of the charging port and weak texture, improving the reliability of charging docking.

[0102] Furthermore, the method provided in this embodiment of the invention further includes the following steps: Step b1: Obtain the confidence level of the first depth corresponding to the key point in the initial depth map.

[0103] Specifically, the confidence level of the first depth of key points can be extracted from the confidence map output by the Unimatch depth estimation algorithm, providing a quantitative basis for subsequent dynamic weight allocation and solving the problem of subjective weight setting in traditional methods.

[0104] For example, the Unimatch depth estimation algorithm outputs a reliability assessment of the depth value of each pixel in the initial depth map, with a value range of [0,1]. The closer the value is to 1, the more reliable the depth value is (such as the edge area of ​​the charging port terminal), and the closer the value is to 0, the less reliable the depth value is (such as metal reflection, rain and fog obscured areas). It is calculated and generated by the algorithm through indicators such as feature matching consistency and disparity continuity, which will not be elaborated here.

[0105] Step b2: Determine the first weight corresponding to the first depth based on the confidence level of the first depth. The first weight is positively correlated with the confidence level.

[0106] Step b3: Determine the second weight corresponding to the theoretical depth based on the first weight.

[0107] Specifically, the weights of the first depth and the theoretical depth can be dynamically allocated based on the obtained depth confidence. The higher the depth confidence, the greater the weight of the first depth and the smaller the weight of the theoretical depth. The sum of the first weight and the second weight is 1. For example, the Sigmoid function can be used as the mapping relationship between the depth confidence and the first weight to solve the problem of insufficient accuracy caused by the traditional fixed weight correction of depth, so that the depth correction is more in line with the actual reliability of the data.

[0108] Step b4: Determine the calculation weight corresponding to the depth error based on the confidence level of a depth. The calculation weight is positively correlated with the confidence level.

[0109] Specifically, by dynamically adjusting the weight of depth error in calculating pose change based on depth confidence, it ensures that high-confidence depth data dominates optimization and the influence of low-confidence data is weakened, thus solving the optimization bias problem caused by traditional fixed weights.

[0110] For example, the relationship between confidence level and computational weight can be determined by piecewise linear mapping. For instance, when the confidence level is greater than or equal to 0.8 (high confidence), the corresponding computational weight can be set to 1, indicating that the depth error fully participates in the optimization and dominates the depth consistency constraint. When the confidence level is greater than or equal to 0.5 and less than 0.8 (medium confidence), the corresponding computational weight can increase linearly according to a linear function. For example, when the confidence level is 0.66, the computational weight is 0.72. When the confidence level is less than 0.5 (low confidence), the corresponding computational weight can be set to 0.4 to reduce the interference of low-reliability depth error on the optimization. This ensures the dominance of high-confidence data while avoiding the complete failure of low-confidence data, thus balancing optimization accuracy and robustness.

[0111] This invention addresses the problems of traditional weight allocation, such as lack of reliable data and strong subjective assumptions, by introducing depth confidence scores corresponding to key points in the initial depth map. By obtaining the confidence score of the first depth, the first weight is positively correlated with the confidence score, ensuring that the first depth with high confidence plays a greater role in depth correction. Simultaneously, the second weight of the theoretical depth and the calculation weight of the depth error are reasonably determined. For example, in the reflective metal area of ​​the charging port, the first depth confidence score is low, corresponding to a small first weight and a small depth error calculation weight, reducing the interference of unreliable depth data in this area on optimization. This dynamic weight adjustment strategy based on confidence scores allows the optimization process to better reflect the actual data quality, improving the accuracy of depth correction and pose optimization, and enhancing the system's adaptability to complex environments.

[0112] For example, the pose change is calculated using the following formula (1): (1) in, Indicates the change in pose. Indicates the first The image coordinates of each key point in a monocular image. Indicates the first The coordinates of the projection points corresponding to each key point Indicates the first The three-dimensional spatial coordinates of the key points R This indicates that the rotation matrix is ​​used to describe the orientation of the target charging port. t This indicates that the translation vector is used to describe the position of the target charging port. Represents the projection function. This represents the camera intrinsic parameter matrix. Indicates the first The calculation weights corresponding to the depth errors of each key point Indicates the first The depth values ​​corresponding to the key points in the initial depth map. Indicates the first The depth values ​​corresponding to the key points in the corrected depth map.

[0113] This invention maps the 3D spatial coordinates of key points in a preset charging port model to the camera coordinate system to obtain the theoretical depth, thereby correcting the initial depth map and resolving depth errors caused by texture loss and occlusion in binocular matching. Furthermore, by combining depth error and reprojection error to calculate pose change, it comprehensively considers deviations in both depth and image coordinates, avoiding pose deviations caused by a single error metric. This process can accurately locate the charging port, reducing docking failures caused by inaccurate depth and coordinate offsets. It maintains high positioning accuracy even in scenarios with partial occlusion of the charging port and weak texture, improving the reliability of charging docking.

[0114] For example, a detailed implementation scheme of an auxiliary localization algorithm based on the binocular depth estimation algorithm Unimatch and PNP (Perspective-n-Point) is presented. This scheme significantly improves localization accuracy and robustness in complex environments by fusing dense depth estimation and keypoint pose calculation.

[0115] 1. System input and preprocessing.

[0116] Input data: Synchronously acquired stereo RGB image pairs (Left image) (Right figure). Camera intrinsic parameter matrix K and baseline length B (obtained through calibration).

[0117] Preprocessing: Image distortion correction: Image correction based on camera distortion coefficients.

[0118] Illumination normalization: Adaptive histogram equalization (CLAHE) is applied to reduce the impact of illumination variations.

[0119] 2. Unimatch Depth Estimation Module.

[0120] Core steps: (1) Feature Enhancement: Multi-scale features are extracted using the TransformerEncoder to enhance the response to low-texture regions. (Left image: Features) and features of the right figure The generating formula: (2) Positional embedding is introduced to preserve spatial information. Positional embedding is fused with the Transformer encoder through the following steps. Its core goal is to inject spatial positional information into the unordered Transformer features, especially enhancing the model's understanding of spatial relationships in low-texture regions. The specific implementation process is as follows: Position codes are typically generated using learnable parameterization methods or fixed-form trigonometric functions; The fusion of positional encoding and features, specifically the fusion of positional encoding and the Transformer encoder, typically occurs during the input embedding stage. The specific steps are as follows: Feature extraction: Input left and right images and Basic features are extracted using a convolutional backbone network (such as ResNet-50) to obtain the initial feature map. and Multi-scale processing (such as pyramid pooling or downsampling) is performed on the feature map to generate multi-scale features. and .

[0121] Add position encoding: In the input layer of the Transformer encoder, add position encoding. PE Add directly to multi-scale features: (3) If positional encoding is learnable, then and These are independent learning parameters. If the positional encoding is fixed (such as a trigonometric function), it is directly superimposed onto the feature map.

[0122] Transformer encoder processing: enhancing features and The input is a Transformer encoder, which performs self-attention computation and feature interaction. Self-attention mechanism: Captures long-distance dependencies by calculating the similarity between features. Multi-head attention: Uses multiple attention heads in parallel to improve the model's ability to model different spatial relationships.

[0123] The role and advantages of positional encoding: Enhancing the response of low-texture regions: In charging port positioning scenarios, the surface of the charging port is often pure black or metallic, with a simple texture. Traditional convolutional networks struggle to extract effective features in this case, while positional encoding, by explicitly injecting spatial information, helps the model understand the relative positions of different regions in the feature map, improving the feature discrimination ability of low-texture regions. Improving the consistency of multi-scale features: When fusing multi-scale features, positional encoding ensures that features at different scales have a unified spatial reference system, avoiding matching errors caused by scale differences. Supporting geometric constraints: In subsequent disparity calculations, the spatial information implicit in positional encoding helps the model more accurately predict the displacement relationship between pixels, reducing the impact of occlusion and mismatches.

[0124] Global matching and cost volume construction: Calculating the feature similarity matrix It is a four-dimensional tensor, in which the first two dimensions ( H , W ) represents the positions of all pixels in the left image, and the last two dimensions ( H , W () represents the positions of all pixels in the right image.

[0125] (4) in, The left figure is located in space ( i , j The feature vector at the location (such as the feature output by the Transformer encoder). The right figure is located in space ( k , l The eigenvector at (). The position shown in the left figure ( i , j ) and the position of the right figure ( k , l The cosine similarity of the two pairs represents the confidence level of their match.

[0126] The probability distribution is generated using Softmax, and the point with the highest probability is selected as the matching point. The purpose of Softmax is to optimize the similarity matrix. S The specific steps to convert to a probability distribution are as follows: Local window constraints: Applying Softmax directly to the entire four-dimensional tensor would result in an explosion of computational complexity (e.g., for a 512×512 image). S The size of the tensor is =68,719,476,736, In actual implementation, a local search window is usually introduced, that is, for each position in the left figure ( i , jOnly candidate locations within its parallax range in the right figure are considered. k , l ).

[0127] Softmax normalization: Within a local window, similarity values Apply the Softmax function to generate the probability distribution: (5) in, τ It is the temperature scaling factor, which controls the "sharpness" of the probability distribution: a smaller value indicates a higher probability distribution. τ (like τ A value of 0.01 will concentrate the probability at the position of highest similarity, increasing certainty. A larger value... τ It will smooth the probability distribution and increase the robustness of matching. P ( i , j , k , l The left image indicates the location of the left figure. i , j ) and the position in the right figure ( k , l The probability of a match. Softmax normalization ensures that: (6) That is, the sum of the matching probabilities for each position in the left image is 1.

[0128] Parallax and Depth Calculation: Parallax Map D Generated from matching offsets: D=argmaxS The depth map Z is obtained through geometric transformation: (7) in, Z This is a depth map, representing the distance from each pixel to the camera. B The baseline length of the binocular camera (distance between the optical centers of the two cameras). f This refers to the camera's focal length. D This is a disparity map.

[0129] Optimization strategy: Multi-level refinement, iteratively optimizing residual parallax at 1 / 8 and 1 / 4 resolutions to improve edge details. Occlusion handling, utilizing left-right consistency check to mask occluded areas.

[0130] 3. Feature point extraction and matching.

[0131] Key point selection: Feature points are filtered using geometric stability criteria: planar or cylindrical patches (such as SIFT / SURF descriptors) are extracted, excluding regions with monolithic textures. Highly stable points (such as corners and edge intersections) are retained through curvature analysis. Descriptor matching: for the key point {pi} in the left image, the search range is narrowed in the right image by depth constraints: the ORB descriptor is matched using Hamming distance, and mismatches are eliminated using RANSAC.

[0132] 4. PNP pose calculation module.

[0133] The PNP algorithm is used to solve for the pose and obtain the initial pose. The specific implementation process is existing technology and will not be described in detail here. 5. Depth-pose joint optimization, for details please refer to the specific description of step S206 above, and will not be repeated here.

[0134] Step S207: Based on the target pose, align the charging plug with the target charging port and insert it. See the above for details. Figure 1 The relevant description of step S104 shown will not be repeated here.

[0135] In some optional embodiments, the charging control method for the charging device provided in this embodiment of the invention further includes the following steps before performing step S205: Step c1: Input the stereo images into the pre-trained multi-task model based on image quality classification and image quality enhancement to obtain the image quality classification result and the stereo images after image quality enhancement.

[0136] The pre-trained multi-task model based on image quality classification and image quality enhancement is described above and will not be repeated here.

[0137] This invention addresses the problems of large matching errors and low initial depth map accuracy caused by poor original image quality (such as blurring due to rain or fog) in traditional binocular matching by performing multi-task model processing on the binocular images before stereo matching. The multi-task model improves the quality of the binocular images, eliminates the influence of environmental interference on the left and right eye images, and ensures consistency and clarity between the binocular images, providing high-quality input for subsequent stereo matching. Simultaneously, the image quality classification results can help determine the current environment, providing a basis for adjusting subsequent optimization strategies. Compared to directly matching the original binocular images, this processing significantly reduces matching noise, improves the accuracy of the initial depth map, provides a more reliable depth foundation for subsequent alternating optimization updates, and further enhances the accuracy and stability of charging port positioning.

[0138] In practical applications, the main processes of automatic charging of charging equipment include: 1. When the charging plug is in the initial position, a multi-task model based on image quality classification and image quality enhancement identifies the current environment (normal, rainy, foggy, snowy, etc.) and performs noise reduction processing to output a high-quality, clear image. For example, Figure 6 This is a comparison chart of image denoising effects.

[0139] 2. The charging port is first identified using the Deeplabv3+ semantic segmentation algorithm and the YOLOv8 object detection algorithm, and it is determined whether there are any obstacles obstructing it. If there are obstacles obstructing it, a voice prompt is given.

[0140] 3. Coarse positioning: By combining the PNP algorithm with the fixed relative positions of the seven terminals of the charging port, we obtain the approximate position of the charging port and move the charging plug to a position approximately 200mm away from the charging port.

[0141] 4. Further locate the charging port at a nearby position. Due to visual positioning errors in the PNP algorithm, a Unimatch depth estimation algorithm using a binocular camera is combined with the PNP algorithm to obtain the precise location of the charging port. For example, Figure 7 This is a diagram showing the results of binocular depth estimation.

[0142] 5. Move the charging plug to the final alignment position, and slowly move the charging plug along the z-axis; with the cooperation of the pressure sensor, the charging plug is finally successfully inserted into the charging port. For example, Figure 8 A schematic diagram of the main process for automatically charging charging devices.

[0143] This invention, through joint learning of environment state classification and downstream task-driven image enhancement, improves the peak signal-to-noise ratio (PSNR) of the output image by over 8dB and the structural similarity index (SSIM) by over 0.15 in rain, fog, and snow scenes. Furthermore, by embedding a local attention module, it achieves cross-task feature interaction between environment perception and image enhancement, increasing inference speed by 40% compared to traditional Transformers while ensuring real-time performance. In the coarse localization stage: semantic segmentation (Deeplabv3+) and object detection (YOLOv8) are fused to construct a geometrically constrained enhanced PNP algorithm, utilizing the seven terminals of the charging port to fix the spatial relationship and compress the initial localization error to ±20mm. In the fine localization stage: a Unimatch-PNP joint optimization mechanism is proposed, fusing the binocular depth map with the PNP geometric prior depth, achieving sub-millimeter-level pose estimation (translation error <0.5mm, angle error <0.3°) within a 200mm distance, breaking through the limit of insertion accuracy. Dynamically adjusting the image enhancement strategy and localization algorithm parameters reduces the system response latency to 80ms.

[0144] The charging control scheme for the charging device provided by this invention has rain and fog penetration imaging capability: reconstructing high-frequency textures in rain and fog transmittance <50% and outputting enhanced images with MTF >0.2 (compared to MTF≈0 in traditional schemes). Dynamic suppression of metal reflection: capturing the polarization angle distribution of the metal surface through a cross-task attention mechanism reduces the overexposure rate of reflective areas from 78% to 9%. Industrial-grade verification shows that the stainless steel charging port achieves 100% successful positioning at a 60° strong light incident angle. Cross-scale continuous positioning: accuracy is compressed without breaks. In the coarse positioning stage, geometric constraint PNP compresses the initial error of 5 meters from >35mm to ±20mm. In the fine positioning stage, Unimatch-PNP joint optimization achieves 0.5mm positioning within a 200mm distance. By fusing binocular subpixel displacement and geometric prior, the diffraction limit resolution is broken. System entropy reduction closed-loop architecture: a dynamic feedback mechanism, real-time optimization of environmental perception, enhancement, and positioning parameters (80ms / cycle), error chain blocking, image enhancement module reduces segmentation error by 60%, and pose transfer error rate is compressed from >200% to <30%. It also has the ability to adapt to the entire physical environment, and the local attention mechanism enables the inference latency of the multi-task model to be only 35ms (40% faster than global attention). The hardware resource reuse rate is improved by 3 times: environment recognition and image enhancement share 90% of the computation graph.

[0145] This embodiment also provides a charging control device for a charging device, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0146] This invention provides a charging control device for a charging equipment, the charging equipment including: a charging plug, such as... Figure 9 As shown, the device includes: The first processing module 1101 is used to determine the initial pose of the target charging port after controlling the charging plug to move to the preset position of the target charging port, and to acquire a binocular image containing the target charging port. The initial pose is the pose of the target charging port relative to the camera coordinate system corresponding to the binocular image. The second processing module 1102 is used to perform stereo matching on the stereo image to obtain the initial depth map corresponding to the stereo image. The third processing module 1103 is used to alternately optimize and update the initial depth map and initial pose based on the three-dimensional spatial coordinates of key points in the preset charging port model until the preset optimization conditions are met, and then determine the optimized initial pose as the target pose of the target charging port; the preset charging port model and the target charging port are of the same type. The fourth processing module 1104 is used to control the charging plug to be aligned with the target charging port and then inserted based on the target pose.

[0147] In some optional implementations, the preset optimization conditions include: the change in pose of the initial pose before and after optimization meets the preset pose change condition or the number of iterations reaches the preset number of iterations threshold.

[0148] In some optional implementations, the third processing module 1103 includes: The first processing subunit is used to map the three-dimensional spatial coordinates of key points in the preset charging port model to the camera coordinate system based on the initial pose, and project them onto the image plane of the camera coordinate system to obtain the projection point coordinates and theoretical depth corresponding to the key points. The second processing subunit is used to correct the first depth corresponding to the key point in the initial depth map based on the theoretical depth corresponding to the key point, so as to obtain the corrected depth map. The third processing subunit is used to calculate the pose change based on the depth error between the key points in the initial depth map and the corrected depth map, and the reprojection error between the image coordinates of the key points in the monocular image and the projection point coordinates of the key points. The monocular image is one of the images in the binocular image. The fourth processing subunit is used to obtain the optimized initial pose based on the initial pose and pose change, and then call the first subunit to run.

[0149] In some optional implementations, the second processing unit is specifically used to perform a weighted summation of the theoretical depth and the first depth to obtain a corrected first depth value.

[0150] In some optional embodiments, the charging control device of the above-mentioned charging equipment further includes: The second acquisition module is used to acquire the confidence level of key points in the initial depth map corresponding to the first depth. The ninth processing module is used to determine the first weight corresponding to the first depth based on the confidence level of the first depth, and the first weight is positively correlated with the confidence level; The tenth processing module is used to determine the second weight corresponding to the theoretical depth based on the first weight; And / or, the eleventh processing module is used to determine the computational weight corresponding to the depth error based on the confidence level of a depth, wherein the computational weight is positively correlated with the confidence level.

[0151] In some alternative implementations, the pose change is calculated using the following formula:

[0152] in, Indicates the change in pose. Indicates the first The image coordinates of each key point in a monocular image. Indicates the first The coordinates of the projection points corresponding to each key point Indicates the first The three-dimensional spatial coordinates of the key points R This indicates that the rotation matrix is ​​used to describe the orientation of the target charging port. t This indicates that the translation vector is used to describe the position of the target charging port. Represents the projection function. This represents the camera intrinsic parameter matrix. Indicates the first The calculation weights corresponding to the depth errors of each key point Indicates the first The depth values ​​corresponding to the key points in the initial depth map. Indicates the first The depth values ​​corresponding to the key points in the corrected depth map.

[0153] In some optional embodiments, the charging control device of the charging equipment provided in this invention further includes: The first acquisition module is used to acquire a first image containing the target charging port; The fifth processing module is used to input the first image into a pre-trained multi-task model based on image quality classification and image quality enhancement to obtain the image quality classification result and the second image after image quality enhancement. The image quality classification result is related to the environmental state of the target charging port. The sixth processing module is used to determine the target location of the target charging port based on the second image; The seventh processing module is used to control the charging plug to move to a preset position of the target charging port based on the target position. The preset position is a position at a set distance from the target position.

[0154] In some optional implementations, the above-mentioned multi-task model based on image quality classification and image quality enhancement includes: an encoder, a decoder, and an image quality classification branch. The encoder includes a deep feature extraction network, a channel-space joint attention module, a first convolutional layer, a stacking operation layer, and a first feature dimension transformation layer connected in sequence. The input end of the deep feature extraction network is used to receive the first image. The decoder includes a second feature dimension transformation layer, a concatenation layer, a second convolutional layer, and an upsampling layer connected in sequence. The input of the second feature dimension transformation layer is connected to the output of the deep feature extraction network, and the concatenation layer is also connected to the output of the first feature dimension transformation layer. The decoder is used to perform image quality enhancement tasks based on the input image features. The image quality classification branch consists of a C3T module, a third convolutional layer, and a fully connected layer connected in sequence. The input of the C3T module is connected to the output of the feature dimension transformation layer. The image quality classification branch is used to perform image quality classification tasks based on the input image features.

[0155] In some optional implementations, the above-mentioned multi-task model based on image quality classification and image quality enhancement also includes: A first local attention module is set between the channel-space joint attention module and the first convolutional layer; A second local attention module is placed between the second convolutional layer and the upsampling layer.

[0156] In some optional embodiments, the charging control device of the charging equipment provided in this invention further includes: The eighth processing module is used to input the stereo images into the pre-trained multi-task models based on image quality classification and image quality enhancement, respectively, to obtain the image quality classification results and the stereo images after image quality enhancement.

[0157] The further functional descriptions of the above modules and units are the same as those in the corresponding method embodiments described above, and will not be repeated here.

[0158] This invention also provides a charging device for automatically charging electric vehicles, the charging device including a charging plug and a controller. Figure 10 As shown, the controller includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions that execute within the computer device, including instructions stored in or on memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 10 Take a processor 10 as an example.

[0159] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.

[0160] The memory 20 stores instructions executable by at least one processor 10 to cause at least one processor 10 to perform the method shown in the above embodiments.

[0161] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device as shown by a landing page for an app. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, which can be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0162] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0163] The controller also includes a communication interface 30 for the vehicle to communicate with other devices or communication networks.

[0164] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.

[0165] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0166] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A charging control method for a charging device, the charging device comprising: A charging plug, characterized in that the method includes: After controlling the charging plug to move to the preset position of the target charging port, the initial pose of the target charging port is determined, and a binocular image containing the target charging port is acquired. The initial pose is the pose of the target charging port relative to the camera coordinate system corresponding to the binocular image. Perform stereo matching on the binocular images to obtain the initial depth map corresponding to the binocular images; Based on the three-dimensional spatial coordinates of key points in the preset charging port model, the initial depth map and the initial pose are alternately optimized and updated until the preset optimization conditions are met, and the optimized initial pose is determined as the target pose of the target charging port; the preset charging port model is of the same type as the target charging port. Based on the target pose, the charging plug is aligned with the target charging port and then inserted. The initial depth map and the initial pose are alternately optimized and updated based on the three-dimensional spatial coordinates of key points in the preset charging port model, including: Based on the initial pose, the three-dimensional spatial coordinates of key points in the preset charging port model are mapped to the camera coordinate system and projected onto the image plane of the camera coordinate system to obtain the projection point coordinates and theoretical depth corresponding to the key points; Based on the theoretical depth corresponding to the key point, the first depth corresponding to the key point in the initial depth map is corrected to obtain the corrected depth map. Based on the depth error between the key points in the initial depth map and the corrected depth map, and the reprojection error between the image coordinates of the key points in the monocular image and the projection point coordinates of the key points, the pose change is calculated. The monocular image is one of the images in the binocular image. The process involves obtaining an optimized initial pose based on the initial pose and the pose change, and then mapping the three-dimensional spatial coordinates of key points in the preset charging port model to the camera coordinate system based on the initial pose, and projecting them onto the image plane of the camera coordinate system to obtain the projection point coordinates and theoretical depth of the key points.

2. The method according to claim 1, characterized in that, The preset optimization conditions include: the change in the initial pose before and after optimization meets the preset pose change condition or the number of iterations reaches the preset iteration threshold.

3. The method according to claim 1, characterized in that, The step of correcting the first depth corresponding to the key point in the initial depth map based on the theoretical depth corresponding to the key point includes: The corrected first depth value is obtained by weighted summation of the theoretical depth and the first depth.

4. The method according to claim 3, characterized in that, The method further includes: Obtain the confidence level of the first depth corresponding to the key point in the initial depth map; A first weight is determined based on the confidence level of the first depth, and the first weight is positively correlated with the confidence level. The second weight corresponding to the theoretical depth is determined based on the first weight; And / or, determine the computational weight corresponding to the depth error based on the confidence level of the depth, wherein the computational weight is positively correlated with the confidence level.

5. The method according to claim 4, characterized in that, The change in pose is calculated using the following formula: in, Indicates the change in pose. Indicates the first The image coordinates of each key point in a monocular image. Indicates the first The coordinates of the projection points corresponding to each key point Indicates the first The three-dimensional spatial coordinates of the key points R This indicates that the rotation matrix is ​​used to describe the orientation of the target charging port. t This indicates that the translation vector is used to describe the position of the target charging port. Represents the projection function. Represents the camera intrinsic parameter matrix. Indicates the first The calculation weights corresponding to the depth errors of each key point Indicates the first The depth values ​​corresponding to the key points in the initial depth map. Indicates the first The depth values ​​corresponding to the key points in the corrected depth map.

6. The method according to any one of claims 1-5, characterized in that, The method further includes: Obtain a first image containing the target charging port; The first image is input into a pre-trained multi-task model based on image quality classification and image quality enhancement to obtain the image quality classification result and the second image after image quality enhancement. The image quality classification result is related to the environmental state of the target charging port. Based on the second image, the target location of the target charging port is determined; Based on the target position, the charging plug is controlled to move to a preset position of the target charging port, where the preset position is a position at a set distance from the target position.

7. The method according to claim 6, characterized in that, The multi-task model based on image quality classification and image quality enhancement includes: an encoder, a decoder, and an image quality classification branch, wherein... The encoder includes a deep feature extraction network, a channel-space joint attention module, a first convolutional layer, a stacking operation layer, and a first feature dimension transformation layer connected in sequence. The input end of the deep feature extraction network is used to receive the first image. The decoder includes a second feature dimension transformation layer, a splicing layer, a second convolutional layer, and an upsampling layer connected in sequence. The input of the second feature dimension transformation layer is connected to the output of the deep feature extraction network, and the splicing layer is also connected to the output of the first feature dimension transformation layer. The decoder is used to perform image quality enhancement tasks based on the input image features. The image quality classification branch includes a C3T module, a third convolutional layer, and a fully connected layer connected in sequence. The input of the C3T module is connected to the output of the feature dimension transformation layer. The image quality classification branch is used to perform an image quality classification task based on the input image features.

8. The method according to claim 7, characterized in that, The multi-task model based on image quality classification and image quality enhancement also includes: A first local attention module is provided between the channel-space joint attention module and the first convolutional layer; A second local attention module is provided between the second convolutional layer and the upsampling layer.

9. The method according to claim 6, characterized in that, Before performing stereo matching on the stereo images to obtain the initial depth map corresponding to the stereo images, the method further includes: The stereo images are input into a pre-trained multi-task model based on image quality classification and image quality enhancement to obtain the image quality classification result and the stereo images after image quality enhancement.

10. A charging control device for a charging apparatus, the charging apparatus comprising: A charging plug, characterized in that the device comprises: The first processing module is used to determine the initial pose of the target charging port after controlling the charging plug to move to the preset position of the target charging port, and to acquire a binocular image containing the target charging port. The initial pose is the pose of the target charging port relative to the camera coordinate system corresponding to the binocular image. The second processing module is used to perform stereo matching on the stereo image to obtain the initial depth map corresponding to the stereo image. The third processing module is used to alternately optimize and update the initial depth map and the initial pose based on the three-dimensional spatial coordinates of key points in the preset charging port model, until the preset optimization conditions are met, and then determine the optimized initial pose as the target pose of the target charging port; the preset charging port model and the target charging port are of the same type. The fourth processing module is used to control the charging plug to be aligned with the target charging port and then inserted based on the target pose. The third processing module includes: a first processing subunit, used to map the three-dimensional spatial coordinates of key points in the preset charging port model to the camera coordinate system based on the initial pose, and project them onto the image plane of the camera coordinate system to obtain the projection point coordinates and theoretical depth corresponding to the key points; a second processing subunit, used to correct the first depth corresponding to the key points in the initial depth map based on the theoretical depth corresponding to the key points, to obtain the corrected depth map; a third processing subunit, used to calculate the pose change based on the depth error between the key points in the initial depth map and the corrected depth map, and the reprojection error between the image coordinates corresponding to the key points in the monocular image and the projection point coordinates corresponding to the key points, wherein the monocular image is one image in the binocular image; and a fourth processing subunit, used to obtain the optimized initial pose based on the initial pose and the pose change, and to call the first processing subunit to run.

11. A charging device, characterized in that, The charging device includes: a charging plug and a controller, the controller including: A memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, the processor executing the computer instructions to perform the method of any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to perform the method of any one of claims 1 to 9.

Citation Information

Patent Citations

  • Connector device

    CN108075301A

  • Automatic charging processing method and device of charging pile

    CN114298998A

  • Cross-modal matching positioning method and system based on visual image and point cloud map

    CN116823929A

  • Charging interface attitude detection method and device, electronic equipment and storage medium

    CN118247265A

  • Charging plug and charging pile

    CN216033816U