Charging device and charging control method, apparatus and medium thereof
By optimizing the charging port positioning through binocular stereo matching and multi-task modeling, the problem of large positioning error and high latency of the charging robot in complex environments was solved, and sub-millimeter-level precise docking of the charging plug was achieved.
Patent Information
- Application Number
- CN202511406234.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2045-09-29
AI Technical Summary
Existing charging robot systems have significant technical bottlenecks in millimeter-level positioning accuracy, especially in complex environments where the positioning accuracy of the charging port is insufficient. Traditional methods are easily affected by environmental interference and lack closed-loop optimization mechanisms, resulting in large positioning errors and high system latency.
By employing binocular stereo matching combined with a preset charging port model, the initial depth map and pose are alternately optimized, and a multi-task model is used to improve image quality. Furthermore, confidence and weight adjustment strategies are introduced to achieve sub-millimeter-level positioning of the charging port.
Achieving precise alignment of the charging plug in complex environments improves positioning accuracy and environmental robustness, meets the millimeter-level plugging requirements for automatic charging of electric vehicles, and reduces system latency.
Smart Images

Figure CN120863391B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of charging control, in particular to a charging device and a charging control method, device and medium thereof. BACKGROUND
[0002] With the rapid breakthrough of artificial intelligence, robot vision and force sensing fusion technology, and the continuous rise of new energy vehicle ownership, in recent years, the closed-loop scenario of automated valet parking (AVP) + automatic charging has become a key node of the integration of smart cities and new energy strategies because it can reconstruct user experience and energy efficiency. However, the existing charging robot system has significant technical bottlenecks in terms of millimeter-level positioning accuracy. Traditional single dependence on geometric algorithms such as Perspective-n-Point (PNP) is easily disturbed by image noise, and the pose solution error increases dramatically when the charging port is partially blocked or texture is missing; monocular depth estimation algorithm has a larger depth estimation error when the distance is farther and the target is smaller, making it difficult to meet the millimeter-level insertion requirement on a small target (such as a car charging port) at a long distance. Therefore, how to achieve millimeter-level charging port positioning has become a problem to be solved for automatic charging robots. SUMMARY
[0003] Therefore, the present application provides a charging device and a charging control method, device and medium thereof to solve the problem of low positioning accuracy of the charging robot when positioning the charging port in related technologies.
[0004] In a first aspect, the present application provides a charging control method of a charging device, the charging device comprising a charging plug, the method comprising:
[0005] After controlling the charging plug to move to a preset position of a target charging port, determining an initial pose of the target charging port, and obtaining a binocular image containing the target charging port, the initial pose being the pose of the target charging port relative to the camera coordinate system corresponding to the binocular image;
[0006] Performing binocular stereo matching on the binocular image to obtain an initial depth map corresponding to the binocular image;
[0007] Based on the three-dimensional space coordinates of the key points in the preset charging port model, the initial depth map and the initial pose are updated alternately until a preset optimization condition is met, and the optimized initial pose is determined as the target pose of the target charging port; the preset charging port model is of the same type as the target charging port;
[0008] After controlling the charging plug to align with the target charging port based on the target pose, inserting the charging plug into the target charging port.
[0009] The application solves the problems of insufficient depth map precision and large pose solution error caused by environmental interference in traditional charging positioning by determining the initial pose of the charging port and obtaining the initial depth map through binocular stereo matching, alternately optimizing and updating the depth map and the pose, and finally realizing the accurate alignment of the charging plug and the charging port with the optimized target pose. By determining the initial pose first and obtaining the initial depth map based on the binocular image, and then combining the three-dimensional coordinates of the key points of the preset charging port model for the alternately optimized depth map and pose, the depth deviation and the pose error can be dynamically corrected, the limitations of single modal positioning are avoided, the target pose obtained through final optimization can realize sub-millimeter positioning of the charging port, ensure the accurate alignment of the charging plug, and still maintain high docking success rate in complex environments such as rain, fog and low light, greatly improve the environmental robustness and positioning accuracy of the automatic charging equipment, and meet the millimeter-level plug-in requirements of the electric vehicle automatic charging.
[0010] In an optional embodiment, the preset optimization condition includes that the pose change amount of the initial pose before and after optimization meets a preset pose change amount condition or the number of iterations reaches a preset iteration number threshold.
[0011] The application solves the problem of error accumulation caused by the mutual independence of the depth map and the pose solution in the traditional positioning by first correcting the depth map, then optimizing the pose using the corrected depth map, and then performing a closed-loop operation with a high cycle. And each alternately optimized cycle is based on the three-dimensional coordinates of the key points of the preset charging port model. The initial depth map is corrected to eliminate the depth deviation caused by environmental interference, and then the initial pose is optimized based on the corrected depth map. The optimization is terminated through the preset pose change amount condition and the iteration number threshold, which can ensure the optimization accuracy and avoid system delay caused by excessive iteration. Compared with the traditional single positioning scheme, the positioning error can be gradually compressed, the final pose accuracy is significantly improved, the system real-time performance is considered, and the demand for fast response of the charging equipment is met.
[0012] In an optional embodiment, the alternately optimized and updated initial depth map and initial pose based on the three-dimensional space coordinates of the key points in the preset charging port model includes:
[0013] The three-dimensional space coordinates of the key points in the preset charging port model are mapped to the camera coordinate system based on the initial pose, and projected to the image plane of the camera coordinate system to obtain the projection point coordinates and the theoretical depth corresponding to the key points;
[0014] The first depth corresponding to the key points in the initial depth map is corrected based on the theoretical depth corresponding to the key points, to obtain a corrected depth map;
[0015] Calculate a pose change based on a depth error of the key points between the initial depth map and the corrected depth map and a re-projection error between image coordinates corresponding to the key points in a monocular image and the projection point coordinates corresponding to the key points, the monocular image being one of the binocular images;
[0016] Obtain an optimized initial pose based on the initial pose and the pose change, and return to the step of mapping three-dimensional space coordinates of key points in a preset charging port model to the camera coordinate system based on the initial pose, and projecting to the image plane of the camera coordinate system to obtain the projection point coordinates and the theoretical depth corresponding to the key points.
[0017] The present application corrects the initial depth map by mapping the three-dimensional space coordinates of the key points in the preset charging port model to the camera coordinate system and obtaining the theoretical depth, thereby solving the depth error caused by texture loss and occlusion in binocular matching; then combining the depth error and the re-projection error to calculate the pose change, which can comprehensively consider the deviation of depth and image coordinates, and avoid the pose deviation caused by a single error indicator. This process can accurately position the charging port, reduce the docking failure caused by inaccurate depth and coordinate deviation, and still maintain high positioning accuracy in the local occlusion and weak texture scene of the charging port, thereby improving the reliability of the charging docking.
[0018] In an optional implementation, the correcting the first depth corresponding to the key points in the initial depth map based on the theoretical depth corresponding to the key points comprises:
[0019] Weighted sum of the theoretical depth and the first depth to obtain a corrected first depth value.
[0020] The present application corrects the depth value by weighting the theoretical depth and the first depth in the initial depth map, which can dynamically allocate weights according to the reliability of the two, for example, when the first depth is low in reliability due to environmental interference (such as foggy weather), the weight of the theoretical depth is increased, and vice versa. This flexible weighting strategy can make full use of effective depth information, reduce the influence of single depth data deviation on positioning, make the corrected depth value more consistent with the actual scene, provide an accurate depth basis for subsequent pose optimization, further improve the charging port positioning accuracy, and ensure the accuracy of the charging plug docking.
[0021] In an optional implementation, the method further comprises:
[0022] Obtain the confidence of the first depth corresponding to the key points in the initial depth map;
[0023] Determine a first weight corresponding to the first depth based on the confidence of the first depth, the first weight being positively correlated with the confidence;
[0024] The second weight corresponding to the theoretical depth is determined based on the first weight;
[0025] And / or, determine the computational weight corresponding to the depth error based on the confidence level of the depth, wherein the computational weight is positively correlated with the confidence level.
[0026] This invention addresses the problems of traditional weight allocation, such as lack of reliable data and strong subjective assumptions, by introducing depth confidence scores corresponding to key points in the initial depth map. By obtaining the confidence score of the first depth, the first weight is positively correlated with the confidence score, ensuring that the first depth with high confidence plays a greater role in depth correction. Simultaneously, the second weight of the theoretical depth and the calculation weight of the depth error are reasonably determined. For example, in the reflective metal area of the charging port, the first depth confidence is low, corresponding to a small first weight and a small depth error calculation weight, reducing the interference of unreliable depth data in this area on optimization. This dynamic weight adjustment strategy based on confidence scores allows the optimization process to better reflect actual data quality, improves the accuracy of depth correction and pose optimization, and enhances the system's adaptability to complex environments.
[0027] In one alternative implementation, the pose change is calculated using the following formula:
[0028]
[0029] in, Indicates the change in pose. Indicates the first The image coordinates of each key point in a monocular image. Indicates the first The coordinates of the projection points corresponding to each key point Indicates the first The three-dimensional spatial coordinates of the key points R This indicates that the rotation matrix is used to describe the orientation of the target charging port. t This indicates that the translation vector is used to describe the position of the target charging port. Represents the projection function. Represents the camera intrinsic parameter matrix. Indicates the first The calculation weights corresponding to the depth errors of each key point Indicates the first The depth values corresponding to the key points in the initial depth map. Indicates the first The depth values corresponding to the key points in the corrected depth map.
[0030] The application solves the problems of single error consideration and one-sided optimization direction in traditional pose optimization by incorporating the reprojection error and the depth error into a unified pose change optimization target. In the formula, the reprojection error considers the deviation between the image coordinates of the key points and the projection point coordinates, and the depth error considers the deviation between the depth before and after correction by combining with the dynamically calculated weight. The two collaborative optimizations can comprehensively correct the pose deviation. At the same time, the introduction of the calculation weight can adjust the influence degree of the depth error in the optimization according to the depth confidence, avoiding the interference of low confidence depth data on the optimization result. The formula makes the pose optimization have a clear mathematical basis, can accurately calculate the pose adjustment amount, ensures that the optimized pose is more consistent with the actual situation, significantly improves the charging port positioning accuracy, and reduces the failure rate of charging docking.
[0031] In an optional embodiment, the method further comprises:
[0032] obtaining a first image containing the target charging port;
[0033] inputting the first image into a pre-trained multi-task model based on picture quality classification and picture quality improvement to obtain a picture quality classification result and a second image with improved picture quality, the picture quality classification result being related to an environmental state in which the target charging port is located;
[0034] determining a target position of the target charging port based on the second image;
[0035] controlling a charging plug to move to a preset position of the target charging port based on the target position, the preset position being a position with a set distance from the target position.
[0036] The application solves the problems of split of environmental perception and image enhancement and positioning failure caused by poor image quality in complex environments in traditional positioning by introducing a multi-task model to process the first image in the early stage of charging positioning. The multi-task model can output the picture quality classification result (reflecting the environmental state) and the second image with improved quality at the same time, which not only clearly reflects the current environment (such as rain and fog), but also eliminates the influence of environmental interference on the image, providing a high-quality image basis for subsequent determination of the target position of the charging port. By determining the target position based on the second image and controlling the charging plug to move to the preset position, the misjudgment of the target position caused by the blurring and low contrast of the original image can be avoided, and the accuracy of target positioning in complex environments can be greatly improved, laying a foundation for subsequent accurate docking.
[0037] In an optional implementation, the multi-task model based on the picture quality classification and the picture quality improvement comprises an encoder, a decoder, and a picture quality classification branch.
[0038] The decoder comprises a second feature dimension transformation layer, a splicing layer, a second convolutional layer, and an up-sampling layer connected in sequence, an input end of the second feature dimension transformation layer is connected with an output end of the deep feature extraction network, the splicing layer is further connected with an output end of the first feature dimension transformation layer, and the decoder is used for performing a picture quality improvement task based on an input image feature.
[0039] The picture quality classification branch comprises a C3T module, a third convolutional layer, and a full connection layer connected in sequence, an input end of the C3T module is connected with an output end of the feature dimension transformation layer, and the picture quality classification branch is used for performing a picture quality classification task based on an input image feature.
[0040] The present application solves the problems of single function of a traditional single-task model and high system delay caused by multi-model series connection through the collaborative design of the encoder, the decoder, and the picture quality classification branch. The encoder extracts deep image features through the deep feature extraction network, the channel-space joint attention module, and the like in sequence, thereby providing high-quality feature support for subsequent tasks. The decoder performs a picture quality improvement task based on the features, thereby outputting a clear image. The picture quality classification branch realizes environmental state classification through the C3T module and the like. The three share the encoder features, thereby avoiding repeated calculation and greatly reducing the inference delay. Compared with the traditional independent image enhancement model and classification model, the structure can simultaneously complete two tasks, thereby improving the processing efficiency, and the collaborative optimization of the modules makes the image quality improvement and the classification result more accurate, thereby providing efficient and reliable image preprocessing support for the charging positioning.
[0041] In an optional implementation, the multi-task model based on the picture quality classification and the picture quality improvement further comprises:
[0042] A first local attention module is arranged between the channel-space joint attention module and the first convolutional layer.
[0043] A second local attention module is arranged between the second convolutional layer and the up-sampling layer.
[0044] The application solves the problem of low task precision caused by insufficient attention to key area features of the charging port and complex background interference of the traditional model by adding first and second local attention modules in the multi-task model. The first local attention module is located between the channel-space joint attention module and the first convolution layer, which can enhance the feature expression of the charging port key area (such as the terminal) in the features extracted by the encoder and suppress background noise. The second local attention module is located between the second convolution layer and the up-sampling layer, which can focus on optimizing the details of the charging port area during the image quality improvement process of the decoder, so that the charging port features in the output image are clearer. Through the local attention mechanism, the model can focus on the key area, improve the accuracy of picture quality classification and the pertinence of image quality improvement, provide more accurate image information for subsequent positioning, and further improve the charging positioning accuracy.
[0045] In an optional embodiment, before performing binocular stereo matching on the binocular image to obtain an initial depth map corresponding to the binocular image, the method further comprises:
[0046] The binocular image is input into a pre-trained multi-task model based on picture quality classification and picture quality improvement respectively to obtain a picture quality classification result and a binocular image after picture quality improvement.
[0047] The application solves the problem of large matching error and low initial depth map precision caused by poor original image quality (such as rain and fog blur) in the traditional binocular matching by performing multi-task model processing on the binocular image before binocular stereo matching. By improving the quality of the binocular image through the multi-task model, the influence of environmental interference on the left and right images is eliminated, the consistency and clarity between the binocular images are ensured, and high-quality input is provided for subsequent stereo matching. At the same time, the picture quality classification result can assist in judging the current environment and provide a basis for subsequent optimization strategy adjustment. Compared with directly matching the original binocular image, this processing can greatly reduce the matching noise, improve the precision of the initial depth map, provide a more reliable depth basis for subsequent alternating optimization and update, and further improve the accuracy and stability of the charging port positioning.
[0048] In a second aspect, the application provides a charging control device of a charging device, the charging device comprising a charging plug, and the device comprising:
[0049] A first processing module is configured to determine an initial pose of the target charging port relative to a camera coordinate system corresponding to the binocular image and acquire a binocular image containing the target charging port after controlling the charging plug to move to a preset position of the target charging port, wherein the initial pose is the pose of the target charging port relative to the camera coordinate system corresponding to the binocular image.
[0050] A second processing module is configured to perform binocular stereo matching on the binocular image to obtain an initial depth map corresponding to the binocular image.
[0051] The third processing module is configured to update the initial depth map and the initial pose alternately based on the three-dimensional space coordinates of the key points in the preset charging port model until a preset optimization condition is met, and determine the optimized initial pose as a target pose of a target charging port.
[0052] The fourth processing module is configured to control the charging plug to be inserted into the target charging port after the charging plug is aligned with the target charging port based on the target pose.
[0053] The third processing module comprises: a first processing subunit configured to map the three-dimensional space coordinates of the key points in the preset charging port model to a camera coordinate system based on the initial pose, and project the three-dimensional space coordinates of the key points to an image plane of the camera coordinate system to obtain projection point coordinates and theoretical depths corresponding to the key points; a second processing subunit configured to correct first depths corresponding to the key points in the initial depth map based on the theoretical depths corresponding to the key points to obtain a corrected depth map; a third processing subunit configured to calculate a pose change amount based on a depth error between the initial depth map and the corrected depth map, and a re-projection error between image coordinates corresponding to the key points in a monocular image and the projection point coordinates corresponding to the key points, the monocular image being one of the binocular images; and a fourth processing subunit configured to obtain an optimized initial pose based on the initial pose and the pose change amount, and call the first processing subunit to run.
[0054] In a third aspect, the present application provides a charging device, the charging device comprising: a charging plug and a controller, the controller comprising:
[0055] The memory and the processor are connected to each other in communication, the memory stores computer instructions, and the processor executes the computer instructions to perform the method in the first aspect and any optional implementation thereof.
[0056] In a fourth aspect, the present application provides a computer readable storage medium, the computer readable storage medium storing computer instructions, and the computer instructions are used to make a computer execute the method provided in the first aspect or any implementation thereof.
[0057] The present application has the following beneficial effects:
[0058] This invention determines the initial pose of the charging port and uses an initial depth map obtained through binocular stereo matching. It then alternately optimizes and updates the depth map and pose, finally using the optimized target pose to achieve precise alignment between the charging plug and the charging port. This effectively solves the problems of insufficient depth map accuracy and large pose calculation errors caused by environmental interference in traditional charging positioning. By first determining the initial pose and obtaining the initial depth map based on binocular images, and then combining the 3D coordinates of key points of the preset charging port model for alternating optimization of the depth map and pose, it can dynamically correct depth deviation and pose error, avoiding the limitations of single-modal positioning. This allows the final optimized target pose to achieve sub-millimeter-level positioning of the charging port, ensuring precise alignment and insertion of the charging plug. It maintains a high docking success rate even in complex environments such as rain, fog, and low light, significantly improving the environmental robustness and positioning accuracy of automatic charging equipment and meeting the millimeter-level insertion requirements of electric vehicle automatic charging. Attached Figure Description
[0059] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0060] Figure 1 This is a flowchart of a charging control method for a charging device according to an embodiment of the present invention;
[0061] Figure 2 This is a flowchart of a charging control method for another charging device according to an embodiment of the present invention;
[0062] Figure 3 This is a schematic diagram of the network structure of a multi-task model based on image quality classification and image quality improvement according to an embodiment of the present invention.
[0063] Figure 4 This is a schematic diagram of the structure of a local attention module according to an embodiment of the present invention;
[0064] Figure 5 This is a schematic diagram of the C3T module according to an embodiment of the present invention;
[0065] Figure 6 These are comparison images of image denoising effects according to embodiments of the present invention;
[0066] Figure 7 This is a diagram illustrating the binocular depth estimation effect according to an embodiment of the present invention;
[0067] Figure 8 This is a schematic diagram of the main process of automatic charging of the charging device according to an embodiment of the present invention;
[0068] Figure 9 is a structural schematic diagram of a charging control device of a charging device according to an embodiment of the application;
[0069] Figure 10 is a structural schematic diagram of a controller of a charging device according to an embodiment of the application. DETAILED DESCRIPTION
[0070] To make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0071] The existing charging robot system has significant technical bottlenecks in adaptability to complex environments and millimeter-level positioning accuracy:
[0072] Lack of environmental robustness:
[0073] Under bad weather such as rain, fog and snow or low light conditions, the image collected by the traditional vision system is severely degraded (blurred, low contrast), which leads to a cliff-like performance drop of the recognition model based on deep learning (such as semantic segmentation and target detection);
[0074] Insufficient positioning accuracy:
[0075] Single dependence on geometric algorithms (such as PNP) is easily disturbed by image noise, and the pose solution error increases dramatically when the charging port is partially blocked or texture is missing;
[0076] Monocular depth estimation algorithms are difficult to meet the millimeter-level plugging requirement on small targets at a long distance (such as a car charging port);
[0077] System-level coordination defects:
[0078] The existing solutions lack a closed-loop optimization mechanism from environmental perception to image enhancement to positioning decision, and cannot dynamically adjust the processing strategy according to the real-time environmental state.
[0079] Although the current mainstream solutions attempt to introduce image enhancement or fuse multiple sensors to improve the positioning robustness, there are still two major limitations:
[0080] Enhancement and recognition are separated: traditional image enhancement methods (such as histogram equalization and dehazing algorithms based on physical models) and subsequent recognition and positioning modules are designed independently, and the enhanced image may not be suitable for task requirements;
[0081] Multi-source data is not deeply coupled: visual and depth information are often simply concatenated, without establishing a joint optimization mechanism between geometric priors and deep learning estimation, resulting in insufficient positioning accuracy in complex scenes.
[0082] How to break through the environmental interference and precision bottleneck and realize "all-weather, all-scene, millimeter-level" charging port positioning has become a key technical challenge restricting automatic charging robots.
[0083] In actual electric vehicle automatic charging scenes, millimeter-level precision of charging plugging depends on high-precision 6D pose estimation (3D position + 3D rotation) of the charging port. However, complex environmental degradation and cross-scale positioning error accumulation lead to double challenges for existing solutions:
[0084] 1. Perception failure under environmental disturbance: scenes such as rain, fog and snow cause image blurring and contrast reduction, and traditional visual algorithms cannot simultaneously realize environment state discrimination and task-oriented image enhancement, resulting in subsequent segmentation / detection failure;
[0085] 2. Single modality positioning bottleneck:
[0086] Pure geometric methods (such as PNP) rely on charging port texture features, and in occluded or weak texture scenes, the pose solution error is more than 10 cm;
[0087] Monocular depth estimation has a ranging error of more than 5% for small targets (charging ports), and binocular depth algorithms have improved accuracy but lack geometric constraints and are easily disturbed by matching noise;
[0088] 3. System delay due to module fragmentation: environmental perception, image enhancement, and positioning decision are independently run, without establishing a closed-loop optimization mechanism, resulting in a processing delay of more than 200 ms.
[0089] In view of the defects of the prior art, the embodiment of the present application fuses a multi-task model of cross-domain self-adaptation and image quality evaluation, aiming to build high-robustness visual positioning and cross-scene self-adaptation ability, and to overcome the "last millimeter" precise positioning problem in complex environments, providing key infrastructure support for the automatic driving ecosystem.
[0090] According to the embodiment of the present application, a charging control method for a charging device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0091] A charging control method of a charging device is provided in the embodiment. The charging device is an automatic charging device such as an automatic charging robot, etc. The charging device includes a charging plug for being inserted into a charging port on a vehicle to charge the vehicle. The method is applied to a controller of the charging device, such as a single-chip microcomputer, a MCU, etc. Figure 1 A flowchart of the charging control method of the charging device according to the embodiment of the present application is shown in FIG. 2, which includes the following steps: Figure 1
[0092] In step S101, after the charging plug is controlled to move to a preset position close to the target charging port, the initial pose of the target charging port is determined, and a binocular image containing the target charging port is obtained.
[0093] The preset position is a close position of about 200 mm away from the target charging port, which is calibrated in advance, and is convenient for subsequent accurate positioning. The initial pose is the pose of the target charging port relative to the camera coordinate system corresponding to the binocular image. Specifically, the initial pose is the 6D pose (including 3D position information and 3D rotation information) of the target charging port relative to the camera coordinate system corresponding to the binocular image, which is used to describe the position and orientation of the charging port in the three-dimensional space and is the basis for subsequent positioning optimization. The camera coordinate system is a three-dimensional coordinate system with the optical center of the left camera of the binocular camera as the origin, the X-axis along the horizontal right of the camera, the Y-axis vertically downward, and the Z-axis along the optical axis of the camera forward. It is the reference coordinate system for pose solving. In addition, in actual application, the left camera can be replaced by the right camera, and the present application is not limited thereto. The binocular image is two images containing the target charging port, which are simultaneously captured by the binocular camera (including left and right two synchronously captured RGB cameras) carried on the charging device.
[0094] Specifically, the determination method of the initial pose of the target charging port can be realized by the charging port pose determination algorithm in the prior art, which will not be described one by one. For example, the initial pose can be obtained by the following steps: a binocular image containing the target charging port is obtained by the binocular image acquisition carried on the vehicle, then the charging port is segmented by using one of the monocular images in the binocular, the charging port position is detected after the segmentation, and the interference of obstacles (such as fallen leaves, dust, etc.) is excluded; the initial pose of the target charging port relative to the camera coordinate system is calculated by using the PNP algorithm combined with the relative geometric constraints of the 7 fixed terminals in the preset charging port model (the charging port terminals are standardized structures, and the three-dimensional relative positions of the 7 terminals are known), which is denoted as R0 and t0, wherein R0 is a 3x3 rotation matrix describing the orientation of the charging port, and t0 is a 3x1 translation vector describing the position of the charging port.
[0095] In step S102, the binocular image is subjected to binocular stereo matching to obtain an initial depth map corresponding to the binocular image.
[0096] Specifically, binocular stereo matching is to calculate the disparity between pixels (the distance between the left pixel and the corresponding pixel in the right image in the horizontal direction) by finding the matching relationship of corresponding pixel points in binocular images to obtain depth information. The existing binocular stereo matching algorithm can be used to convert the binocular image into a disparity map to obtain the depth map corresponding to the binocular image according to the disparity map. Exemplarily, in the embodiment of the present application, the Unimatch depth estimation algorithm is taken as an example for description, which is a binocular depth estimation algorithm based on Transformer. Through multi-scale feature enhancement and global matching strategy, the matching accuracy of low texture and occluded areas is improved.
[0097] Exemplarily, the specific implementation process of the above step S102 includes:
[0098] The binocular image is preprocessed: firstly, based on the camera distortion coefficient (previously obtained through calibration), the image is de-distorted to eliminate the image distortion caused by lens optical distortion; secondly, the image is processed by adaptive histogram equalization (CLAHE) to reduce the influence of uneven illumination in rain, fog and low light environment and improve the image contrast.
[0099] Subsequently, the Unimatch depth estimation algorithm is used to perform binocular stereo matching: first, the multi-scale features of the binocular image (left image feature FL and right image feature FR) are extracted by the Transformer encoder, and the position encoding (generated by the trigonometric function, which retains the pixel space position information) is introduced to enhance the feature expression of low texture areas (such as the pure black surface of the charging port); secondly, the feature similarity matrix S (dimension HxWxHxW, H and W are the height and width of the image respectively, S (i,j,k,l) represents the cosine similarity of the left image (i,j) pixel and the right image (k,l) pixel) is calculated, and the matching probability distribution is generated within the local search window (horizontal direction ±90 pixels, based on the size of the charging port preset) by the Softmax function (temperature coefficient τ=0.01, controlling the sharpness of the probability distribution), and the maximum probability point is selected as the corresponding matching point; thirdly, the disparity map D (D (i,j)=|i-k|, k is the horizontal coordinate of the matching point in the right image) is calculated according to the matching point, and then the initial depth map is converted from the disparity map through the geometric formula Z=(Bxf) / D (where B is the baseline length of the binocular camera, which is previously calibrated as 120mm; f is the focal length of the camera, which is previously calibrated as 800 pixels).
[0100] Step S103, based on the three-dimensional space coordinates of the key points in the preset charging port model, the initial depth map and the initial pose are updated alternately, until the preset optimization condition is met, and the optimized initial pose is determined as the target pose of the target charging port.
[0101] wherein the preset charging port model is of the same type as the target charging port. It refers to a three-dimensional model that is completely consistent with the type of the target charging port, and contains three-dimensional coordinates of all structures of the charging port, wherein the three-dimensional space coordinates of the 7 terminals are known key parameters, which are the reference for optimization. By cyclically performing the two steps of "depth optimization" and "pose optimization", the depth deviation and the pose error are gradually corrected, and the collaborative optimization of the two is realized.
[0102] Exemplarily, the above-mentioned preset optimization condition includes two termination criteria, one is the change threshold of the pose before and after optimization (such as the change of the rotation angle <0.1°, the change of the translation distance <0.05mm), and the other is the maximum number of iterations (preset as 5 times), and the optimization is stopped when any criterion is met.
[0103] In step S104, the charging plug is inserted after being aligned with the target charging port based on the target pose.
[0104] Specifically, by the target pose (R, t) obtained according to step S103, the pose of the charging plug is adjusted by the robot motion control system, so that it is accurately aligned with the target charging port, that is, the center axis of the charging plug coincides with the center axis of the charging port (position alignment), and the orientation of the charging plug is consistent with the orientation of the charging port (attitude alignment), and finally the insertion action is completed. The robot motion control system is a system for controlling the motion of a charging device such as a charging robot arm, which can calculate the motion parameters (such as angle, speed) of each joint of the robot arm according to the target pose, so as to realize the accurate displacement and attitude adjustment of the charging plug.
[0105] The embodiment of the application determines the initial pose of the charging port, and updates the depth map and the pose alternately by the initial depth map obtained by binocular stereo matching, and finally realizes the accurate alignment of the charging plug and the charging port by the optimized target pose, effectively solving the problem of insufficient depth map accuracy and large pose calculation error caused by environmental interference in traditional charging positioning. By determining the initial pose first and obtaining the initial depth map based on the binocular image, and then combining the key point three-dimensional coordinates of the preset charging port model to alternately optimize the depth map and the pose, the depth deviation and the pose error can be dynamically corrected, the limitations of single modal positioning are avoided, the target pose obtained by the final optimization can realize sub-millimeter level positioning of the charging port, and the charging plug can be accurately aligned and inserted. In complex environments such as rain, fog and low light, it can still maintain a high docking success rate, greatly improves the environmental robustness and positioning accuracy of the automatic charging equipment, and meets the millimeter-level plug-in requirements of the automatic charging of electric vehicles.
[0106] The embodiment also provides a charging control method of a charging device, such as an automatic charging robot, the charging device comprising a charging plug for being inserted into a charging port on a vehicle to charge the vehicle. Figure 2 The charging control method of the charging device according to the embodiment of the application is shown in a flowchart as shown in Figure 2 The flowchart comprises the following steps.
[0107] In step S201, a first image containing a target charging port is acquired.
[0108] Specifically, one monocular image in the binocular image containing the target charging port acquired by the image acquisition device, such as the binocular camera, carried by the charging device (hereinafter referred to as charging robot) can be directly affected by environmental light, weather conditions and other factors.
[0109] In step S202, the first image is input into a pre-trained multi-task model based on picture quality classification and picture quality improvement to obtain a picture quality classification result and a second image after picture quality improvement.
[0110] The picture quality classification result is related to the environmental state in which the target charging port is located. For example, the environmental state includes sunny, rainy, foggy, snowy, low light, and the corresponding picture quality classification result can also be sunny, rainy, foggy, snowy, low light. This is only an example, and the application is not limited thereto.
[0111] Specifically, the pre-trained multi-task model of picture quality classification and picture quality improvement is used to simultaneously perform environmental state classification and image quality enhancement on the first image, solve the problem of fragmented function and low processing efficiency of the traditional single model, and output the picture quality classification result (reflecting the environmental state) and the second image after quality improvement. The second image after quality improvement by the multi-task model is significantly optimized in terms of clarity, contrast and noise suppression compared with the first image, and can be directly used for subsequent charging port target position determination.
[0112] For example, the model training data set of the model contains charging port images in five types of environments, such as sunny, rainy, foggy, snowy and low light, to ensure the generalization ability of the model. The specific training process of the model can be realized by referring to the existing model training process, which will not be described here. Further, the picture quality classification result, i.e. the environmental state label output by the above-mentioned model, can contain eight categories, such as “sunny”, “light rain”, “heavy rain”, “thin fog”, “thick fog”, “light snow”, “heavy snow” and “low light (<10 lux)”, which are used to reflect the environmental state in which the target charging port is located, and provide a basis for subsequent positioning parameter adjustment.
[0113] Further, as shown in Figure 3 The multi-task model based on image quality classification and image quality improvement includes an encoder 301, a decoder 302 and an image quality classification branch 303. The encoder 301 includes a deep feature extraction network, a channel-spatial joint attention module, a first convolutional layer, a stacking operation layer and a first feature dimension transformation layer connected in sequence. The input end of the deep feature extraction network is configured to receive a first image.
[0114] The decoder 302 includes a second feature dimension transformation layer, a splicing layer, a second convolutional layer and an up-sampling layer connected in sequence. The input end of the second feature dimension transformation layer is connected with the output end of the deep feature extraction network. The splicing layer is further connected with the output end of the first feature dimension transformation layer. The decoder is configured to perform an image quality improvement task based on the input image features.
[0115] The image quality classification branch 303 includes a C3T module, a third convolutional layer and a full connection layer connected in sequence. The input end of the C3T module is connected with the output end of the feature dimension transformation layer. The image quality classification branch is configured to perform an image quality classification task based on the input image features.
[0116] The embodiment of the present application solves the problem of single function of the traditional single-task model and high system delay caused by multi-model series connection through the collaborative design of the encoder, the decoder and the image quality classification branch. The encoder extracts deep image features through the deep feature extraction network, the channel-spatial joint attention module and the like in sequence, thereby providing high-quality feature support for subsequent tasks. The decoder performs an image quality improvement task based on the features, thereby outputting a clear image. The image quality classification branch realizes environment state classification through the C3T module and the like. The three share the encoder features, thereby avoiding repeated calculation and greatly reducing the inference delay. Compared with the traditional independent image enhancement model and classification model, the structure can complete two tasks at the same time, thereby improving the processing efficiency. In addition, the modules are optimized in collaboration, thereby making the image quality improvement and classification results more accurate, thereby providing efficient and reliable image preprocessing support for charging positioning.
[0117] In some optional embodiments, the multi-task model based on image quality classification and image quality improvement further includes:
[0118] A first local attention module is arranged between the channel-spatial joint attention module and the first convolutional layer.
[0119] A second local attention module is arranged between the second convolutional layer and the up-sampling layer.
[0120] The embodiment of the application solves the problem of low task precision caused by insufficient attention to key area features of the charging port and complex background interference of the traditional model by adding first and second local attention modules in the multi-task model. The first local attention module is located between the channel-space joint attention module and the first convolution layer, which can enhance the feature expression of the charging port key area (such as the terminal) in the features extracted by the encoder and suppress background noise. The second local attention module is located between the second convolution layer and the up-sampling layer, which can focus on optimizing the details of the charging port area during the image quality improvement process of the decoder, so that the charging port features in the output image are clearer. Through the local attention mechanism, the model can focus on the key area, improve the accuracy of picture quality classification and the pertinence of image quality improvement, provide more accurate image information for subsequent positioning, and further improve the charging positioning accuracy.
[0121] Specifically, the existing image quality processing model of the intelligent charging robot has three core bottlenecks: first, the traditional single-task model (image enhancement or environment classification alone) has fragmented functions, and multiple models need to be connected in series, resulting in a processing delay of more than 200 ms, which cannot meet the real-time requirements; second, the feature extraction process lacks targeted enhancement of the key area of the charging port (such as the terminal edge), and background noise easily interferes with effective features, resulting in weak feature discrimination ability in rain, fog, and low-light environments; third, the encoder and the decoder lack sufficient feature interaction, and the high-level semantic features and the low-level detail features are not fully fused, resulting in problems such as edge blur and texture distortion in the enhanced image. In addition, the existing model does not introduce a local attention mechanism, and the attention degree of the features of the local target such as the charging port is insufficient, further restricting the image quality improvement effect and the environment classification accuracy, and it is difficult to adapt to the harsh requirements of the automatic charging scene.
[0122] To solve the above problems, the multi-task model provided in the embodiment of the application adopts an end-to-end architecture of "encoder-decoder-classification branch", the encoder is responsible for extracting multi-scale features of the first image, the decoder performs a picture quality improvement task based on the encoder features to output a second image, and the picture quality classification branch performs an environment state classification task based on the encoder features; at the same time, first and second local attention modules (i.e., CASS local attention modules) are arranged at key nodes of the encoder and the decoder to strengthen the key features of the charging port. The model input is a first image of 1280x720x3 (RGB), and the output is a classification result of 8 types of environment states (such as sunny, heavy rain, dense fog, etc.) and a second image of 1280x720x3, and the model inference speed is ≥30fps, which meets the real-time processing requirements. Specifically, the structure of the local attention module is as shown in Figure 4 .
[0123] Specifically, in Figure 3The depth feature extraction network of the encoder 301 mentioned above is a convolutional neural network based on ResNet-50. By introducing dilated convolution (dilation rate=2) to replace part of the ordinary convolution, the receptive field is expanded without reducing spatial resolution. It can extract multi-scale features from low-level (edges, texture) to high-level (semantics) features of the image, and uses dilated convolution to extract the depth information features of the image. Dilated convolution can effectively capture multi-scale information in the image while maintaining spatial resolution and avoiding information loss caused by pooling operations. The channel-space joint attention module, abbreviated as CA_ODS module, is used to enhance the discriminative ability of features. By jointly optimizing features through channel attention and spatial attention, important regions and features are highlighted, noise and redundant information are suppressed, and the quality and robustness of features are further improved. The feature channel weights are calculated through channel attention (SENet mechanism) to highlight effective channel features; the spatial attention map is generated through spatial attention (CBAM mechanism) to suppress background noise and achieve accurate feature selection. The first local attention module (CASS) includes an LS Block (Local Spatial Feature Extraction Block) and a TSM (Temporal Shifting Module), which can focus on the local features of the charging port and enhance the expression of key features such as terminal edges and contours. The first convolutional layer uses two 3×3 convolutions with a stride of 1 and padding of 1, further extracting deep semantic information from the output features of the first local attention module. The activation function used is SiLU. It should be noted that... Figure 3 The first local attention module is not shown. Stacked operation layer ( Figure 3 (Abbreviated as stacking): The multi-channel features output from the first convolutional layer are stacked according to channel dimension to increase the number of feature channels and retain richer feature information. For example, 256-channel features are stacked into 512 channels. First feature dimension transformation layer: Uses 1×1 convolution, with one layer, stride = 1, and no padding. It is used to adjust the channel dimension of the stacked features to adapt to the input requirements of subsequent decoder concatenation layers and classification branches. For example, 512-channel features are reduced to 256 channels to further extract deeper semantic features. Dimensionality reduction or expansion of features reduces computational complexity while retaining key information.
[0124] Furthermore, the second feature dimension transformation layer of the decoder 302 above employs a 1×1 convolution, consistent with the structure of the first feature dimension transformation layer. It is used to adjust the low-level feature channel dimension of the deep feature extraction network output, matching it with the number of high-level feature channels in the encoder output, facilitating subsequent splicing and fusion. (Splicing layer) Figure 3In the middle, it is called splicing: the high-level features output by the encoder are spliced with the low-level features adjusted by the decoder in the channel dimension. The low-level features contain rich detail information, while the high-level features contain semantic information. The combination of the two can achieve a balance between semantics and details, realize the fusion of semantic features and detail features, and improve the texture and edge quality of the enhanced image. The second convolutional layer: uses 3x3 convolution, 3 layers, step = 1, padding = 1, and the activation function is SiLU function, which is used to refine the fused features after splicing and eliminate redundant information generated in the feature fusion process. The second local attention module: has the same structure as the first local attention module and is arranged between the second convolutional layer and the upsampling layer, which is used to enhance the detail features (such as terminal texture and edge profile) of the charging port in the fused features, and avoid the loss of details in the upsampling process. It should be noted that the second local attention module is not shown in the Figure 3 The upsampling layer ( Figure 3 In the middle, it is called upsample four times): uses transpose convolution to realize the enlargement of the feature map size, and the transpose convolution kernel size = 4x4, step = 2, padding = 1. Through multi-stage upsampling, the feature map is restored to the original image resolution such as (1280x720).
[0125] Further, the C3T module of the above picture quality classification branch 303: combines the C3 module (residual bottleneck structure) and the Transformer attention mechanism module, retains local features (such as fog droplet texture and raindrop trajectory) through residual connection, captures global features (such as the distribution of the whole image fog concentration) through the multi-head self-attention of Transformer, and enhances the discriminability of environmental state features. Exemplarily, the specific network structure of the C3T module is as shown in Figure 5 , wherein the activation function of the activation block is SiLU function.
[0126] The third convolutional layer: uses 1x1 convolution kernel convolution, 1 layer, step = 1, no padding, which is used to compress the channel dimension of the C3T module output feature, and reduce the calculation amount of the subsequent fully connected layer. The fully connected layer: contains 2 layers of fully connected network, the first layer is the hidden layer (dimension = 128), and the activation function is SiLU function; the second layer is the output layer (dimension = 8), which corresponds to 8 kinds of environmental states. Through the Softmax activation function, the probability value of each category is output, and the category with the maximum probability is the picture quality classification result, such as sunny: 0.01, heavy rain: 0.02, heavy fog: 0.95, low light: 0.02); select the "heavy fog" with the maximum probability as the picture quality classification result, and complete the environmental state classification task.
[0127] In addition, in practical applications, the deep feature extraction network described above can be replaced by: ordinary convolution + multi-scale pooling (such as Pyramid Pooling Module, PPM). The main purpose of the dilated convolution is to capture multi-scale information, and the pyramid pooling module can also achieve similar functions, and may be more flexible in some scenarios. The first feature dimension transformation layer and the second feature dimension transformation layer described above can be replaced by: other dimension reduction or dimension increasing operations, such as GroupConvolution or Depthwise Separable Convolution.
[0128] Further, the first and second local attention modules described above both adopt the CASS Block structure, which includes the LS Block and the TSM module. Among them, the LS Block structure adopts the 3x3 depth separable convolution (the number of convolution kernels = the number of input channels, groups = the number of input channels), the step = 1, the padding = 1, and cooperates with the SiLU activation function to extract the local spatial features of the input features, focus on the charging port area (the receptive field size = 31x31, which is suitable for the size of the charging port in the feature map), and suppress the background noise.
[0129] The TSM module: divides the input feature map into 3 parts along the channel dimension (channel ratio 1:1:1), among which two parts are shifted by 1 pixel along the horizontal and vertical directions respectively, and the third part remains unchanged. Then, the three feature maps are spliced along the channel dimension, fused by 1x1 convolution, and the spatial correlation of the features is enhanced to avoid local feature isolation. Among them, the LS Block effectively extracts the local spatial information of the input feature map to compensate for the lack of local modeling capability. The CASS Block is the core module of the model, which undergoes a series of processing in the input stage, so that the network can learn deeper and richer feature representations, and at the same time, through batch normalization, the training and inference process is kept efficient and stable. The input features first undergo conventional convolution and batch normalization, then enter the LS module and the CASS module, and then pass through the fusion feature and TSM channel transfer module to enhance the spatial features and temporal features.
[0130] The embodiment of the application synchronously realizes picture quality improvement and environment classification through the integrated architecture of the encoder-decoder-classification branch, reduces the processing delay, speeds up the inference speed, improves the environment classification accuracy, solves the problem of traditional single-task model fragmentation, meets the dual requirements of real-time and accuracy, and sets the first and second local attention modules (CASS modules) to focus on the charging port area, extracts local spatial features and enhances spatial correlation, so that the response value of key features such as the charging port terminal edge and texture is increased by 30%-40%, and the response value of background noise is reduced by 20%-30%; and has good effect in rain, fog, low light and other scenes.
[0131] Exemplarily, taking the multi-task model as shown in Figure 3 the whole model running process is described as follows:
[0132] 1. First, the 1280x720 size image of each batch enters the model;
[0133] 2. The deep information features feature1 of the image are extracted through the hole convolution network;
[0134] 3. The feature feature2 is obtained by extracting the feature through the channel-space joint attention module;
[0135] 4. The feature feature3 is obtained by 3x3 convolution, stacking and 1x1 convolution of the feature feature2;
[0136] 5. The feature1 for image quality classification is obtained by the C3T module, convolution and full connection layer of the feature feature3;
[0137] 6. The feature feature4 is obtained by 1x1 convolution operation of the feature feature1;
[0138] 7. The feature feature5 is obtained by splicing the feature feature4 and the feature feature3 up-sampled by 4 times;
[0139] 8. The generated high-resolution image is obtained by 3x3 convolution and 4 times up-sampling of the feature feature5.
[0140] Step S203, determining the target position of the target charging port based on the second image.
[0141] The target position refers to the pixel coordinates (u, v) of the charging port center in the image coordinate system and the actual distance (d) from the charging port center to the camera, which can completely describe the position of the charging port in the three-dimensional space, and provide a basis for the subsequent movement of the charging plug.
[0142] Specifically, the target position of the charging port can be determined by a fusion algorithm of semantic segmentation + target detection, solving the problem of large positioning error of traditional single detection algorithm in complex environment, and ensuring the accuracy of the target position.
[0143] Illustratively, the charging port area in the image can be accurately segmented by a Deeplabv3+ semantic segmentation model, outputting a charging port area mask (only keeping the charging port pixels, and setting the background pixels to 0), calculating the pixel center coordinates of the charging port area through the mask, then using a YOLOv8 target detection model to obtain the bounding box coordinates of the charging port, calculating the center pixel coordinates of the bounding box, and finally weighting and fusing the pixel center coordinates of the charging port area obtained by semantic segmentation and the center pixel coordinates of the bounding box obtained by target detection to obtain the final pixel coordinates of the center of the charging port. The actual distance from the center of the charging port to the camera is calculated through the pixel size of the charging port in the second image and the camera intrinsic parameter, and the final pixel coordinates of the center of the charging port, i.e. the actual distance from the center of the charging port to the camera, is determined as the target position of the target charging port.
[0144] Step S203, based on the target position, control the charging plug to move to the preset position of the target charging port.
[0145] The preset position is a position with a set distance from the target position. The distance is to avoid collision between the plug and the vehicle, and to provide a close enough observation distance for subsequent fine positioning. Illustratively, the set distance is 200mm. The target position is converted from the image coordinate system to the charging robot base coordinate system, and then the charging plug is moved to the preset position, so that it is 200mm away from the target charging port, completing the coarse positioning of the target charging port.
[0146] The embodiment of the application solves the problem of split between environment perception and image enhancement in traditional positioning, and positioning failure caused by poor image quality in complex environment by introducing a multi-task model to process the first image in the early stage of charging positioning. The multi-task model can output the picture quality classification result (reflecting the environment state) and the second image with improved quality at the same time, which not only clearly indicates the current environment (such as rain, fog), but also eliminates the influence of environmental interference on the image, providing a high-quality image basis for subsequent determination of the target position of the charging port. By determining the target position based on the second image and controlling the charging plug to move to the preset position, the misjudgment of the target position caused by the original image blur and low contrast can be avoided, and the accuracy of target positioning in complex environment is greatly improved, laying a foundation for subsequent accurate docking.
[0147] Step S204, after controlling the charging plug to move to the preset position of the target charging port, determining the initial pose of the target charging port and obtaining a binocular image containing the target charging port. The initial pose is the pose of the target charging port relative to the camera coordinate system corresponding to the binocular image. For details, see the above description of the embodiment of the application. Figure 1The description of step S101 shown in the figure will not be repeated here.
[0148] Step S205, binocular stereo matching is performed on the binocular image to obtain an initial depth map corresponding to the binocular image. For details, refer to the above description such as Figure 1 The description of step S102 shown in the figure will not be repeated here.
[0149] Step S206, based on the three-dimensional space coordinates of the key points in the preset charging port model, the initial depth map and the initial pose are updated alternately until the preset optimization condition is met, and the optimized initial pose is determined as the target pose of the target charging port; the preset charging port model and the target charging port are of the same type.
[0150] Specifically, the preset optimization condition includes that the pose change amount of the initial pose before and after optimization meets the preset pose change amount condition or the iteration number reaches the preset iteration number threshold.
[0151] The embodiment of the application solves the problem of mutual independence of depth map and pose calculation in traditional positioning, which is easy to cause error accumulation, by first correcting the depth map, then optimizing the pose using the corrected depth map, and then performing a loop operation. And each time the alternating optimization loop is based on the three-dimensional coordinates of the key points in the preset charging port model, the initial depth map is first corrected to eliminate the depth deviation caused by environmental interference, and then the initial pose is optimized based on the corrected depth map, and the optimization is terminated through the preset pose change amount condition and the iteration number threshold, which can not only ensure the optimization accuracy, but also avoid the system delay caused by excessive iteration. Compared with the traditional single positioning scheme, the positioning error can be gradually compressed, the final pose accuracy is significantly improved, and the system real-time performance is also considered, which meets the demand of fast response of the charging equipment.
[0152] Further, the step S206 of alternately updating the initial depth map and the initial pose based on the three-dimensional space coordinates of the key points in the preset charging port model includes:
[0153] Step a1, based on the initial pose, the three-dimensional space coordinates of the key points in the preset charging port model are mapped to the camera coordinate system, and projected to the image plane of the camera coordinate system to obtain the projection point coordinates and the theoretical depth corresponding to the key points.
[0154] The preset charging port model is a three-dimensional digital model completely consistent with the structure of a target charging port (such as a Type-C charging port) and contains three-dimensional space coordinates of seven fixed terminals (key points) of the charging port. By obtaining the three-dimensional space coordinates of the seven key points in the preset charging port model, the key point coordinates are mapped to the camera coordinate system, and the Z-axis coordinates of the key points in the camera coordinate system are the theoretical depths of the key points. Then, the three-dimensional coordinates of the key points in the camera coordinate system are converted into two-dimensional pixel coordinates by a projection function and a camera intrinsic matrix to obtain the above projection point coordinates. The conversion of the three-dimensional space coordinates from the model coordinate system to the camera coordinate system and the further projection to the image plane of the camera coordinate system are existing technologies, and will not be described in detail here.
[0155] In step a2, the first depth corresponding to the key point in the initial depth map is corrected based on the theoretical depth corresponding to the key point to obtain a corrected depth map.
[0156] The initial depth map is a distance map from all pixel points to the camera calculated by binocular stereo matching of left and right images taken by a binocular camera. The distance corresponding to the key point in the initial depth map is the first depth, but the first depth may have errors (such as rain and fog, which can cause inaccurate binocular matching, resulting in a larger or smaller first depth. The theoretical depth calculated in step a1 is converted from the three-dimensional space coordinates of the key points in the preset charging port model, so the accuracy of the theoretical depth is highly reliable. Taking the seven fixed terminals of the charging port as an example, the depth value corresponding to each fixed terminal in the initial depth map, i.e., the first depth, can be extracted from the initial depth map. Thus, the first depth with possible depth errors in the initial depth map is corrected using the highly accurate theoretical depth, and the corrected depth values of all key points are updated to the initial depth map to obtain a corrected depth map, so that the depth values of each key point in the corrected depth map are accurate, and the depth estimation error is reduced.
[0157] Specifically, by taking the above theoretical depth as a benchmark, the first depth corresponding to the key points in the initial depth map is corrected, the depth deviation of the initial depth map caused by matching noise and environmental interference is eliminated, and a higher-precision corrected depth map is generated, providing reliable depth data support for subsequent pose optimization. Specifically, by correcting the measured depth (first depth) affected by environmental interference to a depth closer to the true value, for example, in a rain and fog scene, the first depth of the key points in the initial depth map may deviate by 5mm, and after correction, the deviation can be reduced to within 0.5mm to eliminate depth measurement error. And because the key points are the core area of the charging port, after correcting the depth of the key points, the accuracy of the entire depth map can be improved (such as by interpolating the correction logic of the key points to the surrounding pixels), providing reliable depth data for subsequent more accurate pose calculation and improving the overall accuracy of the depth map. Even in low light, rain and fog, and other complex environments, because there is a theoretical depth as a benchmark, the problem of inaccurate depth map can be avoided, the pain point of traditional binocular depth estimation environment difference precision collapse is solved, and the environmental robustness is enhanced.
[0158] Further, the step a2 obtains the corrected first depth value by weighted sum of the theoretical depth and the first depth.
[0159] The present application corrects the depth value by weighting the theoretical depth and the first depth in the initial depth map, and dynamically allocates the weight according to the reliability of the two, for example, when the first depth is affected by environmental interference (such as foggy weather) and has low reliability, the weight of the theoretical depth is increased, and vice versa. This flexible weighting strategy can make full use of effective depth information, reduce the influence of single depth data deviation on positioning, make the corrected depth value more consistent with the actual scene, provide accurate depth basis for subsequent pose optimization, further improve the charging port positioning accuracy, and ensure the accuracy of the charging plug docking.
[0160] Step a3, based on the depth error of the key points between the initial depth map and the corrected depth map, and the re-projection error between the image coordinates corresponding to the key points in the monocular image and the projection point coordinates corresponding to the key points, the pose change is calculated. The monocular image is one of the binocular images.
[0161] Wherein, the depth error refers to the difference between the first depth of the same key point in the initial depth map and the depth in the corrected depth map (reflecting the change before and after the depth correction, and also reflecting the error size of the initial depth); the image coordinates corresponding to the key point in the monocular image are the 2D positions of the key point actually found from the monocular image; the re-projection error is the distance between the actual image coordinates and the projection point coordinates (which should be theoretically in the position) calculated in step a1 (for example, the actual image coordinates are (100, 200) pixels, and the projection point coordinates are (102, 201) pixels, and the re-projection error is the distance between the two points).
[0162] Then, by combining the "depth error" and the "re-projection error", the pose change amount is calculated, which can reflect how much the initial pose needs to be adjusted (such as rotating the angle by 0.1°, moving the distance by 0.2mm) to make the depth more accurate and the key point position in the image more consistent with the theoretical position. By considering both "image position error" and "depth error", the problem of "looking correct but actually being far apart" caused by traditional positioning only looking at a single error (such as only looking at the image position and ignoring the depth) is avoided, and the pose adjustment is more comprehensive. By quantifying the pose adjustment amplitude, a clear pose change amount is obtained, rather than adjusting by experience, ensuring that each pose optimization has a precise direction and amplitude, such as accurately calculating that the pose of the charging port needs to be rotated by 0.3° and translated by 0.5mm, rather than roughly adjusting by experience.
[0163] Specifically, by constructing a pose change amount target function that combines depth error and re-projection error, a nonlinear optimization algorithm is used to solve the pose change amount, which realizes accurate correction of the initial pose and solves the one-sided problem of traditional single error index in pose optimization.
[0164] Step a4, based on the initial pose and the pose change amount, an optimized initial pose is obtained, and the above step a1 is returned.
[0165] Specifically, by adjusting the pose according to the pose change amount based on the initial pose, the optimized initial pose can be obtained. By adjusting the initial pose according to the pose change amount, the adjusted and more accurate charging port pose is obtained, which is the optimized initial pose.
[0166] Further, since the pose is composed of a rotation matrix R and a translation vector t, the specific optimization process is to update these two parameters, wherein the rotation matrix update: the initial rotation matrix is R 0, the rotation adjustment amount in the pose change amount corresponds to a small rotation matrix Δ R , and the optimized rotation matrix Rnew = R0 x Δ R (Rotation matrix multiplication represents rotation superposition). Translation vector update: the initial translation vector is t 0, the translation adjustment in the pose change is Δ t , the optimized translation vector tnew = t 0 + Δ t (Translation vector addition represents position superposition).
[0167] After the foregoing depth correction and multi-error optimization, the optimized pose can greatly reduce the error of the initial pose, for example, the initial pose can have a translation error of ±5mm and a rotation error of ±1°, and after optimization, the translation error can be reduced to ±0.5mm and the rotation error can be reduced to ±0.3°, achieving "sub-millimeter" positioning accuracy, which is required for automatic charging of electric vehicles, otherwise the charging plug cannot be inserted or the interface can be damaged.
[0168] The optimized pose is the final instruction for adjusting the position of the charging plug, and provides accurate basis for the final docking. Based on this accurate positioning, the charging equipment can accurately control the movement of the plug to ensure that the plug and the charging port are perfectly aligned, solving the problem of docking failure caused by insufficient positioning accuracy in the prior art. Moreover, if the error after one optimization does not meet the requirements, the optimized initial pose can be taken as a new initial pose, and the foregoing depth correction, error calculation and pose adjustment steps can be repeated until the error is small enough to meet the charging requirements (for example, 3-5 iterations can achieve sub-millimeter accuracy), further ensuring the positioning reliability in complex environments.
[0169] The embodiment of the present application maps the three-dimensional space coordinates of the key points in the preset charging port model to the camera coordinate system and obtains the theoretical depth, thereby correcting the initial depth map, solving the depth error caused by texture loss and occlusion in binocular matching; and then combining the depth error and the re-projection error to calculate the pose change, the deviation in both the depth and the image coordinates can be comprehensively considered, and the pose deviation caused by a single error index can be avoided. The process can accurately position the charging port, reduce the docking failure caused by inaccurate depth and coordinate deviation, and still maintain high positioning accuracy in the case of local occlusion and weak texture of the charging port, thereby improving the reliability of the charging docking.
[0170] Further, the method provided by the embodiment of the present application further includes the following steps:
[0171] Step b1, obtaining the confidence of the first depth of the key points in the initial depth map.
[0172] Specifically, the confidence of the first depth of the key points can be extracted from the confidence map output by the Unimatch depth estimation algorithm, thereby providing a quantitative basis for subsequent dynamic weight allocation, and solving the problem of subjectivity in traditional weight setting.
[0173] Exemplarily, the Unimatch depth estimation algorithm outputs a reliability evaluation of the depth value of each pixel in the initial depth map, with a value range of [0, 1], and the value closer to 1 indicates that the depth value is more reliable (such as the charging port terminal edge area), and the value closer to 0 indicates that the depth value is less reliable (such as the metal reflection, rain and fog shielding area), which is generated by the algorithm through feature matching consistency, disparity continuity and other indicators, which will not be expanded here.
[0174] Step b2, determining a first weight corresponding to the first depth based on the confidence of the first depth, the first weight being positively correlated with the confidence.
[0175] Step b3, determining a second weight corresponding to the theoretical depth based on the first weight.
[0176] Specifically, the weight of the first depth and the theoretical depth can be dynamically allocated according to the obtained depth confidence, the greater the depth confidence, the greater the weight of the corresponding first depth, and the smaller the weight of the corresponding theoretical depth, the sum of the first weight and the second weight being 1, such as by using a Sigmoid function as the mapping relationship between the depth confidence and the first weight, to solve the problem of insufficient accuracy caused by traditional fixed weight correction depth, and make the depth correction more consistent with the actual reliability of the data.
[0177] Step b4, determining a calculation weight corresponding to the depth error based on the confidence of the depth, the calculation weight being positively correlated with the confidence.
[0178] Specifically, by dynamically adjusting the weight of the depth error in calculating the pose change based on the depth confidence, it is ensured that the high-confidence depth data dominates the optimization and the low-confidence data has reduced influence, thereby solving the optimization deviation problem caused by traditional fixed weight.
[0179] Exemplarily, the relationship between the confidence and the calculation weight can be determined by a piecewise linear mapping, such as: when the confidence is greater than or equal to 0.8 and is in a high confidence, the corresponding calculation weight can be set to 1, indicating that the depth error fully participates in the optimization and dominates the depth consistency constraint; when the confidence is greater than or equal to 0.5 and less than 0.8 and is in a medium confidence, the corresponding calculation weight can be linearly increased according to a linear function, such as when the confidence is 0.66, the calculation weight is 0.72; when the confidence is less than 0.5 and is in a low confidence, the corresponding calculation weight can be set to 0.4, to weaken the interference of low-reliability depth error on the optimization, thereby ensuring the dominance of high-confidence data and avoiding the complete failure of low-confidence data, and balancing the optimization accuracy and robustness.
[0180] This invention addresses the problems of traditional weight allocation, such as lack of reliable data and strong subjective assumptions, by introducing depth confidence scores corresponding to key points in the initial depth map. By obtaining the confidence score of the first depth, the first weight is positively correlated with the confidence score, ensuring that the first depth with high confidence plays a greater role in depth correction. Simultaneously, the second weight of the theoretical depth and the calculation weight of the depth error are reasonably determined. For example, in the reflective metal area of the charging port, the first depth confidence score is low, corresponding to a small first weight and a small depth error calculation weight, reducing the interference of unreliable depth data in this area on optimization. This dynamic weight adjustment strategy based on confidence scores allows the optimization process to better reflect the actual data quality, improving the accuracy of depth correction and pose optimization, and enhancing the system's adaptability to complex environments.
[0181] For example, the pose change is calculated using the following formula (1):
[0182] (1)
[0183] in, Indicates the change in pose. Indicates the first The image coordinates of each key point in a monocular image. Indicates the first The coordinates of the projection points corresponding to each key point Indicates the first The three-dimensional spatial coordinates of the key points R This indicates that the rotation matrix is used to describe the orientation of the target charging port. t This indicates that the translation vector is used to describe the position of the target charging port. Represents the projection function. Represents the camera intrinsic parameter matrix. Indicates the first The calculation weights corresponding to the depth errors of each key point Indicates the first The depth values corresponding to the key points in the initial depth map. Indicates the first The depth values corresponding to the key points in the corrected depth map.
[0184] This invention maps the 3D spatial coordinates of key points in a preset charging port model to the camera coordinate system to obtain the theoretical depth, thereby correcting the initial depth map and resolving depth errors caused by texture loss and occlusion in binocular matching. Furthermore, by combining depth error and reprojection error to calculate pose change, it comprehensively considers deviations in both depth and image coordinates, avoiding pose deviations caused by a single error metric. This process can accurately locate the charging port, reducing docking failures caused by inaccurate depth and coordinate offsets. It maintains high positioning accuracy even in scenarios with partial occlusion of the charging port and weak texture, improving the reliability of charging docking.
[0185] Exemplarily, through the detailed implementation scheme of the auxiliary positioning algorithm combined with the binocular depth estimation algorithm Unimatch and PNP (Perspective-n-Point). The scheme significantly improves the positioning accuracy and robustness in complex environments by fusing dense depth estimation and key point pose solving.
[0186] 1. System input and preprocessing.
[0187] Input data: synchronously collected binocular RGB image pair (left image), (right image). Camera intrinsic matrix K and baseline length B (obtained through calibration).
[0188] Preprocessing:
[0189] Image de-distortion: correct the image based on the camera distortion coefficient.
[0190] Light normalization: adaptive histogram equalization (CLAHE) processing to reduce the influence of light changes.
[0191] 2. Unimatch depth estimation module.
[0192] Core steps:
[0193] (1) Feature enhancement: use the TransformerEncoder encoder to extract multi-scale features and enhance the response to low-texture areas. The generation formula of left image feature and right image feature
[0194] (2)
[0195] and introduce position encoding (Positional Embedding) to retain spatial information. Position encoding is fused with the Transformer encoder through the following steps, and its core goal is to inject spatial position information into unordered Transformer features, especially in low-texture areas to enhance the model's understanding of spatial relationships. The following is the specific implementation process:
[0196] The generation of position encoding usually adopts a learnable parameterized method or a fixed form of trigonometric function;
[0197] Fusion of position encoding and features. The fusion of position encoding and the Transformer encoder usually occurs in the input embedding stage, and the specific steps are as follows:
[0198] Feature extraction: input left and right images and Extract basic features through a convolutional backbone network (e.g., ResNet-50) to obtain an initial feature map and Perform multi-scale processing (e.g., pyramid pooling or downsampling) on the feature map to generate multi-scale features and .
[0199] Add position encoding: add position encoding PE directly to the multi-scale features:
[0200] (3)
[0201] If the position encoding is learnable, and are independent learning parameters. If the position encoding is fixed (e.g., trigonometric functions), it is directly superimposed on the feature map.
[0202] Transformer encoder processing: input the enhanced features and into the Transformer encoder for self-attention calculation and feature interaction. Self-attention mechanism: capture long-range dependencies by calculating the similarity between features. Multi-head attention: use multiple attention heads in parallel to improve the model's ability to model different spatial relationships.
[0203] Role and advantages of position encoding: enhance the response of low-texture areas: in the charging port positioning scene, the surface of the charging port is often black or metal, with single texture. At this time, traditional convolutional networks have difficulty extracting effective features, while position encoding explicitly injects spatial information to help the model understand the relative positions of different regions in the feature map, improving the feature discrimination ability of low-texture areas. Improve the consistency of multi-scale features: when fusing multi-scale features, position encoding ensures that features at different scales have a unified spatial reference system, avoiding matching errors caused by scale differences. Support geometric constraints: in subsequent disparity calculation, the spatial information implied by position encoding helps the model more accurately predict the displacement relationship between pixels, reducing the impact of occlusion and mismatch.
[0204] Global matching and cost volume construction: calculate the feature similarity matrix , which is a four-dimensional tensor, where the first two dimensions ( H , W ) represent all pixel positions in the left image, and the last two dimensions ( H , W ) represent all pixel positions in the right image.
[0205] (4)
[0206] in, The left figure is located in space ( i , j The feature vector at the location (such as the feature output by the Transformer encoder). The right figure is located in space ( k , l The eigenvector at (). The position shown in the left figure ( i , j ) and the position of the right figure ( k , l The cosine similarity of the two pairs represents the confidence level of their match.
[0207] The probability distribution is generated using Softmax, and the point with the highest probability is selected as the matching point. The purpose of Softmax is to optimize the similarity matrix. S The specific steps to convert to a probability distribution are as follows:
[0208] Local window constraints:
[0209] Applying Softmax directly to the entire four-dimensional tensor would result in an explosion of computational complexity (e.g., for a 512×512 image). S The size of the tensor is =68,719,476,736, In actual implementation, a local search window is usually introduced, that is, for each position in the left figure ( i , j Only candidate locations within its parallax range in the right figure are considered. k , l ).
[0210] Softmax normalization:
[0211] Within a local window, similarity values Apply the Softmax function to generate the probability distribution:
[0212] (5)
[0213] in, τ It is the temperature scaling factor, which controls the "sharpness" of the probability distribution: a smaller value indicates a higher probability distribution. τ (like τ A value of 0.01 will concentrate the probability at the position of highest similarity, increasing certainty. A larger value... τ It will smooth the probability distribution and increase the robustness of matching. P ( i , j ,k , l ) represents the probability that the left image position ( i , j ) matches the right image position ( k , l ). By Softmax normalization, it is ensured that:
[0214] (6)
[0215] That is, the sum of the matching probabilities for each left image position is 1.
[0216] Disparity and depth calculation: Disparity map D Generated from matching offsets: D = argmax S Depth map Z is obtained by geometric conversion:
[0217] (7)
[0218] where, Z is the depth map, representing the distance of each pixel to the camera. B is the baseline length of the binocular camera (the distance between the centers of the two cameras). f is the focal length of the camera. D is the disparity map.
[0219] Optimization strategy: Multi-level refinement, iterative optimization of residual disparity at 1 / 8, 1 / 4 resolution to improve edge details. Occlusion handling, using Left-Right Consistency Check to mask occluded areas.
[0220] 3. Feature point extraction and matching.
[0221] Key point selection:
[0222] Use geometric stability criteria to filter feature points: extract planar or cylindrical surface patches (such as SIFT / SURF descriptors), exclude areas with single texture. Retain high-stability points (such as corner points, edge intersection points) through curvature analysis. Descriptor matching: for left image key points {pi}, narrow the search range in the right image through depth constraints: match ORB descriptors using Hamming Distance, and remove false matches using RANSAC.
[0223] 4. PNP pose solving module.
[0224] PNP algorithm is used to solve the pose to obtain the initial pose. The specific implementation process is the prior art, which will not be described here.
[0225] 5. Depth-pose joint optimization, see the specific description of step S206 above for details, which will not be described here.
[0226] Step S207: Based on the target pose, align the charging plug with the target charging port and insert it. See the above for details. Figure 1 The relevant description of step S104 shown will not be repeated here.
[0227] In some optional embodiments, the charging control method for the charging device provided in this embodiment of the invention further includes the following steps before performing step S205:
[0228] Step c1: Input the stereo images into the pre-trained multi-task model based on image quality classification and image quality enhancement to obtain the image quality classification result and the stereo images after image quality enhancement.
[0229] The pre-trained multi-task model based on image quality classification and image quality enhancement is described above and will not be repeated here.
[0230] This invention addresses the problems of large matching errors and low initial depth map accuracy caused by poor original image quality (such as blurring due to rain or fog) in traditional binocular matching by performing multi-task model processing on the binocular images before stereo matching. The multi-task model improves the quality of the binocular images, eliminates the influence of environmental interference on the left and right eye images, and ensures consistency and clarity between the binocular images, providing high-quality input for subsequent stereo matching. Simultaneously, the image quality classification results can help determine the current environment, providing a basis for adjusting subsequent optimization strategies. Compared to directly matching the original binocular images, this processing significantly reduces matching noise, improves the accuracy of the initial depth map, provides a more reliable depth foundation for subsequent alternating optimization updates, and further enhances the accuracy and stability of charging port positioning.
[0231] In practical applications, the main processes of automatic charging of charging equipment include:
[0232] 1. When the charging plug is in the initial position, a multi-task model based on image quality classification and image quality enhancement identifies the current environment (normal, rainy, foggy, snowy, etc.) and performs noise reduction processing to output a high-quality, clear image. For example, Figure 6 This is a comparison chart of image denoising effects.
[0233] 2. The charging port is first identified using the Deeplabv3+ semantic segmentation algorithm and the YOLOv8 object detection algorithm, and it is determined whether there are any obstacles obstructing it. If there are obstacles obstructing it, a voice prompt is given.
[0234] 3. Coarse positioning: By combining the PNP algorithm with the fixed relative position of the seven terminals of the charging port, we obtain the approximate position of the charging port and move the charging plug to the approximate position of the charging port, which is about 200mm away from the charging port.
[0235] 4. Further positioning of the charging port in the approximate position, due to the visual positioning error of the PNP algorithm. The Unimatch depth estimation algorithm of the binocular camera is combined with the PNP algorithm to obtain the accurate position of the charging port. Exemplarily, Figure 7 is a diagram of the binocular depth estimation effect.
[0236] 5. Move the charging plug to the final alignment position, slowly move the charging plug to the z-axis; with the cooperation of the pressure sensor, the final charging plug is successfully inserted into the charging port. Exemplarily, Figure 8 is a schematic diagram of the main process of the automatic charging of the charging device.
[0237] The Peak signal-to-noise ratio (PSNR) of the image output by the embodiment of the application in the rain, fog and snow scene is improved by more than 8dB, and the Structural Similarity Index (SSIM) is improved by more than 0.15; and the cross-task feature interaction of environment perception and image enhancement is realized by embedding a local attention module, the inference speed is improved by 40% compared with the traditional Transformer, and the real-time performance is guaranteed. In the coarse positioning stage: the PNP algorithm is constructed by fusing semantic segmentation (Deeplabv3+) and target detection (YOLOv8), and the initial positioning error is compressed to ±20mm by using the fixed spatial relationship of the seven terminals of the charging port; in the fine positioning stage: a Unimatch-PNP joint optimization mechanism is proposed, the binocular depth map and the PNP geometric prior depth are fused, and sub-millimeter level pose estimation (translation error <0.5mm, angle error <0.3°) is realized within 200mm distance, which breaks through the limit of plug-in accuracy. Dynamically adjust the image enhancement strategy and positioning algorithm parameters, and the system response delay is reduced to 80ms.
[0238] The charging control scheme of the charging device provided by the application has rain and fog penetration imaging capability: reconstructing high-frequency textures under a rain and fog transmittance < 50% scenario, outputting an enhanced image with MTF > 0.2 (the traditional scheme MTF ≈ 0). Metal reflection dynamic suppression: capturing the polarization angle distribution of the metal surface through the cross-task attention mechanism, reducing the overexposure rate of the reflection area from 78% to 9%, industrial-level verification: the success positioning rate of the stainless steel charging port is 100% under a 60° strong light incidence angle. Cross-scale continuous positioning: the accuracy is compressed without a fault, in the coarse positioning stage: the geometric constraint PNP compresses the 5-meter initial error from > 35 mm to ± 20 mm, in the fine positioning stage: the Unimatch-PNP joint optimization realizes 0.5 mm positioning within a distance of 200 mm. By fusing binocular sub-pixel displacement and geometric priori, the diffraction limit resolution is broken. System entropy reduction closed-loop architecture: dynamic feedback mechanism, real-time optimization of environmental perception, enhancement, and positioning parameters (80 ms / period), error chain blocking, image enhancement module reduces segmentation error by 60%, and pose transmission error rate is compressed from > 200% to < 30%. And it has the ability of self-adaptation to all physical states, the local attention mechanism makes the multi-task model reasoning delay only 35 ms (40% faster than the global attention), and the hardware resource reuse rate is increased by 3 times: the environment recognition and image enhancement share 90% of the computation graph.
[0239] In the embodiment, a charging control device of a charging device is also provided, which is used to implement the above-mentioned embodiments and preferred embodiments, and will not be described again. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, implementation of hardware, or a combination of software and hardware, is also possible and contemplated.
[0240] The embodiment of the application provides a charging control device of a charging device, and the charging device comprises a charging plug, as shown in the figure, and the device comprises: Figure 9
[0241] The first processing module 1101 is configured to determine an initial pose of the target charging port after controlling the charging plug to move to a preset position of the target charging port, and acquire a binocular image containing the target charging port, wherein the initial pose is a pose of the target charging port relative to a camera coordinate system corresponding to the binocular image;
[0242] The second processing module 1102 is configured to perform binocular stereo matching on the binocular image to obtain an initial depth map corresponding to the binocular image;
[0243] The third processing module 1103 is configured to update the initial depth map and the initial pose alternately based on the three-dimensional space coordinates of the key points in the preset charging port model until a preset optimization condition is met, and determine the optimized initial pose as a target pose of the target charging port; the preset charging port model is of the same type as the target charging port.
[0244] The fourth processing module 1104 is configured to control the charging plug to be inserted into the target charging port after the charging plug is aligned with the target charging port based on the target pose.
[0245] In some optional embodiments, the preset optimization condition includes that a pose change amount of the initial pose before and after optimization meets a preset pose change amount condition or an iteration number reaches a preset iteration number threshold.
[0246] In some optional embodiments, the third processing module 1103 includes:
[0247] The first processing subunit is configured to map the three-dimensional space coordinates of the key points in the preset charging port model to the camera coordinate system based on the initial pose, and project the three-dimensional space coordinates of the key points to the image plane of the camera coordinate system to obtain the projection point coordinates and the theoretical depth corresponding to the key points.
[0248] The second processing subunit is configured to correct the first depth of the key points in the initial depth map based on the theoretical depth corresponding to the key points to obtain a corrected depth map.
[0249] The third processing subunit is configured to calculate a pose change amount based on a depth error of the key points between the initial depth map and the corrected depth map and a re-projection error between the image coordinates corresponding to the key points in the monocular image and the projection point coordinates corresponding to the key points, the monocular image being one of the binocular images.
[0250] The fourth processing subunit is configured to obtain the optimized initial pose based on the initial pose and the pose change amount, and call the first subunit to run.
[0251] In some optional embodiments, the second processing unit is specifically configured to perform weighted summation on the theoretical depth and the first depth to obtain a corrected first depth value.
[0252] In some optional embodiments, the charging control device of the charging device further includes:
[0253] The second acquisition module is configured to acquire a confidence degree of the first depth of the key points in the initial depth map.
[0254] The ninth processing module is configured to determine a first weight corresponding to the first depth based on the confidence degree of the first depth, the first weight being positively correlated with the confidence degree.
[0255] The tenth processing module is used to determine the second weight corresponding to the theoretical depth based on the first weight;
[0256] And / or, the eleventh processing module is used to determine the computational weight corresponding to the depth error based on the confidence level of a depth, wherein the computational weight is positively correlated with the confidence level.
[0257] In some alternative implementations, the pose change is calculated using the following formula:
[0258]
[0259] in, Indicates the change in pose. Indicates the first The image coordinates of each key point in a monocular image. Indicates the first The coordinates of the projection points corresponding to each key point Indicates the first The three-dimensional spatial coordinates of the key points R This indicates that the rotation matrix is used to describe the orientation of the target charging port. t This indicates that the translation vector is used to describe the position of the target charging port. Represents the projection function. This represents the camera intrinsic parameter matrix. Indicates the first The calculation weights corresponding to the depth errors of each key point Indicates the first The depth values corresponding to the key points in the initial depth map. Indicates the first The depth values corresponding to the key points in the corrected depth map.
[0260] In some optional embodiments, the charging control device of the charging equipment provided in this invention further includes:
[0261] The first acquisition module is used to acquire a first image containing the target charging port;
[0262] The fifth processing module is used to input the first image into a pre-trained multi-task model based on image quality classification and image quality enhancement to obtain the image quality classification result and the second image after image quality enhancement. The image quality classification result is related to the environmental state of the target charging port.
[0263] The sixth processing module is used to determine the target location of the target charging port based on the second image;
[0264] The seventh processing module is used to control the charging plug to move to a preset position of the target charging port based on the target position. The preset position is a position at a set distance from the target position.
[0265] In some optional embodiments, the multi-task model based on the picture quality classification and the picture quality improvement includes an encoder, a decoder, and a picture quality classification branch.
[0266] The decoder includes a second feature dimension transformation layer, a concatenation layer, a second convolutional layer, and an up-sampling layer connected in sequence, the input end of the second feature dimension transformation layer is connected with the output end of the deep feature extraction network, the concatenation layer is further connected with the output end of the first feature dimension transformation layer, and the decoder is configured to perform a picture quality improvement task based on the input image features.
[0267] The picture quality classification branch includes a C3T module, a third convolutional layer, and a fully connected layer connected in sequence, the input end of the C3T module is connected with the output end of the feature dimension transformation layer, and the picture quality classification branch is configured to perform a picture quality classification task based on the input image features.
[0268] In some optional embodiments, the multi-task model based on the picture quality classification and the picture quality improvement further includes:
[0269] A first local attention module is arranged between the channel-spatial joint attention module and the first convolutional layer.
[0270] A second local attention module is arranged between the second convolutional layer and the up-sampling layer.
[0271] In some optional embodiments, the charging control device of the charging equipment further includes:
[0272] The eighth processing module is configured to input the binocular image into the pre-trained multi-task model based on the picture quality classification and the picture quality improvement respectively to obtain a picture quality classification result and a binocular image after picture quality improvement.
[0273] Further function descriptions of each module and unit are the same as those of the corresponding method embodiments, and will not be described here.
[0274] The embodiment of the present application further provides a charging equipment for automatically charging an electric vehicle, which includes a charging plug and a controller. Figure 10As shown, the controller includes one or more processors 10, memory 20, and interfaces 30 for interconnecting these components. The components communicate using a system bus or other suitable interconnect. The processor can process instructions for execution within the computer device, including instructions stored in the memory or on disk storage to display graphical information for a GUI on an external input / output device, such as a display device coupled to one of the interfaces. In some embodiments, multiple processors and / or multiple buses can be employed as needed to implement the described functionality. Figure 10 The processor 10 is used as an example.
[0275] The processor 10 can be a central processing unit, a network processor, or a combination thereof. The processor 10 can further include a hardware chip. The hardware chip can be an application specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device can be a complex programmable logic device, a field programmable logic device, a general array logic, or any combination thereof.
[0276] The memory 20 stores instructions that are executable by the at least one processor 10 to cause the at least one processor 10 to perform the methods described in the above embodiments.
[0277] The memory 20 can include a program storage area and a data storage area. The program storage area can store an operating system, application programs required by at least one function, and the like. The data storage area can store data created by the use of the computer device according to the display of a small program landing page, and the like. In addition, the memory 20 can include a high-speed random access memory, and can further include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some alternative embodiments, the memory 20 can optionally include a memory that is remotely located with respect to the processor 10, and these remote memories can be connected to the computer device through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0278] The memory 20 can include a volatile memory, such as a random access memory, and can also include a non-volatile memory, such as a flash memory, a hard disk, or a solid state disk. The memory 20 can further include a combination of the above-mentioned types of memories.
[0279] The controller further includes a communication interface 30 for communication of the vehicle with other devices or communication networks.
[0280] The embodiments of the present application further provide a computer readable storage medium, and the method according to the embodiments of the present application can be implemented in hardware, firmware, or recorded in a storage medium, or stored in a remote storage medium or a non-transitory machine readable storage medium and downloaded to a local storage medium through network, so that the method described herein can be processed by such software on a storage medium using a general purpose computer, a special purpose processor, or programmable or special hardware. The storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid state disk, etc. Further, the storage medium can also include a combination of the above-mentioned memories. It can be understood that the computer, the processor, the microprocessor controller, or the programmable hardware includes a storage component that can store or receive software or computer code, when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.
[0281] Part of the present application can be applied as a computer program product, for example, computer program instructions, when executed by a computer, through the operation of the computer, the method and / or technical solutions according to the present application can be invoked or provided. Those skilled in the art should understand that the form of computer program instructions in a computer readable medium includes but is not limited to source files, executable files, installation package files, etc. Correspondingly, the way of computer program instructions executed by computer includes but is not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Here, the computer readable medium can be any available computer readable storage medium or communication medium accessible to the computer.
[0282] Although the embodiments of the present application are described in conjunction with the accompanying drawings, various modifications and changes can be made by those skilled in the art without departing from the spirit and scope of the present application, and such modifications and changes fall within the scope defined by the appended claims.
Claims
1. A charge control method of a charging device, the charging device comprising: The charging plug is characterized in that the method comprises: After the charging plug is controlled to move to a preset position of the target charging port, an initial pose of the target charging port is determined, and a binocular image containing the target charging port is acquired, the initial pose being a pose of the target charging port relative to a camera coordinate system corresponding to the binocular image; Binocular stereo matching is performed on the binocular image to obtain an initial depth map corresponding to the binocular image; Based on three-dimensional space coordinates of key points in a preset charging port model, the initial depth map and the initial pose are sequentially updated alternately until a preset optimization condition is met, and the optimized initial pose is determined as a target pose of the target charging port; the preset charging port model is of the same type as the target charging port; After the charging plug is controlled to align with the target charging port based on the target pose and then be inserted; Based on the three-dimensional space coordinates of the key points in the preset charging port model, the initial depth map and the initial pose are sequentially updated alternately, comprising: Based on the initial pose, the three-dimensional space coordinates of the key points in the preset charging port model are mapped to the camera coordinate system, and are projected to an image plane of the camera coordinate system to obtain projection point coordinates and theoretical depths of the key points; Based on the theoretical depths of the key points, first depths of the key points in the initial depth map are corrected to obtain a corrected depth map; Based on depth errors of the key points between the initial depth map and the corrected depth map, re-projection errors between image coordinates of the key points in a monocular image and the projection point coordinates of the key points, a pose change amount is calculated, the monocular image being one of the binocular images; Based on the initial pose and the pose change amount, an optimized initial pose is obtained, and the step of mapping the three-dimensional space coordinates of the key points in the preset charging port model to the camera coordinate system based on the initial pose, and projecting to the image plane of the camera coordinate system to obtain the projection point coordinates and the theoretical depths of the key points is returned; Based on the theoretical depths of the key points, the first depths of the key points in the initial depth map are corrected, comprising: The theoretical depths and the first depths are weighted and summed to obtain a corrected first depth value; The method further comprises: A confidence of the first depth of the key points in the initial depth map is acquired; Based on the confidence of the first depth, a first weight corresponding to the first depth is determined, the first weight being positively correlated with the confidence; Based on the first weight, a second weight corresponding to the theoretical depth is determined.
2. The method of claim 1, wherein, The preset optimization condition comprises that a pose change amount of the initial pose before and after optimization meets a preset pose change amount condition or an iteration number reaches a preset iteration number threshold.
3. The method of claim 1, wherein, The method further comprises: Based on the confidence of the first depth, a calculation weight corresponding to the depth error is determined, the calculation weight being positively correlated with the confidence.
4. The method of claim 3, wherein, The pose change amount is calculated by the following formula: wherein, denotes a pose change amount, denotes the image coordinate of the th key point in the monocular image, denotes the projection point coordinate of the th key point, denotes the three-dimensional space coordinate of the th key point, R denotes a rotation matrix for describing the orientation of the target charging port, t denotes a translation vector for describing the position of the target charging port, denotes a projection function, denotes a camera intrinsic matrix, denotes the calculation weight corresponding to the th key point depth error, denotes the depth value of the th key point in the initial depth map, denotes the depth value of the th key point in the corrected depth map.
5. The method according to any one of claims 1 to 4, characterized in that, The method further comprises: A first image containing the target charging port is acquired; inputting the first image into a pre-trained multi-task model based on picture quality classification and picture quality improvement to obtain a picture quality classification result and a second image after picture quality improvement, the picture quality classification result being related to an environmental state in which the target charging port is located; determining a target position of the target charging port based on the second image; controlling the charging plug to move to a preset position of the target charging port, the preset position being a position with a set distance from the target position.
6. The method of claim 5, wherein, The multi-task model based on picture quality classification and picture quality improvement comprises an encoder, a decoder, and a picture quality classification branch. The encoder comprises a deep feature extraction network, a channel-spatial joint attention module, a first convolutional layer, a stacking operation layer, and a first feature dimension transformation layer connected in sequence, and an input end of the deep feature extraction network is configured to receive the first image. The decoder comprises a second feature dimension transformation layer, a splicing layer, a second convolutional layer, and an up-sampling layer connected in sequence, an input end of the second feature dimension transformation layer is connected with an output end of the deep feature extraction network, the splicing layer is further connected with an output end of the first feature dimension transformation layer, and the decoder is configured to perform a picture quality improvement task based on input image features. The picture quality classification branch comprises a C3T module, a third convolutional layer, and a fully connected layer connected in sequence, an input end of the C3T module is connected with an output end of the feature dimension transformation layer, and the picture quality classification branch is configured to perform a picture quality classification task based on input image features.
7. The method of claim 6, wherein, The multi-task model based on picture quality classification and picture quality improvement further comprises: a first local attention module is arranged between the channel-spatial joint attention module and the first convolutional layer; a second local attention module is arranged between the second convolutional layer and the up-sampling layer.
8. The method of claim 5, wherein, Before performing binocular stereo matching on the binocular image to obtain an initial depth map corresponding to the binocular image, the method further comprises: inputting the binocular image into a pre-trained multi-task model based on picture quality classification and picture quality improvement respectively to obtain a picture quality classification result and a binocular image after picture quality improvement.
9. A charging control device of a charging apparatus, the charging apparatus comprising: A charging plug, characterized in that the device comprises: a first processing module configured to determine an initial pose of the target charging port and acquire a binocular image containing the target charging port after controlling the charging plug to move to a preset position of the target charging port, the initial pose being a pose of the target charging port relative to a camera coordinate system corresponding to the binocular image; a second processing module configured to perform binocular stereo matching on the binocular image to obtain an initial depth map corresponding to the binocular image; a third processing module configured to sequentially perform alternating optimization and update on the initial depth map and the initial pose based on three-dimensional space coordinates of key points in a preset charging port model until a preset optimization condition is met, and determine an optimized initial pose as a target pose of the target charging port, the preset charging port model being of the same type as the target charging port. The fourth processing module is configured to control the charging plug to be inserted into the target charging port after the charging plug is aligned with the target charging port based on the target pose. The third processing module includes: a first processing subunit, configured to map three-dimensional space coordinates of key points in a preset charging port model to a camera coordinate system based on the initial pose, and project the three-dimensional space coordinates to an image plane of the camera coordinate system to obtain projection point coordinates and theoretical depths corresponding to the key points; a second processing subunit, configured to correct first depths corresponding to the key points in an initial depth map based on the theoretical depths corresponding to the key points to obtain a corrected depth map; a third processing subunit, configured to calculate a pose change based on a depth error between the key points in the initial depth map and the corrected depth map and a re-projection error between image coordinates corresponding to the key points in a monocular image and the projection point coordinates corresponding to the key points, the monocular image being one of the binocular images; and a fourth processing subunit, configured to obtain an optimized initial pose based on the initial pose and the pose change, and call the first processing subunit to run. The second processing unit is specifically configured to perform weighted summation on the theoretical depth and the first depth to obtain a corrected first depth value. The charging control device of the charging equipment further includes: The second acquisition module is configured to acquire a confidence degree of the first depth corresponding to the key points in the initial depth map. The ninth processing module is configured to determine a first weight corresponding to the first depth based on the confidence degree of the first depth, the first weight being positively correlated with the confidence degree. The tenth processing module is configured to determine a second weight corresponding to the theoretical depth based on the first weight.
10. A charging device, characterized by The charging equipment includes a charging plug and a controller, and the controller includes: A memory and a processor, which are in communication connection with each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the method in any one of claims 1 to 8.
11. A computer readable storage medium, characterized in that, The computer readable storage medium stores computer instructions, and the computer instructions are used to make a computer execute the method in any one of claims 1 to 8.
Citation Information
Patent Citations
Method and device for estimating pose of electric vehicle charging socket and autonomous charging robot employing the same
US20230102948A1
Pose determination method and apparatus, electronic device, and readable storage medium
WO2023016182A1