Positioning method and device for a body-equipped robot to charge a vehicle
Patent Information
- Application Number
- CN202611289553.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-25
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]采取基于人工辅助的半自主方案时,当车辆停放朝向发生变化时,具身机器人无法利用车型特征预判充电口方位,导致搜索路径规划冗余
建立具身智能主动感知闭环:通过观测置信度评估触发主动视点规划,驱动具身机器人自主移动至最佳观测位姿,解决视觉模糊与遮挡导致的定位失败问题,提升复杂环境下的任务成功率。
Smart Images

Figure CN122808519A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of control system technology, and in particular to a method and device for positioning a vehicle charging using an embodied robot. Background Technology
[0002] With the rapid development of the new energy vehicle industry, automated charging robots, as key equipment connecting smart grids and intelligent transportation, have increasingly broad application prospects. In practical application scenarios of various charging stations, automated charging robots need to autonomously locate target vehicles and complete the docking task of charging guns in unstructured environments.
[0003] Current navigation and positioning solutions for automated charging robots mainly fall into two categories: The first is based on high-precision maps and pre-defined locations, requiring vehicles to park in precisely designated parking spaces at the charging station. The robot then moves to these fixed coordinates according to a fixed map path. The second is a semi-autonomous solution assisted by a human, requiring the driver to manually confirm vehicle type, parking orientation, and other information via mobile software or an in-vehicle terminal. The robot then invokes the corresponding motion planning program based on the driver's instructions.
[0004] However, when embodied robots move to fixed coordinates along fixed map paths to perform tasks, they cannot observe the charging port by adjusting distance or angle like humans do when it is invisible due to distance, poor lighting, or being obscured by a cover, leading to direct task failure. Furthermore, existing visual perception algorithms are mostly based on pure geometric detection or end-to-end regression. If the charging port feature is not directly observed within the field of view, the system often cannot indirectly locate the charging port using logical reasoning from other exterior components of the vehicle (such as headlights or logos).
[0005] When adopting a semi-autonomous solution based on human assistance, the robot cannot predict the location of the charging port using vehicle model characteristics when the vehicle's parking orientation changes, resulting in redundant search path planning. Summary of the Invention
[0006] To address the aforementioned issues, this application proposes a method for embodying robots to locate vehicle charging, comprising: acquiring environmental perception information of the embodying robot in its current observation pose; The environmental perception information is identified to obtain the vehicle model identifier, the location of vehicle exterior components, and the observation confidence level corresponding to the identification result; If the observation confidence level is lower than the preset condition, the robot is controlled to adjust its observation posture, reacquire environmental perception information and perform identification again until the observation confidence level meets the preset condition. If the observation confidence level meets the preset conditions, the spatial transformation relationship associated with the vehicle model identifier is retrieved from the pre-constructed topological model; Based on the spatial transformation relationship and the position of the vehicle exterior components, the location of the charging port can be deduced.
[0007] In one example, the identification of the environmental perception information specifically includes: Multimodal features of color and depth images are extracted, and point cloud features of point cloud data are also extracted; depth images are obtained by aligning with color images, and point cloud data are obtained by transforming depth images. The point cloud features and the image multimodal features are fused using multi-scale attention to obtain joint features; The joint features are classified globally using a global semantic branch network to obtain a vehicle classification probability vector, and the vehicle identifier and classification confidence are determined based on the vehicle classification probability vector. The joint features are identified by a local discriminant branch network to obtain heat maps corresponding to multiple semantic key points. The location of each corresponding appearance component is determined based on each heat map, and the location confidence is calculated based on all heat maps. The classification confidence score and the location confidence score are weighted and fused to obtain the observation confidence score.
[0008] In one example, the classification confidence is determined based on the vehicle model classification probability vector, specifically including: The classification confidence level is obtained by calculating the information entropy of the vehicle classification probability vector; Location reliability is calculated based on all heatmaps, specifically including: The location reliability is obtained by calculating the average of the peak responses of each heatmap.
[0009] In one example, the multimodal feature extraction module includes a first convolutional layer, a first pooling layer, a second convolutional layer, and a second pooling layer; The point cloud feature extraction module includes a first multi-layer perceptron layer, a max pooling layer, and a second multi-layer perceptron layer; The multi-scale attention fusion module includes a feature concatenation layer, a self-attention mechanism layer, and a convolutional layer; The global semantic branch network includes a global average pooling layer, a fully connected layer, and a softmax layer; The local discriminative branch network includes upsampling layers, convolutional layers, and sigmoid layers.
[0010] In one example, the global semantic branch network and the local discriminative branch network are trained using the following composite loss function: Classification loss is used to increase the inter-class distance between different vehicle types in the feature space; Keypoint regression loss is used to calculate the pixel deviation between the predicted keypoint heatmap and the true label. Uncertainty estimation loss is used to constrain the variance of the predicted distribution of the outputs of the global semantic branch network and the local discriminative branch network, so that the network has the ability to quantify the uncertainty of observation.
[0011] In one example, controlling the embodied robot to adjust its observation pose specifically includes: Simulate multiple candidate observation poses in a local grid image; Predict the observation confidence for each candidate observation pose, and calculate the increase between the predicted observation confidence and the actual observation confidence. The candidate observation pose with the largest increase in confidence level is selected as the next best observation point.
[0012] In one example, predicting the observation confidence for each candidate observation pose specifically includes: Based on the pose transformation relationship between the current observation pose and each candidate observation pose, as well as the identified vehicle exterior component information, the observability of each exterior component under the candidate observation pose is predicted; observability is used to characterize the observation effect of each exterior component under the candidate observation pose. Based on observability, predict the observation confidence level corresponding to each candidate observation pose.
[0013] In one example, the topological model is a dynamic semantic knowledge graph, which includes: The node set includes a root node representing the vehicle model identifier, observation nodes representing exterior components, and inference nodes representing the charging port. The edge set includes rigid transformation matrices representing the relationships between nodes; for each vehicle category, the rigid transformation matrix of each exterior component relative to the charging port is pre-stored.
[0014] In one example, the location of the charging port is inferred based on the spatial transformation relationship and the position of the vehicle exterior components, specifically including: Using the camera intrinsic parameters, the pixel coordinates of the appearance component in the image are back-projected into three-dimensional space to obtain the coordinates of the appearance component in the camera coordinate system; Using a pre-calibrated hand-eye transformation matrix, the coordinates of the appearance component in the camera coordinate system are transformed to obtain the coordinates of the appearance component in the embodied robot coordinate system; Based on the rigid body transformation matrix from the appearance component to the charging port stored in the topological model, and the coordinates of the appearance component in the embodied robot coordinate system, the position of the charging port in the embodied robot coordinate system is inferred.
[0015] On the other hand, embodiments of this application provide a positioning device for a robot charging a vehicle, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a positioning method for charging a vehicle as described above.
[0016] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects: Establish an embodied intelligent active perception closed loop: trigger active viewpoint planning by evaluating observation confidence, drive the embodied robot to move autonomously to the optimal observation pose, solve the localization failure problem caused by visual blur and occlusion, and improve the success rate of tasks in complex environments.
[0017] Achieving concealed target reasoning based on topology models: By combining easily observable prominent components of the vehicle's exterior with spatial topological constraints, the precise pose of blind spots or obscured charging ports can be derived. This overcomes the limitations of traditional methods that rely on direct visual detection or manual assistance, and adapts to the uncertainty of vehicle parking orientation, reducing redundancy in search path planning. Attached Figure Description
[0018] To more clearly illustrate the technical solution of this application, some embodiments of this application will be described in detail below with reference to the accompanying drawings, in which: Figure 1 This is a positioning model diagram of an embodied robot charging a vehicle, provided as an embodiment of this application.
[0019] Figure 2 A schematic flowchart illustrating a method for locating a vehicle charging device using an embodied robot, as provided in an embodiment of this application. Figure 3 A diagram of a fine-grained visual perception network structure with uncertainty estimation provided in an embodiment of this application.
[0020] Figure 4 A framework diagram of a positioning system for vehicle charging by an embodied robot provided in an embodiment of this application; Figure 5 This is a schematic diagram of a positioning device for charging a vehicle using a robot, provided as an embodiment of this application. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0022] Some embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0023] Figure 1 This is a positioning model diagram of an embodied robot charging a vehicle, provided as an embodiment of this application.
[0024] The hardware structure of an embodied robot includes at least the following four functional modules, which work together through electrical connections and a data communication bus: Visual perception module: This module is installed at the front of the embodied robot body and is used to collect environmental perception information of the vehicle in real time.
[0025] For example, considering the strong light conditions in outdoor charging scenarios and the richness of vehicle surface textures, the visual perception device uses a binocular RGB-D depth camera. This camera includes two image sensors, left and right, which can simultaneously acquire left and right RGB color images and use stereo matching algorithms to calculate and generate depth images aligned with the color images.
[0026] Furthermore, the depth image is converted into 3D point cloud data. Each pixel in the depth image stores the distance (depth value) from that point to the camera plane. Therefore, by combining the camera intrinsic parameters (focal length, optical center coordinates), the coordinates of any pixel in the depth image and its depth value are calculated to obtain the corresponding 3D spatial coordinates of that pixel.
[0027] Based on this, environmental perception information includes color images, depth images, and point cloud data converted from depth images.
[0028] Motion module: including omnidirectional mobile chassis and underlying controller, which is responsible not only for moving to the work area, but also for actively adjusting the observation perspective based on the decision results during the perception phase.
[0029] Robotic arm execution module: Mounted on the upper platform of the omnidirectional mobile chassis, it performs the final charging gun docking operation. For example, a 6-axis collaborative robotic arm mounted on the mobile chassis.
[0030] Decision computing module: It adopts a high-performance embedded industrial control computer to run a visual network with confidence evaluation, maintain a semantic knowledge graph, and perform computational tasks such as calculating the next observation point, semantic reasoning, and navigation guidance information.
[0031] Based on this, the collaborative relationships between the modules are as follows: The visual perception module perceives color and depth images, converts the depth images into point cloud data, and sends the color images, depth images, and point cloud data to the decision computing module.
[0032] The decision calculation module runs a visual perception network with uncertainty estimation and outputs vehicle model recognition results, exterior component locations, and observation confidence levels.
[0033] If the confidence level is lower than the confidence level threshold, the next best observation point in the exploration mode is generated, and the embodied robot is driven by the motion module to adjust its pose, forming an embodied active perception closed loop.
[0034] Once the confidence level meets the requirements, the decision calculation module uses the semantic graph to infer the location of the charging port, obtains the positioning information of the charging port based on the location, and sends the positioning information to the motion module and the robotic arm execution module.
[0035] The motion module drives the embodied robot chassis to move to the work area, and the robotic arm execution module performs charging gun docking based on the final positioning information.
[0036] In summary, by combining the uncertainty estimation of deep neural networks with semantic knowledge graphs, a closed-loop system of perception-evaluation-motion-reasoning is constructed. The embodied robot is no longer a passive camera carrier, but can achieve autonomous and precise positioning of charging ports for different car models.
[0037] More intuitively, Figure 2 This is a flowchart illustrating a method for locating a vehicle charging station using an embodied robot, as provided in an embodiment of this application.
[0038] Figure 2 The process includes the following steps: S201: Obtain environmental perception information of the embodied robot in its current observation pose.
[0039] The current observed pose refers to the combination of the embodied robot's spatial position and spatial attitude at a certain moment. The spatial position is usually represented by three-dimensional coordinates in the embodied robot coordinate system (or world coordinate system), and the spatial attitude is usually represented by the rotation angle about the vertical axis, the pitch angle about the horizontal axis, and the roll angle, which are used to reflect the embodied robot's orientation and tilt state.
[0040] For example, environmental perception information can be defined as follows:
[0041] in, The visual perception state input (environmental perception information) collected by the embodied robot at time t. High-resolution color images captured by the RGB-D camera of the embodied robot. For the aligned depth image, This is local 3D point cloud data converted from a depth map.
[0042] S202: Identify the environmental perception information to obtain the vehicle model identifier, the location of the vehicle exterior components, and the observation confidence level corresponding to the identification result.
[0043] The specific form of the vehicle model identifier can be a numerical code (such as 001), a category label (such as Type-A), or a vehicle model name. It should be noted that this application does not limit the vehicle model identifier to correspond to a real-world vehicle brand or model, as long as the correct spatial topological relationship can be retrieved from the semantic knowledge graph.
[0044] Vehicle exterior components refer to prominent parts of a vehicle that have stable geometric features and are easily observed from a distance or side view. Examples include the left and right headlights, the vehicle emblem, the left and right side mirrors, and the wheel hub centers.
[0045] The location of a vehicle exterior component refers to the pixel position of a salient component in an image. It should be noted that the location can be two-dimensional pixel coordinates, i.e., coordinates in a color image. In this application, identifying the location of at least one exterior component is sufficient, but identifying multiple exterior components helps improve the robustness and accuracy of subsequent charging port inference.
[0046] The identification result refers to the vehicle type identifier and the location of vehicle exterior components. Observation confidence is a quantitative measure of the reliability of the current identification result, reflecting the credibility of the vehicle type identification and vehicle exterior component location obtained by the embodied robot in the current observation pose. Observation confidence is a scalar value (usually normalized to the [0,1] interval). The higher the value, the clearer and more stable the current environmental perception information, and the more reliable the charging port location inferred subsequently.
[0047] It should be noted that the observation confidence score is output along with the vehicle model identification and the location of the exterior components. That is, the vehicle model identifier, the location of the vehicle exterior components, and the observation confidence score are three results output in parallel by the same recognition model.
[0048] For recognition models, deep learning network models with uncertainty estimation (such as CNN, Transformer, etc.) can be used, but other methods can also be used, such as: Classic machine learning methods, for example, extracting HOG, LBP or Haar features and then inputting them into random forest or support vector machine for classification and detection, using the probability estimate output by the classifier as the confidence level.
[0049] Geometric and prior knowledge reasoning: For example, identifying vehicle models by measuring vehicle size proportions through point clouds or depth maps, locating headlights through color segmentation and symmetry analysis, and assessing confidence through occlusion proportions and edge continuity.
[0050] S203: If the observation confidence level is lower than the preset condition, control the embodied robot to adjust the observation posture, reacquire environmental perception information and perform identification again until the observation confidence level meets the preset condition.
[0051] Embossed robots possess embodied active perception capabilities, meaning that when the perception reliability of the charging port under the current observation pose is insufficient, the embodied robot can autonomously adjust its own position and / or posture to obtain clearer and more reliable environmental perception information.
[0052] The preset condition can be a pre-defined confidence threshold. For example, a fixed threshold T (such as T=0.75) can be set when the observed confidence level... At that point, the current perception is considered unreliable. Furthermore, preset conditions can include judgments across multiple dimensions. For example, it may require not only that the observation confidence level be higher than a confidence threshold, but also that at least two or more external components are detected.
[0053] Adjusting the observation pose refers to the embodied robot obtaining a better observation perspective of the target vehicle by changing its spatial position and / or spatial attitude.
[0054] Adjustment strategies may include the following: Planned adjustments: Based on principles such as maximizing information gain, improving confidence in prediction, or minimizing motion costs, the next optimal observation point is calculated in the local map, and then the embodied robot is driven to move to that point. For example, if the current distance is too far, forward movement is planned; if the current viewpoint is skewed, detour to the front or side of the vehicle is planned.
[0055] Trial adjustment: Make trial moves according to the preset search pattern (such as right-forward-left-turn, etc.) until the observation confidence improves.
[0056] Rule-based adjustments: For example, if no headlights are detected, lower the height and move closer to the front of the vehicle.
[0057] After the embodied robot reaches a new observation pose, it first stops moving and stabilizes before triggering the visual perception module to collect data again. It should be noted that images are acquired after the robot has stopped moving to avoid motion blur that could degrade the perception quality.
[0058] The process of adjusting pose, re-acquiring data, and re-identifying can be repeated multiple times, forming a perception and motion closed loop. Each iteration makes decisions based on the latest environmental perception information.
[0059] It should be noted that if, after several adjustments, the observation confidence level is still lower than the preset condition when the maximum number of attempts is reached (e.g., 3 times), the attempt can be abandoned and an error will be reported.
[0060] S204: If the observation confidence level meets the preset conditions, retrieve the spatial transformation relationship associated with the vehicle model identifier from the pre-constructed topological model.
[0061] The topology model is a pre-stored spatial relationship database that records the relative spatial positions of various exterior components on a vehicle and the charging port for different vehicle models.
[0062] It should be noted that the core function of the topology model is to return the spatial transformation relationship from each exterior component of a given vehicle model to the charging port, given a vehicle model identifier. The specific implementation of the topology model can be a semantic knowledge graph, a relational database, a JSON file, or other data structures.
[0063] The associated spatial transformation relationship refers to the mathematical relationship that describes the spatial positional transformation between vehicle exterior components and the charging port, which is bound to a specific vehicle model identifier. This relationship can be expressed as a rigid body transformation matrix (including rotation and translation), for example, the transformation matrix from the left front headlight coordinate system to the charging port coordinate system.
[0064] It should be noted that a vehicle model identifier can be associated with multiple spatial transformation relationships.
[0065] The retrieval process is as follows: based on the vehicle model identifier, perform a search, query, or index operation in the topology model to obtain one or more pre-stored spatial transformation relationships.
[0066] It should be noted that if the record corresponding to the current vehicle model identifier does not exist in the topology model, an error can be reported.
[0067] S205: Based on the spatial transformation relationship and the position of the vehicle exterior components, infer the position of the charging port.
[0068] The reasoning process refers to using spatial transformation relationships and the positions of vehicle exterior components, through mathematical coordinate transformation calculations, to obtain the position of the charging port in the embodied robot's coordinate system, thereby achieving indirect location of invisible or obscured charging ports. For example, the chain rule can be used to reason about the position of the charging port in the embodied robot's coordinate system.
[0069] Furthermore, the position of the charging port in the embodied robot coordinate system is converted into the navigation target position in the world coordinate system and sent to the chassis motion controller, thereby driving the embodied robot to move to the charging operation area.
[0070] It should be noted that, although the embodiments in this application are based on... Figure 2Steps S201 to S205 will be described sequentially, but this does not mean that steps S201 to S205 must be performed in a strict order. The reason this embodiment follows this order is... Figure 2 The order in which steps S201 to S205 are described is provided to facilitate understanding of the technical solutions of the embodiments of this application by those skilled in the art. In other words, in the embodiments of this application, the order of steps S201 to S205 can be appropriately adjusted according to actual needs.
[0071] pass Figure 2 The method establishes an embodied intelligent active perception closed loop: by triggering active viewpoint planning through observation confidence assessment, the embodied robot is driven to move autonomously to the optimal observation pose, which solves the localization failure problem caused by visual blur and occlusion and improves the success rate of tasks in complex environments.
[0072] Achieving concealed target reasoning based on topology models: By combining easily observable prominent components of the vehicle's exterior with spatial topological constraints, the precise pose of blind spots or obscured charging ports can be derived. This overcomes the limitations of traditional methods that rely on direct visual detection or manual assistance, and adapts to the uncertainty of vehicle parking orientation, reducing redundancy in search path planning.
[0073] based on Figure 2 In addition to the method described herein, this application also provides some specific implementation schemes and extended schemes of the method, which will be further described below.
[0074] In one example, a fine-grained visual perception network with uncertainty estimation is constructed for environmental perception information recognition. This involves building a deep neural network with a global semantic branch and a local discriminative branch, and embedding uncertainty quantization modules at the output of each branch.
[0075] Based on this, the identification of the environmental perception information includes the following steps: Step 1: Extract the multimodal features of the color image and depth image, and extract the point cloud features of the point cloud data; the depth image is obtained by aligning it with the color image, and the point cloud data is obtained by transforming the depth image.
[0076] Multimodal features are extracted by performing convolution and pooling operations on color and depth images. Color images provide appearance semantics (such as headlight color and logo shape), while depth images provide spatial distance information.
[0077] By applying multilayer perceptron (MLP) and max pooling operations to point cloud data, sparse yet accurate 3D geometric features are extracted. Point cloud features can compensate for the accuracy degradation of depth images in distant or low-texture regions, enhancing the network's understanding of spatial structure.
[0078] It should be noted that, regarding the feature extraction network structure, the multimodal feature extraction network can use SwinTransformer as the backbone network to extract multi-scale feature maps corresponding to color images, thereby capturing semantic information at different levels. The point cloud feature extraction network can use PointNet to extract permutation-invariant global geometric features from disordered point clouds.
[0079] For example, the structure of a multimodal feature extraction network includes a first convolutional layer, a second pooling layer, and so on. The first convolutional layer is used to extract low-level local features of the input image, the first pooling layer is used for downsampling to expand the receptive field, the second convolutional layer is used to extract mid-level semantic features, and the second pooling layer is used for further downsampling to obtain a more global feature representation. This stacked structure enables the network to extract multi-level features from color and depth images, ranging from details to semantics.
[0080] The point cloud feature extraction network consists of a first MLP layer, a max pooling layer, and a second MLP layer. The first MLP layer is used to independently enhance the features of each point, the max pooling layer is used to perform max pooling on all points to obtain a global point cloud feature vector, and the second MLP layer is used to further transform the global feature vector.
[0081] Step 2: Perform multi-scale attention fusion of the point cloud features and the image multimodal features to obtain joint features.
[0082] The multi-scale attention fusion process includes steps such as feature concatenation, self-attention capture, and convolutional layer extraction. The specific process can be summarized as follows: For each scale, image features and point cloud features of the same spatial resolution are concatenated along the channel dimension to obtain a joint feature map with an increased number of channels. This step achieves preliminary alignment and merging of the two modalities.
[0083] The concatenated feature map is input into the self-attention mechanism module. This mechanism captures long-distance dependencies between any two locations in the feature map, thus enhancing the feature's representational power. For example, for each location in the feature map, its relevance weights to all other locations are calculated. Based on these weights, the features of all locations are weighted and aggregated, ensuring that the new features at each location incorporate the global contextual information of the entire feature map. Therefore, when a part of a component (such as a headlight) is partially occluded, the network can utilize information from other visible areas to assist in localization.
[0084] The feature map enhanced by self-attention is then passed through one or more convolutional layers (typically 1×1 or 3×3 convolutions). These convolutional layers adjust the number of channels to accommodate the input requirements of subsequent global classification heads or keypoint regression heads, thus obtaining fused features at this scale. This further mixes information from different locations and channels to extract higher-order semantic features.
[0085] By integrating features from multiple scales, a joint feature is obtained. The joint feature simultaneously encodes information such as image texture, color, depth information (from color and depth maps), precise geometric structure information of the point cloud (from point cloud branches), global contextual relationships (from the self-attention mechanism), and multi-scale information (from feature pyramids of different resolutions).
[0086] Based on this, depth images and point cloud data are used to provide geometric constraints and are fused with multi-scale features of color images for attention fusion to improve the prediction accuracy of 2D key point locations under conditions of texture blur, illumination variation, or partial occlusion.
[0087] Step 3: Perform global classification on the joint features based on the global semantic branch network to obtain the vehicle classification probability vector, and determine the vehicle identifier and classification confidence based on the vehicle classification probability vector.
[0088] Global classification is performed on the joint features to determine which of several pre-defined vehicle models the currently observed vehicle belongs to. Global classification means that the network outputs global attributes of the entire image (or the entire point cloud scene), rather than information from local regions.
[0089] The global semantic branch network can adopt the following structure: one or more global average pooling layers to compress the spatial dimension to 1×1 to obtain a global feature vector; one or more fully connected layers to map the global feature vector to the space of the number of categories N to obtain the original score; the last layer usually uses the Softmax or Sigmoid function to transform the original score and output a probability vector.
[0090] It should be noted that the vehicle classification probability vector is an N-dimensional vector, where N is the total number of vehicle categories pre-supported by the system. The vehicle identifier is a discrete label derived from the classification probability vector, and the category with the highest probability is taken as the recognition result. The classification confidence score quantifies the reliability of the current classification result; it is not directly equal to the maximum value in the vehicle classification probability vector, but is calculated based on the uncertainty of the entire probability distribution.
[0091] For example, information entropy can be used to quantify the uncertainty of classification. The formula for calculating classification confidence is as follows:
[0092] in, For classification confidence, Let N be the predicted probability of the i-th type of vehicle, and N be the total number of vehicle types.
[0093] Step 4: Perform key point identification on the joint features based on the local discriminant branch network to obtain heat maps corresponding to multiple semantic key points, determine the position of the corresponding appearance component based on each heat map, and calculate the location confidence based on all heat maps.
[0094] The local discriminative branch network predicts joint features and outputs a heatmap of multiple semantic key points.
[0095] It should be noted that a heatmap is a two-dimensional probability distribution map used to represent the probability that each pixel in the distribution map belongs to a certain key point. The location of the key point is obtained by finding the peak value.
[0096] Each heatmap corresponds to a specific vehicle exterior component (such as the left front headlight, vehicle logo, left front wheel hub, etc.), and the value of each pixel in the heatmap represents the probability or confidence level that the pixel is a key point of that component.
[0097] For each vehicle exterior component to be detected, only one key point with clear semantics is defined (usually the geometric center, corner point, or most prominent feature point of the component), rather than the component's contour point set or multiple key points. This design simplifies subsequent spatial transformation relationships. If multiple key points are defined for the same component, a transformation matrix needs to be stored separately for each key point, increasing complexity and being unnecessary.
[0098] The local discriminative branch neural network can adopt the following structure: Upsampling layers increase the spatial resolution of the feature map to the size of the target heatmap (e.g., the same size as the original image, or a fixed size such as 64×64). Upsampling enables subsequent convolutional layers to perform fine pixel-level predictions in a high-resolution space.
[0099] The convolutional layer, through learnable convolutional kernels, further spatially filters the upsampled features, eliminating checkerboard effects or discontinuities introduced by upsampling and enhancing local details. An embedded spatial attention mechanism focuses on easily identifiable areas such as headlights and grilles. This convolutional layer essentially implements an attention function, adaptively enhancing the response of areas containing key components while suppressing background or irrelevant regions. Specifically, the convolutional layer assigns different importance to different spatial locations using learned weights, causing the network to pay more attention to the pixel locations that contribute most to the localization task.
[0100] The last layer typically uses the Sigmoid activation function to compress the output value of each pixel to the [0,1] range, representing the probability that the pixel is a key point.
[0101] It should be noted that the number of channels in the feature map output by the convolutional layer is equal to the number of semantic key points K, thus ultimately outputting K heatmaps.
[0102] The peak position of the heatmap represents the location of the exterior component in the heatmap coordinate system. For this location, the pixel coordinates in the heatmap coordinate system need to be mapped back to the resolution of the original input image proportionally. These mapped coordinates are the location of the vehicle exterior component. If the heatmap peak appears in multiple locations (multiple pixel values are equal), the unique coordinates can be determined by averaging or using the center point.
[0103] Location reliability is used to quantify the reliability of the local network localization result. Location reliability is directly determined by the peak response intensity of multiple heatmaps. The reliability of localization is evaluated by extracting the peak response intensity of the heatmaps. If there is occlusion or blurred features, the peak value will decrease significantly.
[0104] Specifically, the formula for calculating location reliability is as follows:
[0105] in, For location-based reliability, K is the number of keypoints. It is the peak value of the heatmap of the kth key point.
[0106] It can be seen that when the peak values of the heatmaps of all components are very high (close to 1), the location confidence is close to 1, indicating that the overall location is very reliable.
[0107] When the peak value of the heatmap of one or more components is low (e.g., the component is occluded or too far away), the positioning reliability will decrease accordingly, indicating that the overall positioning is unreliable, and the embodied robot should be triggered to actively adjust the observation pose.
[0108] Step 5: Weight and fuse the classification confidence score and the location confidence score to obtain the observation confidence score.
[0109] Simultaneously, global and local features are concatenated and output through a fully connected layer to obtain the final vehicle model identifier and the location of vehicle exterior components (keypoint pixel coordinates).
[0110] The formula for calculating the weighted fusion is as follows:
[0111] in, The weighting coefficients for classification confidence. The weighting coefficients for location reliability.
[0112] An example of a fine-grained visual perception network architecture diagram with uncertainty estimation, such as... Figure 3As shown.
[0113] exist Figure 3 The network structure includes a multimodal feature backbone, point cloud feature extraction, multi-scale attention fusion module, global classification head, key point regression head, and weighted fusion module.
[0114] The multimodal feature backbone includes convolutional layer 1, pooling layer 1, convolutional layer 2, and pooling layer 2. Point cloud feature extraction includes an MLP layer, a max pooling layer, and another MLP layer. The multi-scale attention fusion module includes feature concatenation, a self-attention mechanism, and convolutional layers. The global classification head includes global average pooling, a fully connected layer, and a softmax layer, while the keypoint regression head includes an upsampling layer, a convolutional layer, and a sigmoid layer.
[0115] In one example, the global semantic branch network and the local discriminative branch network constitute a multi-task deep neural network. To enable this network to simultaneously and accurately identify vehicle models, precisely locate vehicle exterior components, and perform reliability assessments on its predictions, this application employs a composite loss function for end-to-end training. This composite loss function consists of a weighted sum of three types of sub-loss functions, corresponding to the three core capabilities the network needs to learn.
[0116] Specifically, the classification loss is applied to the output of the global semantic branch network (i.e., the vehicle classification probability vector). The design goal is to increase the inter-class distance between different vehicle categories in the feature space, making the feature representations of different vehicles as far apart as possible, thereby improving the accuracy of fine-grained vehicle identification. The classification loss encourages the network to map samples of the same vehicle type to nearby regions in the feature space, while mapping samples of different vehicles to regions that are far apart from each other. This loss function, which increases the distance, can effectively handle situations where vehicle types are highly similar in appearance, improving the robustness of classification.
[0117] For example, the classification loss uses additive angular margin loss to optimize the inter-class distance in the feature space, thereby improving vehicle model recognition accuracy. The expression is as follows:
[0118] in, For classifying losses, denoted as the angle between the feature vector and the center of the true vehicle category, m is the angle interval penalty term, s is the scaling factor, and N is the total number of vehicle categories.
[0119] Keypoint regression loss is applied to the output of the local discriminative branch network (i.e., the heatmap of semantic keypoints). The design goal is to calculate the pixel-wise deviation between the predicted keypoint heatmap and the ground truth labels. In the training data, each ground truth keypoint location (e.g., the center of the left front headlight) is converted into a Gaussian heatmap—a probability distribution radiating outwards from the ground truth coordinates with decaying values. The keypoint regression loss calculates the pixel-wise difference (e.g., mean squared error or cross-entropy loss) between the network's predicted heatmap and this Gaussian heatmap. By minimizing this loss, the network learns to generate high response peaks at the correct spatial locations, thus outputting accurate keypoint locations. Compared to directly regressing coordinates, the heatmap approach preserves spatial uncertainty information, making subsequent location confidence calculations (based on peak intensity) possible.
[0120] For example, the expression for keypoint regression loss is as follows:
[0121] in, For keypoint regression loss, The predicted coordinates of the k-th key point are... Here are the true coordinates, and K is the number of keypoints.
[0122] It should be noted that when and A perception is considered successful when the distance is less than a preset distance threshold and the vehicle type is correctly classified. Correct vehicle type classification means that the vehicle type identifier output by the network matches the actual vehicle type label. When training a deep neural network, perception success can be defined as a criterion for evaluating sample quality.
[0123] The uncertainty estimation loss is applied to the variance of the predicted distribution of the outputs of the global semantic branch network and the local discriminative branch network. The design goal is to constrain the variance of the predicted distribution of the outputs of these two networks, so that the network has the ability to quantify the uncertainty of observations. It serves as an additional loss term during the training phase, acting as a regularization term.
[0124] That is, the aim is to minimize high-confidence errors, endow embodied robots with the ability to quantify perception uncertainty and self-evaluate observation confidence, and enable them to judge the validity of the current perspective based on the prediction variance.
[0125] For example, uncertainty estimation loss introduces KL divergence as a supervision term, teaching the network that information entropy and heatmap peaks can reflect the true uncertainty.
[0126] Based on this, in order to accurately distinguish vehicle models and evaluate perceived quality, the expression for the verification loss function can be as follows:
[0127] in, , Hyperparameters used to measure the weight of each task.
[0128] In summary, a composite loss function is used to evaluate perceived reliability to avoid blind decision-making.
[0129] In one example, when the observation confidence level is lower than a preset condition, the embodied robot needs to actively adjust its observation pose to obtain clearer and more reliable perception information.
[0130] Based on this, controlling the embodied robot to adjust its observation pose includes the following steps: Step 1: Simulate multiple candidate observation poses in a local grid map.
[0131] In the local grid map maintained by the embodied robot, multiple candidate observation poses are generated using grid sampling or random sampling within the allowed range of motion, centered on the current observation position (e.g., 2 meters in front, 1 meter to the left and right, and a heading angle of ±30°). Each candidate observation pose includes the embodied robot's position coordinates and orientation angle.
[0132] Step 2: Predict the observation confidence for each candidate observation pose, and calculate the increase between the predicted observation confidence and the observed confidence.
[0133] For each candidate observation pose, estimate the classification confidence and location confidence that can be obtained under that pose, and thus obtain the predicted observation confidence.
[0134] It should be noted that for each candidate observation pose, it is necessary to predict how much the observation confidence can be improved if the embodied robot moves to that pose and collects environmental perception information.
[0135] The confidence level is increased by calculating the difference between the predicted observation confidence level and the current observation confidence level.
[0136] Specifically, predicting the observation confidence level for each candidate observation pose includes the following process: Step 21: Based on the pose transformation relationship between the current observation pose and each candidate observation pose, as well as the identified vehicle exterior component information, predict the observability of each exterior component under the candidate observation pose.
[0137] Among them, observability is used to characterize the observation effect of each appearance component under the candidate observation pose.
[0138] In this example, for each candidate observation pose, the observation confidence is not calculated after the robot actually moves to the candidate observation pose and acquires an image. Instead, the observation effect under the candidate observation pose is estimated in advance based on the current observation results.
[0139] Specifically, based on the pose transformation relationship between the current observation pose and the candidate observation pose, and combined with the identified three-dimensional positions of the vehicle exterior components and the camera imaging parameters, the imaging position of each exterior component under the candidate observation pose is predicted.
[0140] In this process, the three-dimensional coordinates of the currently observed external components are transformed from the camera coordinate system under the current pose to the robot coordinate system under the current pose (i.e., the current position and orientation of the embodied robot are used as a reference).
[0141] Then, using the relative pose transformation relationship between the current observation pose and the candidate observation pose (i.e., the rotation matrix and translation vector from the current pose to the candidate observation pose), the coordinates of the aforementioned external components are transformed from the robot coordinate system under the current pose to the robot coordinate system under the candidate observation pose. It should be noted that this transformation refers to the description of the coordinates under different robot pose reference frames, not a change in the robot coordinate system itself. The robot's body coordinate system remains fixed (e.g., with the chassis center as the origin and the front as the X-axis); only its position and orientation relative to the world coordinate system change.
[0142] Finally, using the fixed relative pose (hand-eye transformation matrix) between the camera and the robot body under the candidate observation pose, the coordinates of the appearance parts in the robot coordinate system under the candidate observation pose are projected onto the camera imaging plane under the candidate observation pose to obtain the predicted imaging position.
[0143] For each candidate observation pose, at least one of the following factors is evaluated: First, field of view coverage, which characterizes whether the exterior component is within the effective field of view of the camera; second, observation angle, which characterizes the degree of matching between the camera's observation direction and the surface orientation of the exterior component; and third, occlusion, which characterizes the degree to which the exterior component is occluded by the environment or vehicle structure in the candidate observation pose.
[0144] Specifically, based on the predicted imaging position (pixel coordinates) of the appearance component, it is determined whether it falls within the camera's effective image area. If the imaging position is within the image width and height range, it is considered to be within the field of view, and the field of view coverage factor is set to 1; otherwise, it is 0. For multiple components, the average of the field of view factors of all components can be taken.
[0145] The observation angle is defined as the angle between the vector pointing from the camera's optical center to the center of the component and the direction of the normal to the component's surface. The smaller the angle, the better the observation effect. This angle is calculated and converted into an observation angle score between 0 and 1 according to a preset mapping relationship (such as linear or Gaussian decay). The score is 1 when the angle is 0° and 0 when the angle reaches 90°.
[0146] Occlusion is determined using depth information or 3D point clouds: A virtual ray is emitted from the camera's optical center towards the center of the exterior component. If the ray intersects with other objects before reaching the component, occlusion is considered to exist. An occlusion factor between 0 and 1 is determined based on the occlusion area or the proportion of the ray that is blocked, with 1 for no occlusion and 0 for complete occlusion.
[0147] Step 22: Based on observability, predict the observation confidence level corresponding to each candidate observation pose.
[0148] Based on the aforementioned field of view coverage, observation angle, and occlusion conditions, the positional confidence of candidate observation poses can be predicted; and combined with the stability of the current vehicle model recognition results (e.g., the information entropy of the vehicle model classification probability vector), the observation confidence of candidate observation poses can be comprehensively evaluated, thereby obtaining the observation confidence corresponding to each candidate observation pose.
[0149] Specifically, the location reliability is obtained by weighting and summing the field coverage score, observation angle score, and occlusion score.
[0150] It should be noted that the preset weighting coefficients can be calibrated according to the actual scenario, such as setting them all to 1 / 3. Furthermore, all scores are normalized to the [0,1] interval. A higher location reliability indicates that more accurate and reliable positioning results can be obtained for the candidate observation pose.
[0151] Stability can be quantified by calculating the information entropy of the vehicle classification probability vector. A smaller entropy value indicates a more certain vehicle identification result (high stability); a larger entropy value indicates a more ambiguous identification result (low stability). The entropy value can be converted into a stability score. For example, the stability score expression is as follows:
[0152] in, For stability scoring, H is the entropy value, and N represents the total number of vehicle types. This makes... ∈[0,1], and the higher the value, the more stable the recognition.
[0153] Finally, the predicted location confidence score and the vehicle model recognition stability score are weighted and fused to obtain the observation confidence score for the candidate observation pose. It should be noted that the preset weights can each be set to 0.5.
[0154] Step 3: Select the candidate observation pose with the largest confidence increase as the next best observation point.
[0155] Using the principle of maximizing information gain, select the pose that maximizes the confidence gain.
[0156] In one example, the topology model is a dynamic semantic knowledge graph. A knowledge graph is a structured knowledge representation method that stores the spatial topological relationships between vehicle exterior components and charging ports for different vehicle models in the form of a graph. Its core components include a set of nodes and a set of edges.
[0157] The node set includes three types of nodes, each representing a type of entity: Root node: Represents the vehicle model identifier. Each vehicle model identifier corresponds to a unique root node. The root node is the entry point for all spatial topology information of that vehicle model in the entire graph.
[0158] Observation nodes: These represent exterior components, such as the left front headlight, right front headlight, vehicle logo, left front wheel hub, and left rearview mirror. These components have stable geometric features and are easily observed by the visual system from a distance or from a side view. Each exterior component corresponds to one observation node.
[0159] Inference node: Represents the charging port, typically including the center point of the charging port and its normal vector. The inference node is the final target that needs to be located.
[0160] Each edge in the edge set connects an observation node to an inference node (or the root node to an observation node) and is accompanied by a rigid body transformation matrix. A rigid body transformation refers to rotating and translating an object in three-dimensional space without changing its shape, size, or the distances between its internal points. This transformation preserves the object's rigidity. To mathematically describe the rotation and translation, a 4×4 homogeneous coordinate matrix is typically used for the rigid body transformation matrix. This matrix can simultaneously represent rotation (the 3×3 submatrix in the upper left corner) and translation (the 3×1 subvector in the upper right corner).
[0161] This matrix describes the rotation and translation transformation relationship from the coordinates of the observation node (exterior component) to the inference node (charging port). For example, For observation nodes, Let E be the inference node and E be the set of edges. For the i-th type of vehicle, its spatial topological relation set is defined as follows:
[0162] in, Let be the rigid body transformation matrix of the j-th observation node relative to the center of the charging port. It should be noted that the rigid body transformation matrix is pre-stored separately for each vehicle model. The positional relationship of the same exterior component (such as the left front headlight) relative to the charging port may be different for different vehicle models, so a separate set of transformation matrices needs to be stored for each vehicle model.
[0163] When adapting to a new car model, simply add a root node (the new car model identifier) to the knowledge graph and add edges (i.e., rigid body transformation matrices) between each observation node and inference node under that car model. There's no need to retrain the visual perception network. This embodies the decoupled design philosophy. That is, the decoupled graph architecture only needs to update parameters to adapt to new car models, without retraining the network.
[0164] Furthermore, based on the spatial transformation relationship and the positions of the vehicle exterior components, the location of the charging port is deduced, including the following steps: Step 1: Using the camera intrinsic parameters, back-project the pixel coordinates of the appearance component in the image to three-dimensional space to obtain the coordinates of the appearance component in the camera coordinate system.
[0165] To obtain the coordinates of this component in real 3D space, back projection is required using camera intrinsics and the depth value corresponding to that pixel. Camera intrinsics include focal length and optical center coordinates. The depth value can be obtained from depth images or point cloud data, representing the distance from the spatial point corresponding to that pixel to the camera plane.
[0166] For example, the back projection expression is as follows:
[0167] Among them, the two-dimensional pixel coordinates of the appearance component are (u,v), and the focal length coordinates are (u,v). The coordinates of the optical center are ( ). Z is the depth value. The three-dimensional coordinates of the appearance component in the camera coordinate system are ( ).
[0168] Step 2: Using a pre-calibrated hand-eye transformation matrix, transform the coordinates of the appearance component in the camera coordinate system to obtain the coordinates of the appearance component in the embodied robot coordinate system.
[0169] Body-bound robots typically define their own coordinate system (e.g., with the center of the chassis as the origin, the front as the X-axis, the left as the Y-axis, and the top as the Z-axis).
[0170] The camera is mounted in a fixed position on the embodied robot, and the relative pose between the two is obtained through hand-eye calibration.
[0171] Based on this, the coordinates of the external parts detected by the camera are multiplied by the hand-eye transformation matrix to obtain the coordinates of the external parts in the body robot coordinate system.
[0172] The expression is as follows:
[0173] in, Let these be the coordinates of the external components in the body robot's coordinate system. These are the coordinates of the exterior components in the camera coordinate system. This is the hand-eye transformation matrix.
[0174] Step 3: Based on the rigid body transformation matrix from the appearance component to the charging port stored in the topology model, and the coordinates of the appearance component in the embodied robot coordinate system, deduce the position of the charging port in the embodied robot coordinate system.
[0175] Based on this, the visually observed exterior components are mapped to a semantic graph, and the estimated position of the charging port in the embodied robot coordinate system is calculated according to the chain rule, as shown in the following expression:
[0176] in, Let be the pose transformation matrix of the charging port reference coordinate system relative to the robot's coordinate system. Let be the pose transformation matrix of the vehicle reference coordinate system relative to the embodied robot coordinate system. This is the rigid body transformation matrix of the charging port reference coordinate system relative to the vehicle reference coordinate system.
[0177] Furthermore, the position of the charging port in the embodied robot coordinate system is converted into a navigation target point in the world coordinate system.
[0178] The navigation guidance action vector can be set as follows:
[0179] in, This refers to the three-dimensional coordinate position of the charging port in the world coordinate system. This indicates the target orientation angle that the chassis of the embodied robot needs to be adjusted. The target point represents the behavior mode. When in exploration mode, the target point is the next best observation point; when in operation mode, the target point is the location of the charging port determined by semantic inference.
[0180] More intuitively, Figure 4 This is a framework diagram of a positioning system for charging a vehicle using an embodied robot, provided as an embodiment of this application.
[0181] exist Figure 4 In this context, the working process of the embodied robot includes the following: 1) Initial perception and evaluation: The embodied robot (automatic charging robot) collects RGB images, depth maps and 3D point cloud data, and inputs them into a fine-grained visual perception network to obtain vehicle model ID, vehicle exterior component locations (obtained through key point heatmaps) and observation confidence.
[0182] 2) Embodied Active Exploration: If the observation confidence is less than the confidence threshold (e.g., due to excessive distance or angle causing key points to be blurred), active viewpoint planning is executed. Based on the principle of maximizing information gain, the system calculates the next optimal observation point in the local map, drives the chassis to actively move to the optimal observation point for secondary observation, until the observation confidence is greater than the confidence threshold.
[0183] 3) Semantic Graph Reasoning: Mapping key points of visually observed exterior components to a semantic graph. Calculating the estimated position of the charging port in the embodied robot coordinate system based on the chain rule.
[0184] 4) Convert the position of the charging port in the robot's coordinate system into a navigation target point in the world coordinate system, send it to the chassis motion controller, and drive the robot to move to the charging work area.
[0185] Based on the same idea, some embodiments of this application also provide devices and non-volatile computer storage media corresponding to the above methods.
[0186] Figure 5 A schematic diagram of a positioning device for vehicle charging by an embodied robot provided in this application embodiment includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform a positioning method for charging a vehicle by an embodied robot as described above.
[0187] Some embodiments of this application provide a non-volatile computer storage medium for locating a vehicle charging device using an embodied robot, which stores computer-executable instructions capable of executing any of the above-described methods for locating a vehicle charging device using an embodied robot.
[0188] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device and medium embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the description of the method embodiments.
[0189] The devices and media provided in this application are one-to-one with the methods. Therefore, the devices and media also have similar beneficial technical effects as their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.
[0190] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0191] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the technical principles of this application should fall within the protection scope of this application.
Claims
1. A method for locating a vehicle charging station using an embodied robot, characterized in that, The method includes: Acquire environmental perception information of the embodied robot in its current observation pose; The environmental perception information is identified to obtain the vehicle model identifier, the location of vehicle exterior components, and the observation confidence level corresponding to the identification result; If the observation confidence level is lower than the preset condition, the robot is controlled to adjust its observation posture, reacquire environmental perception information and perform identification again until the observation confidence level meets the preset condition. If the observation confidence level meets the preset conditions, the spatial transformation relationship associated with the vehicle model identifier is retrieved from the pre-constructed topological model; Based on the spatial transformation relationship and the position of the vehicle exterior components, the location of the charging port can be deduced.
2. The method according to claim 1, characterized in that, The identification of the environmental perception information specifically includes: Multimodal features of color and depth images are extracted, and point cloud features of point cloud data are also extracted; depth images are obtained by aligning with color images, and point cloud data are obtained by transforming depth images. The point cloud features and the image multimodal features are fused using multi-scale attention to obtain joint features; The joint features are classified globally using a global semantic branch network to obtain a vehicle classification probability vector, and the vehicle identifier and classification confidence are determined based on the vehicle classification probability vector. The joint features are identified by a local discriminant branch network to obtain heat maps corresponding to multiple semantic key points. The location of each corresponding appearance component is determined based on each heat map, and the location confidence is calculated based on all heat maps. The classification confidence score and the location confidence score are weighted and fused to obtain the observation confidence score.
3. The method according to claim 2, characterized in that, The classification confidence level is determined based on the vehicle classification probability vector, specifically including: The classification confidence level is obtained by calculating the information entropy of the vehicle classification probability vector; Location reliability is calculated based on all heatmaps, specifically including: The location reliability is obtained by calculating the average of the peak responses of each heatmap.
4. The method according to claim 2, characterized in that, The multimodal feature extraction module includes a first convolutional layer, a first pooling layer, a second convolutional layer, and a second pooling layer; The point cloud feature extraction module includes a first multi-layer perceptron layer, a max pooling layer, and a second multi-layer perceptron layer; The multi-scale attention fusion module includes a feature concatenation layer, a self-attention mechanism layer, and a convolutional layer; The global semantic branch network includes a global average pooling layer, a fully connected layer, and a softmax layer; The local discriminative branch network includes upsampling layers, convolutional layers, and sigmoid layers.
5. The method according to claim 2, characterized in that, The global semantic branch network and the local discriminative branch network are trained using the following composite loss function: Classification loss is used to increase the inter-class distance between different vehicle types in the feature space; Keypoint regression loss is used to calculate the pixel deviation between the predicted keypoint heatmap and the true label. Uncertainty estimation loss is used to constrain the variance of the predicted distribution of the outputs of the global semantic branch network and the local discriminative branch network, so that the network has the ability to quantify the uncertainty of observation.
6. The method according to claim 1, characterized in that, Controlling the embodied robot to adjust its observation pose specifically includes: Simulate multiple candidate observation poses in a local grid image; Predict the observation confidence for each candidate observation pose, and calculate the increase between the predicted observation confidence and the actual observation confidence. The candidate observation pose with the largest increase in confidence level is selected as the next best observation point.
7. The method according to claim 6, characterized in that, Predicting the observation confidence for each candidate observation pose, specifically including: Based on the pose transformation relationship between the current observation pose and each candidate observation pose, as well as the identified vehicle exterior component information, the observability of each exterior component under the candidate observation pose is predicted; observability is used to characterize the observation effect of each exterior component under the candidate observation pose. Based on observability, predict the observation confidence level corresponding to each candidate observation pose.
8. The method according to claim 1, characterized in that, The topological model is a dynamic semantic knowledge graph, which includes: The node set includes a root node representing the vehicle model identifier, observation nodes representing exterior components, and inference nodes representing the charging port. The edge set includes rigid transformation matrices representing the relationships between nodes; for each vehicle category, the rigid transformation matrix of each exterior component relative to the charging port is pre-stored.
9. The method according to claim 8, characterized in that, Based on the spatial transformation relationship and the positions of the vehicle exterior components, the location of the charging port is inferred, specifically including: Using the camera intrinsic parameters, the pixel coordinates of the appearance component in the image are back-projected into three-dimensional space to obtain the coordinates of the appearance component in the camera coordinate system; Using a pre-calibrated hand-eye transformation matrix, the coordinates of the appearance component in the camera coordinate system are transformed to obtain the coordinates of the appearance component in the embodied robot coordinate system; Based on the rigid body transformation matrix from the appearance component to the charging port stored in the topological model, and the coordinates of the appearance component in the embodied robot coordinate system, the position of the charging port in the embodied robot coordinate system is inferred.
10. A positioning device for a robot charging a vehicle, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which enables the at least one processor to perform a positioning method for charging a vehicle by an embodied robot as described in any one of claims 1-9.