Roadside laser radar camera system adaptive calibration method based on reinforcement learning

Through a reinforcement learning-based method training agent to adjust the calibration matrix of the lidar and camera, the coordinate conversion relationship problem caused by sensor position changes is solved, adaptive calibration is achieved, and the accuracy and reliability of data fusion are improved.

CN120275939AInactive Publication Date: 2025-07-08VANJEE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311868647.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Changes in sensor position affect the coordinate conversion relationship between lidar and camera, resulting in reduced data fusion accuracy and reliability. Traditional calibration methods cannot adapt to changes in sensor position.

Method used

Using a reinforcement learning-based method, the calibration matrix between the lidar and the camera is adjusted through agent training, and the calibration process is optimized by the benefit function to achieve adaptive calibration.

Benefits of technology

In the case of sensor position change, maintain the sensor's accurate calibration parameters to improve the quality and accuracy of data fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120275939A_ABST
    Figure CN120275939A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of laser radar and camera calibration, and provides a road side laser radar camera system adaptive calibration method and device based on reinforcement learning, and the method comprises the steps: obtaining the scanning data of a laser radar and the image data of a camera within a preset time period after the pose of the laser radar and the camera is changed; the learning behavior of the intelligent agent is executed, the learning behavior of the intelligent agent is overlapped in the direction of increasing the benefit function until the benefit value of the benefit function is maximum, the learning behavior is to adjust the calibration matrix between the laser radar and the camera, and the benefit function is used for representing the benefit of coordinate conversion between the laser radar and the camera. According to the calibration method provided by the invention, the calibration process of the roadside laser radar and the camera is analogous to an agent behavior and income process in reinforcement learning, and even if the position is changed due to external factors, the sensor can still adaptively adjust the calibration parameters thereof, so that the quality and accuracy of data fusion are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of lidar and camera calibration, and particularly relates to an adaptive calibration method and device for a roadside lidar-camera system based on reinforcement learning. Background Art

[0002] In the fields of autonomous driving and intelligent transportation, lidar and cameras are important sensors that can provide accurate environmental perception information. However, the poses of these sensors may change due to various reasons. For example, roadside sensor units are usually installed on roadside single poles or gantries, etc., and are easily affected by external factors, resulting in changes in position. In addition, they may also be affected by natural environmental factors. The change in the sensor pose will affect the coordinate transformation relationship between the lidar and the camera, and the change in the calibration parameters will affect the data fusion, reducing the accuracy and reliability of the lidar-camera system. Summary of the Invention

[0003] An embodiment of this application provides an adaptive calibration method and device for a roadside lidar-camera system based on reinforcement learning, which can solve the problem that the sensor cannot automatically adjust its calibration parameters when the pose and position of the sensor change.

[0004] In a first aspect, an embodiment of this application provides an adaptive calibration method for a roadside lidar-camera system based on reinforcement learning, including:

[0005] After the poses of the lidar and the camera change, obtain the scan data of the lidar and the image data of the camera within a preset time period;

[0006] Execute the learning behavior of the agent, and iterate the learning behavior of the agent in the direction of making the benefit function increase until the benefit value of the benefit function is the largest. The learning behavior is to adjust the calibration matrix between the lidar and the camera, and the benefit function is used to represent the benefit of the coordinate transformation between the lidar and the camera.

[0007] In a possible implementation manner, after obtaining the scan data of the lidar and the image data of the camera within a preset time period, the method further includes:

[0008] Parse the target-level feature data of all targets within the preset time period from the scan data and the image data;

[0009] Determine the loss value of the loss function according to the target-level feature data corresponding to the lidar and the camera respectively under the current calibration matrix, and determine the benefit value of the benefit function according to the loss value of the loss function.

[0010] In a possible implementation, parsing the target-level feature data of all targets within the preset time period from the scan data and the image data includes:

[0011] Using the current calibration matrix, converting the scan data of the lidar into the camera coordinates of the camera to obtain the mapped data of the scan data in the camera coordinates;

[0012] Performing target-level feature parsing on the mapped data and the image data of the camera to obtain the first target-level feature data corresponding to the image data and the second target-level feature data corresponding to the mapped data;

[0013] Among them, determining the loss value of the loss function according to the target-level feature data corresponding to the lidar and the camera respectively under the current calibration matrix includes: determining the loss value of the loss function according to the first target-level feature data and the second target-level feature data.

[0014] In a possible implementation, using the current calibration matrix to perform target-level feature parsing on the mapped data and the image data of the camera to obtain the first target-level feature data corresponding to the image data and the second target-level feature data corresponding to the mapped data includes:

[0015] Determining the first target-level feature data and the second target-level feature data from the image data of the camera and the mapped data through a target recognition algorithm.

[0016] In a possible implementation, determining the loss value of the loss function according to the first target-level feature data and the second target-level feature data includes:

[0017] For each target at each moment, calculating the difference between the feature dimensions in the first target-level feature data and the second target-level feature data;

[0018] According to the differences between the feature dimensions, combining with the loss function, determining the loss value of the loss function.

[0019] In a possible implementation, the method further includes:

[0020] Based on the static features of the fixed markers, extracting the static features of the fixed markers before and after the pose change of the lidar and the camera; the static features include: plane features and scene identifier features;

[0021] Determining the loss value of the loss function according to the differences between the feature dimensions, combining with the loss function, includes:

[0022] Generate the first loss value of the first loss function according to the differences of each feature dimension and in combination with the first loss function;

[0023] Generate the second loss value of the second loss function according to the static features of the fixed marker before and after the pose change of the lidar and the camera and in combination with the second loss function;

[0024] Calculate the first product of the first loss value and the first preset loss weight, calculate the second product of the second loss value and the second preset loss weight, and take the sum of the first product and the second product as the loss value of the loss function.

[0025] In a possible implementation, the first loss function is:

[0026] where t represents the time, T is the preset time period, n represents different targets, N is the number of targets, k1, k2, k3, and k4 are corresponding preset weights, d is the distance, v is the speed, θ is the heading angle, and c is the category.

[0027] In a possible implementation, the second loss function is:

[0028] where k1 and k2 are preset weights, t represents the time, T is the preset time period, n represents different targets, N is the number of targets, h is the road surface normal vector, and s is the slope of the road surface marking line.

[0029] Second, an adaptive calibration device for a roadside lidar-camera system based on reinforcement learning provided by an embodiment of the present application includes:

[0030] An acquisition module, after the pose of the lidar and the camera changes, acquires the scan data of the lidar and the image data of the camera within a preset time period;

[0031] A reinforcement learning module, executes the learning behavior of the agent, and iterates the learning behavior of the agent along the direction that makes the benefit function smaller until the benefit value of the benefit function is the largest. The learning behavior is to adjust the calibration matrix between the lidar and the camera, and the benefit function is used to represent the benefit of the coordinate transformation between the lidar and the camera.

[0032] Third, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the method described above is implemented.

[0033] Advantages of the present application

[0034] This application provides an adaptive calibration method, device, and storage medium for a roadside lidar-camera system based on reinforcement learning, mainly including the following steps: First, obtain the scan data of the lidar and the image data of the camera within a preset time period; then, use the reinforcement learning algorithm to train the agent so that it can automatically adjust the calibration matrix between the lidar and the camera. During the training process, the agent will adjust its behavior according to the value of the benefit function to maximize the benefit value of the benefit function. The benefit function is used to represent the benefit of the coordinate transformation between the lidar and the camera, and the larger its value, the higher the accuracy of the coordinate transformation. In this way, this solution can achieve the adaptive calibration of the lidar and the camera under changing environmental conditions.

[0035] The calibration method provided by this application analogizes the calibration process of the roadside lidar and the camera to the process of the agent's behavior and reward in reinforcement learning, enabling the roadside sensor to adaptively adjust its calibration parameters to adapt to changes in the external environment. In this way, even when the position changes due to external factors, the sensor can still maintain accurate calibration parameters, thereby improving the quality and accuracy of data fusion. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0037] Figure 1 is a schematic flowchart of an adaptive calibration method for a lidar and a camera based on reinforcement learning provided by an embodiment of this application;

[0038] Figure 2 is a schematic flowchart of determining the benefit value of the benefit function according to the scan data of the lidar and the image data of the camera within a preset time period provided by an embodiment of this application;

[0039] Figure 3 is a schematic flowchart of a method for extracting target-level feature data provided by an embodiment of this application;

[0040] Figure 4 is a schematic flowchart of determining the loss value of the loss function by combining static features provided by an embodiment of this application;

[0041] Figure 5 is a schematic structural diagram of a device for adaptive calibration of a roadside lidar-camera system based on reinforcement learning provided by an embodiment of this application;

[0042] Figure 6It is a schematic structural diagram of a terminal device provided by an embodiment of the present application. Detailed implementation manners

[0043] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0044] It should be understood that when used in the specification and appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0045] It should also be understood that the term "and / or" used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0046] As used in the specification and appended claims of the present application, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" can be interpreted as meaning "once determined", "in response to determining", "once detecting [the described condition or event]", or "in response to detecting [the described condition or event]" depending on the context.

[0047] In addition, in the description of the specification and appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0048] The reference to "one embodiment" or "some embodiments" etc. described in the specification of the present application means that specific features, structures, or characteristics described in connection with that embodiment are included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0049] In the fields of autonomous driving and intelligent transportation, lidar and cameras are important sensors that can provide accurate environmental perception information. However, the poses of these sensors may change for various reasons. For example, roadside sensor units are usually installed on roadside single poles or gantries, etc., and are vulnerable to external factors, resulting in changes in position. In addition, they may also be affected by natural environmental factors. The change in the sensor pose will affect the coordinate transformation relationship between the lidar and the camera, and the change in the calibration parameters will affect the data fusion, reducing the accuracy and reliability of the lidar and camera system.

[0050] To solve this problem, it is usually necessary to calibrate the lidar and the camera to determine the coordinate transformation relationship between them. Traditional calibration methods are usually carried out in a specific scenario, requiring specific calibration boards and calibration procedures, which cannot achieve online calibration of the sensors, and also require a large amount of manpower and time.

[0051] In actual situations, since the poses of the lidar and the camera may change at any time, traditional calibration methods are not applicable to real-world scenarios where the poses change. Therefore, an adaptive calibration method is needed to dynamically adjust the calibration relationship between the lidar and the camera.

[0052] In this context, this application provides an adaptive calibration method for lidar and camera based on reinforcement learning, which mainly includes the following steps: First, obtain the scan data of the lidar and the image data of the camera within a preset time period; then, use the reinforcement learning algorithm to train the intelligent agent so that it can automatically adjust the calibration matrix between the lidar and the camera. During the training process, the intelligent agent will adjust its behavior according to the value of the benefit function to maximize the benefit value of the benefit function. The benefit function is used to represent the benefit of the coordinate transformation between the lidar and the camera, and the larger its value, the higher the accuracy of the coordinate transformation. In this way, this technology can achieve the adaptive calibration of the lidar and the camera under changing environmental conditions.

[0053] To better understand the present invention, first, a brief introduction to the intelligent agent is given.

[0054] An intelligent agent is a software entity that has the abilities of perception, thinking, action, self-adaptation, and learning, and can actively interact with the environment. The intelligent agent can replace humans to complete some complex tasks, improving work efficiency and service quality. At the same time, the intelligent agent can also help humans solve some difficult problems, such as large-scale data processing, complex problem-solving, etc.

[0055] In reinforcement learning, an agent interacts with the environment and learns control strategies through trial and error. The process of the agent's behavior and rewards includes: the agent executes an action in the environment; after the action is executed, the environment transitions to a new state; for this new state, the environment gives a reward signal, which can be a positive reward or a negative reward; the agent executes a new action according to the new state and the reward feedback from the environment according to a certain strategy; through reinforcement learning, the agent can know what action to take in what state to obtain the maximum reward; the agent uses the reward signal to update and improve its strategy to obtain better results in the future environment; repeat the above process until the agent reaches a satisfactory strategy or reaches a preset learning cycle.

[0056] To illustrate the technical solution of this application, specific embodiments will be used to illustrate below.

[0057] Refer to Figure 1 The process of an embodiment of the lidar and camera adaptive calibration method based on reinforcement learning shown, by way of example and not limitation, includes the following steps:

[0058] Step S100: After the poses of the lidar and the camera change, obtain the scan data of the lidar and the image data of the camera within a preset time period;

[0059] Step S200: Execute the learning behavior of the agent, and iterate the learning behavior of the agent in the direction of making the benefit function increase until the benefit value of the benefit function is the largest. The learning behavior is to adjust the calibration matrix between the lidar and the camera, and the benefit function is used to represent the benefit of the coordinate transformation between the lidar and the camera.

[0060] In order to better understand the present invention, first, a brief introduction to the agent will be given.

[0061] An intelligent agent refers to a software entity that has the abilities of perception, thinking, action, self - adaptation and learning, and can actively interact with the environment. The intelligent agent can replace humans to complete some complex tasks, improve work efficiency and service quality. At the same time, the intelligent agent can also help humans solve some difficult problems, such as large - scale data processing, complex problem solving, etc.

[0062] In reinforcement learning, an agent interacts with an environment and learns a control policy through trial and error. The process of the agent's behavior and rewards includes: the agent performs an action in the environment; after the action is performed, the environment transitions to a new state; for this new state, the environment gives a reward signal, which can be a positive reward or a negative reward; the agent executes a new action according to the new state and the reward feedback from the environment according to a certain policy; through reinforcement learning, the agent can know what action to take in what state to obtain the maximum reward; the agent uses the reward signal to update and improve its policy to obtain better results in the future environment; repeat the above process until the agent reaches a satisfactory policy or reaches a preset learning period.

[0063] In an intelligent transportation system, roadside sensors (such as lidar and cameras) are usually deployed on single poles or gantries beside the road. Due to the influence of natural environmental factors (such as wind, rain, etc.), the poses of these sensors may change. This application provides a lidar and camera adaptive calibration method based on reinforcement learning. This method uses a reinforcement learning algorithm to train an agent to automatically adjust the calibration matrix between the lidar and the camera to adapt to the change of the sensor pose and improve the perception performance and decision-making accuracy of the intelligent transportation system.

[0064] In this application, the lidar and the camera are used as agents for deep reinforcement learning, and the calibration matrix is used as the behavior model of the agent to construct a calibration model of the lidar and the camera in deep reinforcement. To achieve the adaptive calibration between the lidar and the camera, it is necessary to define a reward model for the calibration process. The reward model mainly gives the benefit function of the calibration matrix transformation during the process of the agent optimizing the calibration matrix. Through reinforcement learning, the agent can know what action to take in what state to obtain the maximum reward; the agent uses the reward signal to update and improve its policy to obtain better results in the future environment.

[0065] Specifically, in step S100, the perception data of the lidar and the camera in different poses within a preset time period is obtained for subsequent agent learning and coordinate transformation; further, in step S200, a reinforcement learning algorithm is used to train the agent. The agent can learn and optimize by interacting with the environment. In this application, the agent optimizes the benefit of coordinate transformation by adjusting the calibration matrix between the lidar and the camera. During the training process, the agent will adjust its behavior according to the value of the benefit function to maximize the benefit value of the benefit function.

[0066] It should be noted that the benefit function is used to evaluate the accuracy and reliability of coordinate transformation. The larger its value, the better the effect of coordinate transformation. By iterating the learning behavior of the agent in the direction that makes the benefit function increase until the benefit value of the benefit function is maximized, the agent will learn the optimal calibration matrix parameters to adapt to the change of the sensor pose and improve the benefit of coordinate transformation.

[0067] The calibration method provided by this application analogizes the calibration process of the roadside lidar and the camera to the process of the agent's behavior and reward in reinforcement learning, enabling the roadside sensors to adaptively adjust their calibration parameters to adapt to the changes in the external environment. In this way, even when the position changes due to external factors, the sensors can still maintain accurate calibration parameters, thereby improving the quality and accuracy of data fusion.

[0068] See Figure 2 , in a possible implementation manner, after obtaining the scan data of the lidar and the image data of the camera within a preset time period, the method further includes:

[0069] Step A100: Parse the target-level feature data of all targets within the preset time period from the scan data and the image data;

[0070] Step A200: Determine the loss value of the loss function according to the target-level feature data corresponding to the lidar and the camera respectively under the current calibration matrix, and determine the benefit value of the benefit function according to the loss value of the loss function.

[0071] In step A100, by parsing the scan data of the lidar and the image data of the camera, the target-level feature data of all targets is extracted. Among them, the target-level feature data may include information such as the shape, size, and position of the target, reflecting the physical attributes and spatial relationships of the target, and providing basic data for subsequent coordinate transformation and benefit evaluation.

[0072] It should be noted that the benefit function is used to evaluate the benefit of coordinate transformation. The larger its value, the better the effect of coordinate transformation. The benefit function can combine the loss function and comprehensively consider factors such as the matching degree of different target-level feature data and the transformation error to evaluate the overall effect of coordinate transformation. Exemplarily, the relationship between the benefit function and the loss function may be inverse, that is, the smaller the loss value of the loss function, the larger the benefit value of the benefit function. In addition, other evaluation methods can also be used to evaluate the loss function to obtain the benefit function value. For example, the Actor-Critic method can be used to score the loss function to obtain the benefit function value.

[0073] See Figure 3 , in a possible implementation manner, step A100 includes:

[0074] Step A101: Use the current calibration matrix to convert the scan data of the lidar into the camera coordinates of the camera, and obtain the mapped data of the scan data in the camera coordinates;

[0075] Step A102: Perform target-level feature analysis on the mapped data and the image data of the camera to obtain first target-level feature data corresponding to the image data and second target-level feature data corresponding to the mapped data;

[0076] Correspondingly, step A200 includes: determining the loss value of the loss function according to the first target-level feature data and the second target-level feature data.

[0077] It should be noted that in order to obtain the target-level feature data of all targets within a preset time period, it is first necessary to convert the scan data of the lidar into the camera coordinates of the camera through step A101. This step uses the current calibration matrix, which defines the coordinate transformation relationship between the lidar and the camera. Through the current calibration matrix, the scan data of the lidar can be converted from the coordinate system of the lidar to the coordinate system of the camera to achieve data alignment and mapping.

[0078] Furthermore, in step A102, target-level feature analysis is performed on the mapped data and the image data of the camera to obtain first target-level feature data corresponding to the image data and second target-level feature data corresponding to the mapped data. Exemplarily, the target-level feature data may include information such as the shape, size, and position of the target, and the present application does not limit this.

[0079] In a possible implementation manner, step A102 includes:

[0080] Determine the first target-level feature data and the second target-level feature data from the image data of the camera and the mapped data through a target recognition algorithm.

[0081] It should be noted that the target recognition algorithm is a branch of deep learning algorithms, which is mainly used to recognize and extract the targets of interest from images. The target recognition algorithm usually uses a convolutional neural network (CNN) as the basic architecture, extracts features in the image through multiple layers of convolution and pooling operations, and uses a fully connected layer or a softmax layer for classification. The target recognition algorithm can recognize various objects in the image, such as faces, vehicles, pedestrians, etc. Commonly used target recognition algorithms include: (1) Faster R-CNN is a classic deep learning-based target detection algorithm, which is characterized by introducing a Region Proposal Network (RPN) and a ROIPooling layer in the network, improving the accuracy and efficiency of target detection; (2) YOLO (You Only Look Once) is a target detection algorithm based on a single forward pass, which is characterized by fast speed and high accuracy, and is suitable for real-time application scenarios; (3) SSD (Single Shot MultiBox Detector) is a multi-box detection-based target detection algorithm, which is characterized by being able to detect multiple targets in the image simultaneously and having good robustness. This application does not limit the target recognition algorithm.

[0082] In this embodiment, the first target-level feature data is extracted from the image data of the camera by using the target recognition algorithm, and the second target-level feature data is extracted from the mapping data. Both the first target-level data and the second target-level data may include: distance, speed, heading angle, and category. In some embodiments, the first target-level feature data and the second target-level feature data may both include other types of data, and this application does not limit this.

[0083] In a possible implementation manner, determining the loss value of the loss function according to the first target-level feature data and the second target-level feature data includes:

[0084] Step B1: For each target at each moment, calculate the difference between the feature dimensions in the first target-level feature data and the second target-level feature data;

[0085] Step B2: According to the differences between the feature dimensions, in combination with the loss function, determine the loss value of the loss function.

[0086] It should be noted that within a preset time period, the lidar and the camera will collect multiple target-level feature data of multiple targets at multiple moments. For all the collected data, it is necessary to calculate the difference between the first target-level feature data and the second target-level feature data. According to the differences between the feature dimensions and in combination with the loss function, the specific loss value can be calculated, and this loss value reflects the accuracy of the current calibration matrix. The agent will also further determine the benefit value of the benefit function according to this loss value and continuously improve its behavior according to the benefit value.

[0087] Refer to Figure 4 , in a possible implementation, the method further includes:

[0088] Step C1: Based on the static features of the fixed marker, extract the static features of the fixed marker before and after the pose change of the lidar and the camera; the static features include plane features and scene identifier features;

[0089] Determining the loss value of the loss function according to the difference of each feature dimension and in combination with the loss function includes:

[0090] Step C2: Generate a first loss value of the first loss function according to the difference of each feature dimension and in combination with the first loss function;

[0091] Step C3: Generate a second loss value of the second loss function according to the static features of the fixed marker before and after the pose change of the lidar and the camera and in combination with the second loss function;

[0092] Step C4: Calculate a first product of the first loss value and a first preset loss weight, calculate a second product of the second loss value and a second preset loss weight, and use the sum of the first product and the second product as the loss value of the loss function.

[0093] In this embodiment, first, according to the difference of each feature dimension in the first target-level feature data and the second target-level feature data, in combination with the first loss function, generate a first loss value of the first loss function; then, according to the static features of the fixed marker before and after the pose change of the lidar and the camera, in combination with the second loss function, generate a second loss value of the second loss function; finally, calculate the loss value of the final loss function in combination with the weights of the first loss value and the second loss value.

[0094] It should be noted that in the construction of the reinforcement learning input in this embodiment, the feature effects of different levels are considered, and the target-level feature data and the static feature data at the feature level are jointly used as the input to enhance the accuracy of the result. Among them, the static data at the feature level includes the plane features and scene identifier features of the fixed marker. Exemplarily, the plane feature can be the road surface normal vector, and the scene identifier feature can be the slope of the road surface marking line. In some embodiments, it may also include stationary feature objects in other lane scenes, and the present application does not limit this.

[0095] In a possible implementation, the first loss function is:

[0096] Where T is a preset time period, N is the target quantity, k1, k2, k3, and k4 are corresponding preset weights, d is the distance, v is the speed, θ is the heading angle, and c is the category.

[0097] In a possible implementation, the second loss function is:

[0098] Where k1 and k2 are preset weights, T is a preset time period, N is the target quantity, h is the road surface normal vector, and s is the slope of the road surface marking line.

[0099] An apparatus for adaptive calibration of a roadside lidar-camera system based on reinforcement learning provided by an embodiment of the present application is referred to Figure 5 , and includes:

[0100] An acquisition module 301, which acquires the lidar scan data and the camera image data within a preset time period after the poses of the lidar and the camera change;

[0101] A reinforcement learning module 302, which executes the learning behavior of the agent and iterates the learning behavior of the agent in the direction of making the benefit function smaller until the benefit value of the benefit function is the largest. The learning behavior is to adjust the calibration matrix between the lidar and the camera, and the benefit function is used to represent the benefit of the coordinate transformation between the lidar and the camera.

[0102] The present application provides a lidar and camera adaptive calibration device based on reinforcement learning. The device trains an agent through a reinforcement learning algorithm to automatically adjust the calibration matrix between the lidar and the camera to adapt to the change of the sensor pose, and improve the perception performance and decision-making accuracy of the intelligent transportation system.

[0103] In the present application, the lidar and the camera are used as the agents of deep reinforcement learning, and the calibration matrix is used as the behavior model of the agent to construct the calibration model of the lidar and the camera in deep reinforcement. To achieve the adaptive calibration between the lidar and the camera, a reward model for the calibration process needs to be defined. The reward model mainly gives the benefit function of the calibration matrix transformation during the process of the agent optimizing the calibration matrix. Through reinforcement learning, the agent can know what actions it should take in what state to obtain the maximum reward; the agent uses the reward signal to update and improve its strategy to obtain better results in the future environment.

[0104] Specifically, first, the acquisition module 301 acquires the perception data of the lidar and the camera at different poses within a preset time period for subsequent agent learning and coordinate transformation. Further, the reinforcement learning module 302 is used to train the agent. The agent can learn and optimize by interacting with the environment. In this application, the agent optimizes the benefit of coordinate transformation by adjusting the calibration matrix between the lidar and the camera. During the training process, the agent will adjust its behavior according to the value of the benefit function to maximize the benefit value of the benefit function.

[0105] It should be noted that the benefit function is used to evaluate the accuracy and reliability of coordinate transformation. The larger its value, the better the effect of coordinate transformation. By iteratively adjusting the learning behavior of the agent in the direction that makes the benefit function increase until the benefit value of the benefit function is maximized, the agent will learn the optimal calibration matrix parameters to adapt to the changes in the sensor pose and improve the benefit of coordinate transformation.

[0106] The calibration device provided in this application analogizes the calibration process of the roadside lidar and the camera to the process of the agent's behavior and reward in reinforcement learning, enabling the roadside sensors to adaptively adjust their calibration parameters to adapt to changes in the external environment. In this way, even when the position changes due to external factors, the sensors can still maintain accurate calibration parameters, thereby improving the quality and accuracy of data fusion.

[0107] Figure 6 FIG. 10 is a schematic structural diagram of a terminal device provided in an embodiment of the present application. The terminal device 400 includes: at least one processor 401 ( Figure 6 only one processor is shown), a memory 402, and a computer program 403 stored in the memory 402 and executable on the at least one processor 401. When the processor 401 executes the computer program 403, the steps in the calibration method embodiment described above are implemented.

[0108] The terminal device 400 may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor 401 and a memory 402. Those skilled in the art can understand that Figure 6 merely examples of the terminal device 400, which do not constitute a limitation on the terminal device 400, may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0109] The so-called processor 401 may be a Central Processing Unit (CPU), and this processor 401 may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0110] In some embodiments, the memory 402 may be an internal storage unit of the terminal device 400, such as the hard disk or memory of the terminal device 400. In some other embodiments, the memory 402 may also be an external storage device of the terminal device 400, such as a plug-in hard disk equipped on the terminal device 400, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 402 may also include both the internal storage unit of the terminal device 400 and the external storage device. The memory 402 is used to store an operating system, application programs, a BootLoader, data, and other programs, such as the program code of the computer program, etc. The memory 402 may also be used to temporarily store data that has been output or is to be output.

[0111] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be elaborated herein.

[0112] The embodiments of the present application also provide a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments can be implemented.

[0113] The embodiments of the present application provide a computer program product. When the computer program product runs on a mobile terminal, the mobile terminal can implement the steps in the above method embodiments when executed.

[0114] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above method embodiments of the present application can be completed by instructing related hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps in the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device capable of carrying the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium may not be an electrical carrier signal and a telecommunication signal.

[0115] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0116] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in this document can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.

[0117] In the embodiments provided in the present application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0118] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0119] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A method for adaptive calibration of a roadside lidar-camera system based on reinforcement learning, characterized in that Including: After the poses of the lidar and the camera change, obtaining the scan data of the lidar and the image data of the camera within a preset time period; Performing the learning behavior of the agent, and iterating the learning behavior of the agent in the direction that makes the benefit function increase until the benefit value of the benefit function is maximized. The learning behavior is to adjust the calibration matrix between the lidar and the camera, and the benefit function is used to represent the benefit of the coordinate transformation between the lidar and the camera.

2. The method according to claim 1, wherein After obtaining the scan data of the lidar and the image data of the camera within a preset time period, the method further includes: Parsing the target-level feature data of all targets within the preset time period from the scan data and the image data; Determining the loss value of the loss function according to the target-level feature data corresponding to the lidar and the camera respectively under the current calibration matrix, and determining the benefit value of the benefit function according to the loss value of the loss function.

3. The method according to claim 2, characterized in that, The parsing the target-level feature data of all targets within the preset time period from the scan data and the image data includes: Using the current calibration matrix to convert the scan data of the lidar into the camera coordinates of the camera, and obtaining the mapped data of the scan data in the camera coordinates; Performing target-level feature parsing on the mapped data and the image data of the camera to obtain the first target-level feature data corresponding to the image data and the second target-level feature data corresponding to the mapped data; Among them, the determining the loss value of the loss function according to the target-level feature data corresponding to the lidar and the camera respectively under the current calibration matrix includes: determining the loss value of the loss function according to the first target-level feature data and the second target-level feature data.

4. The method according to claim 3, wherein Using the current calibration matrix to perform target-level feature parsing on the mapped data and the image data of the camera to obtain the first target-level feature data corresponding to the image data and the second target-level feature data corresponding to the mapped data includes: Determining the first target-level feature data and the second target-level feature data from the image data of the camera and the mapped data through a target recognition algorithm.

5. The method according to claim 3, wherein The determining the loss value of the loss function according to the first target-level feature data and the second target-level feature data includes: For each target at each moment, calculating the difference between the feature dimensions in the first target-level feature data and the second target-level feature data; Determining the loss value of the loss function according to the difference between the feature dimensions and in combination with the loss function.

6. The method according to claim 5, wherein The method further includes: Extracting the static features of the fixed marker before and after the pose change of the lidar and the camera based on the static features of the fixed marker; the static features include: plane features and scene marker features; The determining the loss value of the loss function according to the difference between the feature dimensions and in combination with the loss function includes: Generating a first loss value of the first loss function according to the difference between the feature dimensions and in combination with the first loss function. Generate the second loss value of the second loss function according to the static features of the fixed marker before and after the pose change of the lidar and the camera, in combination with the second loss function. Calculate the first product of the first loss value and the first preset loss weight, calculate the second product of the second loss value and the second preset loss weight, and use the sum of the first product and the second product as the loss value of the loss function.

7. The method according to claim 6, wherein The first loss function is: Wherein, t represents the moment, T is a preset time period, n represents different targets, N is the number of targets, k1, k2, k3, and k4 are corresponding preset weights, d is the distance, v is the speed, θ is the heading angle, and c is the category.

8. The method according to claim 6, wherein The second loss function is: Wherein, k1 and k2 are preset weights, t represents the moment, T is a preset time period, n represents different targets, N is the number of targets, h is the road surface normal vector, and s is the slope of the road surface marking line.

9. An apparatus for adaptive calibration of a roadside lidar-camera system based on reinforcement learning, characterized in that, It includes: An acquisition module, after the pose of the lidar and the camera changes, acquire the scan data of the lidar and the image data of the camera within a preset time period. A reinforcement learning module, which executes the learning behavior of the agent and iterates the learning behavior of the agent in the direction that makes the benefit function smaller until the benefit value of the benefit function is maximized. The learning behavior is to adjust the calibration matrix between the lidar and the camera, and the benefit function is used to represent the benefit of the coordinate transformation between the lidar and the camera.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.