Image acquisition and fusion method and device
By adopting multi-camera viewing angles and dynamically adjusting camera layout and image fusion schemes in robot vision acquisition, the problems of limited viewing angles, lack of depth information, difficulty in data labeling and poor environmental adaptability in the prior art are solved, and high-quality data acquisition and fusion are achieved, improving the effect of imitation learning and the success rate of robot tasks.
Patent Information
- Application Number
- CN202510535168.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-27
AI Technical Summary
In the prior art, in robot vision collection, there are problems such as limited perspective, lack of depth information, difficulty in data labeling and poor environmental adaptability, which is difficult to adapt to the needs of different task scenarios.
Using multiple camera perspectives, by calculating alignment and fusion weights, dynamically adjusting the camera layout and image fusion scheme to ensure high-quality acquisition and fusion of image data.
It improves the quality of data acquisition and fusion, enhances the effect of imitation learning, improves the success rate and generalization ability of robot tasks, and is suitable for a variety of task scenarios.
Smart Images

Figure CN120075372A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of robot vision and image data processing, and particularly relates to a method and device for image acquisition and fusion. Background Art
[0002] In recent years, imitation learning (IL) has been widely applied in the field of robot learning. Especially with the rise of the concept of large models, imitation learning has been significantly improved in the field of robot control. An imitation learning system usually needs to record relevant operation data through a visual perception module and convert it into actions executable by a robot through a policy learning method. However, existing visual acquisition methods often have the following problems:
[0003] 1. Limited perspective: Traditional cameras are fixedly installed in the experimental environment or on the robot body, which has good effects only for a single task and cannot adapt to generalized data acquisition.
[0004] 2. Lack of depth information: Most cameras have certain requirements for acquisition. When the camera is too close to the task target, some parts such as depth cameras will fail.
[0005] 3. Difficult data annotation: Due to unreasonable selection of perspectives, a large amount of invalid information is included in the acquired data, increasing the complexity of post-processing and raising the data annotation cost.
[0006] 4. Poor environmental adaptability: Existing solutions lack flexible camera arrangement schemes and are difficult to adapt to different task scenarios.
[0007] If multiple camera perspectives are adopted, the need for image data fusion is inevitable in data usage. At the same time, different task scenarios and requirements for image information are different, and maintaining the same fixed camera perspective and image data fusion scheme will not be applicable to all task scenarios. Summary of the Invention
[0008] In view of the deficiencies of the prior art, the present application proposes a method and device for image acquisition and fusion, which adopt multiple camera perspectives, can adjust the placement position of the cameras at any time, automatically adjust the image fusion scheme according to different tasks, improve the quality of data acquisition and fusion, and thus enhance the effect of imitation learning.
[0009] In a first aspect, the present application proposes a method for image acquisition and fusion, including:
[0010] In imitation learning, obtaining video image sequences of multiple fixed cameras;
[0011] According to the video image sequences of each fixed camera, calculating the alignment weight of each fixed camera;
[0012] Taking the corresponding fixed camera with the highest alignment weight value as the reference camera, align the video image sequences of other fixed cameras with the video image sequence of the reference camera according to the time standard of the reference camera;
[0013] Calculate the fusion weight of each fixed camera according to the loss function of the action in imitation learning;
[0014] According to the fusion weight of each fixed camera, splice the aligned video image sequences of multiple fixed cameras to obtain the fused image sequence.
[0015] The calculation of the alignment weight for each fixed camera according to the video image sequence of each fixed camera includes:
[0016] Calculate the clarity and noise level of the image sequence according to the video image sequence of each fixed camera;
[0017] Calculate the comprehensive score of each fixed camera according to the clarity and noise level of the image sequence;
[0018] Calculate the alignment weight of each fixed camera according to the comprehensive score of each fixed camera.
[0019] The calculation formula for the comprehensive score of each fixed camera is as follows: ;
[0020] Wherein, is the comprehensive score of the i-th fixed camera, is the noise level of the image sequence of the i-th fixed camera, is the clarity of the image sequence of the i-th fixed camera, is the minimum value.
[0021] The calculation formula for calculating the alignment weight of each fixed camera according to the comprehensive score of each fixed camera is as follows: ;
[0022] Wherein, is the alignment weight of the i-th fixed camera, exp is the power operation of the natural logarithm e, is the sensitivity coefficient for controlling the weight distribution, is the comprehensive score of the i-th fixed camera, is the comprehensive score of the j-th fixed camera, is the bias of the i-th fixed camera, is the bias of the j-th fixed camera, and M is the number of fixed cameras.
[0023] The calculation formula for calculating the fusion weight of each fixed camera according to the loss function of the action in imitation learning is as follows: ; ;
[0024] Among them, is the loss function for the action of the i-th fixed camera in imitation learning, is the predicted action of the i-th fixed camera in the j-th imitation learning, is the action of the expert demonstration in the j-th imitation learning of the i-th fixed camera, and N is the size of the batch in imitation learning, is the fusion weight in the training of the (t + 1)-th round of imitation learning of the i-th fixed camera, and exp is the power operation of the natural logarithm e, is the sensitivity coefficient for controlling the weight distribution, and M is the number of fixed cameras, is the loss function for the action of the i-th fixed camera in the t-th round of imitation learning training, is the loss function for the action of the j-th fixed camera in the t-th round of imitation learning training.
[0025] Splicing the video image sequences of multiple aligned fixed cameras according to the fusion weight of each fixed camera to obtain a fused image sequence, including:
[0026] Splicing the video image sequences of multiple aligned fixed cameras to obtain a spliced image sequence;
[0027] In the spliced image sequence, smear the video image sequence of the aligned fixed camera with the smallest value of the fusion weight, and use the smeared spliced image sequence as the fused image sequence.
[0028] In a second aspect, the present application proposes an image acquisition and fusion device, including:
[0029] An image acquisition module, configured to acquire video image sequences of multiple fixed cameras in imitation learning;
[0030] An alignment weight calculation module, configured to calculate the alignment weight of each fixed camera according to the video image sequence of each fixed camera;
[0031] An image alignment module, configured to use the fixed camera corresponding to the highest value of the alignment weight as a reference camera, and align the video image sequences of other fixed cameras with the video image sequence of the reference camera according to the time standard of the reference camera;
[0032] A fusion weight calculation module, configured to calculate the fusion weight of each fixed camera according to the loss function of the action in imitation learning;
[0033] An image fusion module, configured to splice the video image sequences of multiple fixed cameras after alignment according to the fusion weights of each fixed camera, to obtain a fused image sequence.
[0034] The image acquisition module includes a plurality of fixed brackets. Fixed cameras are installed on the fixed brackets, and the height and angle of the fixed cameras are adjusted through the fixed brackets. The plurality of fixed cameras are respectively used to acquire video image sequences from the front view, top view, side view, and oblique view.
[0035] In a third aspect, the present application provides an electronic device, including: one or more processors, and a memory for storing instructions. When the instructions are executed by the one or more processors, the one or more processors are caused to execute the method for image acquisition and fusion as described above.
[0036] In a fourth aspect, the present application provides a computer-readable storage medium storing executable instructions, which when executed cause a processor to execute the method for image acquisition and fusion as described above.
[0037] Advantageous effects:
[0038] The present application provides a method and device for image acquisition and fusion. Through an adjustable fixed camera position design, while ensuring image quality, a stable and referenceable time alignment benchmark is provided, thereby reducing the complexity of subsequent image processing and improving the quality of the input data for the strategy. At the same time, combined with an image smear enhancement mechanism with random camera perspectives, the robustness and generalization ability of the strategy model to diverse inputs are further enhanced. Description of the Drawings
[0039] Figure 1 is a flowchart of a method for image acquisition and fusion according to an embodiment of the present application;
[0040] Figure 2 is a schematic block diagram of a device for image acquisition and fusion according to an embodiment of the present application;
[0041] Figure 3 is a schematic diagram of the implementation of the image acquisition module according to an embodiment of the present application;
[0042] Wherein, 1 - fixed bracket, 2 - fixed camera, 3 - mobile robot platform. Detailed Embodiments
[0043] The following will further describe in detail the specific embodiments of the present application with reference to the drawings and embodiments.
[0044] Embodiment 1:
[0045] This embodiment provides a method for image acquisition and fusion, as shown in Figure 1As shown in the figure, it includes:
[0046] Step S1: In imitation learning, obtain video image sequences of multiple fixed cameras;
[0047] In this embodiment, in order to optimize the camera layout so that it can comprehensively cover the key areas of the expert operation and ensure data integrity, multiple fixed cameras are used to take pictures from perspectives such as front view, top view, side view, and oblique view to ensure that the operation area is covered without dead angles. Some cameras can be installed on the robotic arm, and through teleoperation, ensure that the task object always appears within the camera's field of view while performing the task, as Figure 3 shown. The main perspective camera is installed on the fixed bracket 1 with adjustable pose. This fixed bracket 1 can be adjusted with multiple degrees of freedom, and can accurately adjust the height, inclination angle, and orientation of the camera according to the task requirements to ensure the best shooting angle, while maintaining stability to reduce the impact of jitter and perspective shift on data acquisition. This embodiment uses multiple RGB-D cameras, such as Intel RealSense, Azure Kinect, etc., and at the same time, through the adjustment of the camera position, ensure the simultaneous acquisition of high-quality RGB and depth information to improve the quality of 3D data.
[0048] Step S2: According to the video image sequences of each fixed camera, calculate the alignment weight of each fixed camera, including:
[0049] Step S2.1: According to the video image sequences of each fixed camera, calculate the clarity and noise level of the image sequences;
[0050] Step S2.2: According to the clarity and noise level of the image sequences, calculate the comprehensive score of each fixed camera. The calculation formula is as follows: ;
[0051] Where, is the comprehensive score of the i-th fixed camera, is the noise level of the image sequence of the i-th fixed camera, is the clarity of the image sequence of the i-th fixed camera, is the minimum value to prevent the denominator from being zero.
[0052] Among them, the clarity parameter S i =Var(Laplacian(I)), the noise level , where, Var is the variance of the image to be calculated, Laplacian is a Laplacian operator, an image processing filter used to highlight edges and details in the image. A clear image usually contains rich edge information and a strong Laplacian response. I is the image matrix, and its element is the intensity value of the pixel at the coordinate position. is the k-th smoothed region in the image, where K represents the number of smoothed regions randomly selected in the image .
[0053] Step S2.3: Calculate the alignment weight of each fixed camera according to the comprehensive score of each fixed camera. The calculation formula is as follows: ;
[0054] where is the alignment weight of the i-th fixed camera, exp is the power operation of the natural logarithm e is the sensitivity coefficient for controlling the weight distribution. The larger the value of the sensitivity coefficient, the more sensitive it is to the score difference, making the weight more biased towards the camera with the highest score (closer to hard decision); the smaller the value of the sensitivity coefficient, the more evenly the weight is distributed, and it will not be overly biased towards the highest scorer (closer to smooth fusion). In other words: for example, if the three weights are 0.1, 0.1, and 0.8 respectively, multiplying them all by 10 makes the difference larger, and multiplying them all by 0.1 makes the difference smaller. is the comprehensive score of the i-th fixed camera is the comprehensive score of the j-th fixed camera is the bias of the i-th fixed camera is the bias of the j-th fixed camera. Its function is to still select the main camera when the scoring parameter of the main camera is slightly lower than the comprehensive scores of the other cameras. M is the number of fixed cameras.
[0055] Step S3: Take the fixed camera corresponding to the highest alignment weight value as the reference camera, and align the video image sequences of other fixed cameras with the video image sequence of the reference camera according to the time standard of the reference camera;
[0056] In the prior art, usually a certain camera is designated as the reference camera, and other cameras are aligned with the reference camera. However, in practical applications, this cannot meet different task scenarios because different task scenarios have different requirements for image information. Every time the task scenario is changed, the selection of the reference camera needs to be adjusted. Therefore, this application proposes a mobile alignment method, making the method of this application applicable to different task scenarios.
[0057] In this embodiment, for the alignment weight of each fixed camera, the weight of the main view has the highest value. Using the high-quality main view camera as the reference, align the data of other cameras to make the data obtained at the same time point for all views consistent, improving the quality of the picture data. The pictures of each camera are given a comprehensive score according to the clarity parameter and the noise level , and according to the comprehensive scores of the three cameras, select the view through the dynamic smoothing normalization (softmax) weight parameter for data synchronization by comparison.
[0058] Step S4: Calculate the fusion weight of each fixed camera according to the loss function of the action in imitation learning. The calculation formula is as follows: ; ;
[0059] where, is the loss function of the action of the i-th fixed camera in imitation learning, is the predicted action of the i-th fixed camera in the j-th imitation learning, is the action demonstrated by the expert in the j-th imitation learning of the i-th fixed camera. N is the size of the batch in imitation learning, is the fusion weight of the i-th fixed camera in the training of the (t + 1)-th round of imitation learning. exp is the power operation of the natural logarithm e, is the temperature parameter, is the loss function of the action of the i-th fixed camera in the training of the t-th round of imitation learning, is the loss function of the action of the j-th fixed camera in the training of the t-th round of imitation learning. M is the number of cameras.
[0060] Step S5: According to the fusion weight of each fixed camera, splice the aligned video image sequences of multiple fixed cameras to obtain a fused image sequence, including:
[0061] Splice the aligned video image sequences of multiple fixed cameras to obtain a spliced image sequence;
[0062] In the spliced image sequence, smear the video image sequence of the aligned fixed camera with the smallest numerical value of the fusion weight, and use the smeared spliced image sequence as the fused image sequence.
[0063] In the prior art, if the fused image does not meet the requirements, it is usually thought to modify the fusion algorithm to make the fusion algorithm more suitable for different scene requirements. However, in this application, the fusion weight is adjusted to make the fusion result robust and applicable to different scene requirements.
[0064] In this embodiment, the picture data is spliced at the original data level, randomly selected at the cycle times, randomly selected on different cameras through different weight parameters, and the original data is smeared, so as to adjust the data set and improve the robustness of the trained model.
[0065] This embodiment proposes a method for image acquisition and fusion. Multiple fixed cameras are used to obtain video image sequences from the front view, top view, side view, and oblique view respectively. And alignment weights are used to align the video image sequences of other fixed cameras with the video image sequence of the reference camera. Finally, fusion weights are used to splice the aligned video image sequences of multiple fixed cameras to obtain a fused image sequence. This application obtains high-quality visual data, which helps to improve the model training effect of imitation learning, increases the success rate of robot tasks, and enhances the generalization ability of robot task operations. The method of this application can be applied to a variety of task scenarios, and the flexible camera placement scheme is applicable to different industrial production, medical assistance, home service and other scenarios. At the same time, it reduces the situation of manual adjustment, reduces the data processing cost. For sharing picture data sets between different devices, due to the rationality of the position of the main perspective camera, it improves the data acquisition quality, and the optimized camera position design improves the data integrity, reducing the occlusion and information loss problems.
[0066] Embodiment 2:
[0067] This embodiment proposes an apparatus for image acquisition and fusion, as Figure 2 shown, including: an image acquisition module, an alignment weight calculation module, a fusion weight calculation module, and an image fusion module. The image acquisition module is connected to the alignment weight calculation module, the alignment weight calculation module is connected to the fusion weight calculation module, and the fusion weight calculation module is connected to the image fusion module;
[0068] The image acquisition module is used to obtain video image sequences of multiple fixed cameras in imitation learning;
[0069] The alignment weight calculation module is used to calculate the alignment weight of each fixed camera according to the video image sequence of each fixed camera;
[0070] The image alignment module is used to take the fixed camera corresponding to the highest value of the alignment weight as the reference camera, and align the video image sequences of other fixed cameras with the video image sequence of the reference camera according to the time standard of the reference camera;
[0071] The fusion weight calculation module is used to calculate the fusion weight of each fixed camera according to the loss function of the action in imitation learning;
[0072] The image fusion module is used to splice the aligned video image sequences of multiple fixed cameras according to the fusion weight of each fixed camera to obtain a fused image sequence.
[0073] The image acquisition module includes a plurality of fixed brackets 1. A fixed camera 2 is installed on the fixed bracket 1, and the height and angle of the fixed camera 2 are adjusted through the fixed bracket 1. The plurality of fixed cameras 2 are used to respectively obtain video image sequences from the front view, top view, side view, and oblique view, as Figure 3 shown, a plurality of fixed brackets 1 are installed on the mobile robot platform 3, and a fixed camera 2 is installed on the fixed bracket 1.
[0074] Embodiment 3:
[0075] This embodiment provides an electronic device, including: one or more processors, and a memory. The memory is used to store instructions. When the instructions are executed by the one or more processors, the one or more processors are caused to execute the dynamic prediction of the online car-hailing boarding point and the multi-strategy configuration method.
[0076] The electronic device can be a mobile phone, a computer, a tablet computer, etc., including a memory and a processor. A computer program is stored on the memory. When the computer program is executed by the processor, it implements the dynamic prediction of the online car-hailing boarding point and the multi-strategy configuration method as described in the embodiment. It can be understood that the electronic device can also include an input / output (I / O) interface and a communication component.
[0077] Among them, the processor is used to execute all or part of the steps in the dynamic prediction of the online car-hailing boarding point and the multi-strategy configuration method as described in the above embodiment. The memory is used to store various types of data, which can include, for example, instructions of any application program or method in the electronic device, and data related to the application program.
[0078] The processor can be implemented by an application specific integrated circuit (ASIC), a digital signal processor (DSP), a programmable logic device (PLD), a field programmable gate array (FPGA), a controller, a microcontroller, a microprocessor, or other electronic components, and is used to execute the dynamic prediction of the online car-hailing boarding point and the multi-strategy configuration method described in the above embodiment.
[0079] Embodiment 4:
[0080] This embodiment provides a computer-readable storage medium, which stores executable instructions. When the instructions are executed, if they are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium.
[0081] The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method for dynamic prediction and multi-strategy configuration of online car-hailing boarding points described in the various embodiments of the present application.
[0082] The aforementioned storage media include: flash memory, hard disk, multimedia card, card-type memory (for example, SD (Secure Digital Memory Card) or DX (abbreviation of Memory Data Register, MDR) memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, CD, server, APP (Application, abbreviation of application software) application store and other media that can store program verification codes, on which computer programs are stored. When the computer program is executed by the processor, the various steps of the above-mentioned online car-hailing boarding point dynamic prediction and multi-strategy configuration method can be implemented.
[0083] The various embodiments in the present application are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.
[0084] The protection scope of the present application is not limited to the above-mentioned embodiments. Obviously, those skilled in the art can make various changes and modifications to the present disclosure without departing from the scope and spirit of the present disclosure. If these changes and modifications fall within the scope of the claims of the present disclosure and their equivalents, the intention of the present disclosure also includes these changes and modifications.
Claims
1. A method for image acquisition and fusion, characterized in that: include: In imitation learning, multiple video image sequences of fixed cameras are obtained; Calculate the alignment weight of each fixed camera according to the video image sequence of each fixed camera; The corresponding fixed camera with the highest alignment weight is used as the reference camera, and the video image sequences of other fixed cameras are aligned with the video image sequence of the reference camera according to the time standard of the reference camera; Calculate the fusion weight of each fixed camera according to the loss function of the action in imitation learning; According to the fusion weight of each fixed camera, the video image sequences of multiple aligned fixed cameras are spliced to obtain a fused image sequence; The step of calculating the alignment weight of each fixed camera according to the video image sequence of each fixed camera includes: According to the video image sequence of each fixed camera, the clarity and noise level of the image sequence are calculated; Calculate a comprehensive score for each fixed camera based on the clarity and noise level of the image sequence; According to the comprehensive score of each fixed camera, the alignment weight of each fixed camera is calculated.
2. The method for image acquisition and fusion according to claim 1, characterized in that: The comprehensive score of each fixed camera is calculated as follows: ; in, is the comprehensive score of the i-th fixed camera, is the noise level of the image sequence of the i-th fixed camera, is the clarity of the image sequence of the i-th fixed camera, is a minimum value.
3. The method for image acquisition and fusion according to claim 1, characterized in that: According to the comprehensive score of each fixed camera, the alignment weight of each fixed camera is calculated, and the calculation formula is as follows: ; in, is the alignment weight of the i-th fixed camera, exp is the power operation of the natural logarithm e, To control the sensitivity coefficient of weight distribution, is the comprehensive score of the i-th fixed camera, is the comprehensive score of the jth fixed camera, is the bias of the i-th fixed camera, is the offset of the jth fixed camera, and M is the number of fixed cameras.
4. The method for image acquisition and fusion according to claim 1, characterized in that: The fusion weight of each fixed camera is calculated according to the loss function of the action in imitation learning, and the calculation formula is as follows: ; ; in, is the loss function of the action of the i-th fixed camera in imitation learning, is the predicted action in the jth imitation learning of the i-th fixed camera, is the action demonstrated by the expert in the jth imitation learning of the i-th fixed camera, N is the batch size in the imitation learning, is the fusion weight of the i-th fixed camera in the t+1 round of imitation learning training, exp is the power operation of the natural logarithm e, To control the sensitivity coefficient of weight distribution, M is the number of fixed cameras, is the loss function of the action of the i-th fixed camera in the t-th round of imitation learning training, is the loss function of the action of the j-th fixed camera in the t-th round of imitation learning training.
5. The method for image acquisition and fusion according to claim 1, characterized in that: The step of splicing the video image sequences of the aligned multiple fixed cameras according to the fusion weight of each fixed camera to obtain a fused image sequence includes: Stitching the aligned video image sequences of multiple fixed cameras to obtain a stitched image sequence; In the stitched image sequence, the video image sequence of the aligned fixed camera with the smallest fusion weight is smeared, and the stitched image sequence after smearing is used as the fused image sequence.
6. An image acquisition and fusion device, characterized in that: include: An image acquisition module, used to acquire video image sequences of multiple fixed cameras in imitation learning; An alignment weight calculation module, used for calculating the alignment weight of each fixed camera according to the video image sequence of each fixed camera; An image alignment module is used to align the video image sequences of other fixed cameras with the video image sequence of the reference camera according to the time standard of the reference camera, taking the corresponding fixed camera with the highest alignment weight as the reference camera; The fusion weight calculation module is used to calculate the fusion weight of each fixed camera according to the loss function of the action in imitation learning; The image fusion module is used to stitch the video image sequences of multiple aligned fixed cameras according to the fusion weight of each fixed camera to form a fused image sequence.
7. The image acquisition and fusion device according to claim 6, characterized in that: The image acquisition module includes multiple fixed brackets, fixed cameras are installed on the fixed brackets, and the height and angle of the fixed cameras are adjusted by the fixed brackets. The multiple fixed cameras are used to obtain video image sequences from the main view, top view, side view and oblique view respectively.
8. An electronic device, characterized in that: include: One or more processors, and a memory, wherein the memory is used to store instructions, and when the instructions are executed by the one or more processors, the one or more processors execute the image acquisition and fusion method described in any one of claims 1 to 5.
9. A computer-readable storage medium, characterized in that: It stores executable instructions, which, when executed, enable the processor to execute the image acquisition and fusion method described in any one of claims 1 to 5.
Citation Information
Patent Citations
Multi-exposure video fusion method and multi-exposure video fusion device
CN106251365A
Indoor visual navigation method based on causal attention
CN115512214A
Industrial robot image processing method based on image fusion
CN118521859A
Multi-source image fusion method
CN118537230A
Device and method for obtaining image sets
RU2797757C1