Joint compression method and device of visible light image and point cloud data, electronic equipment and storage medium
By converting 3D point cloud data to a 2D coordinate system and combining it with a joint compression model of visible light images and 2D point cloud data, redundant features are eliminated, achieving better joint compression of image and point cloud data. This solves the problem of low compression accuracy in existing technologies and improves data storage and transmission efficiency.
Patent Information
- Application Number
- CN202411409848.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-10
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2044-10-10
AI Technical Summary
The compression accuracy of image data and point cloud data in existing technologies is low, resulting in low efficiency in the storage and transmission of massive amounts of data.
By converting 3D point cloud data to a 2D coordinate system, and combining visible light images and 2D point cloud data, a joint compression model is used for analysis, transformation, fusion, and separation to eliminate redundant features and achieve joint compression of images and point clouds.
By eliminating redundant features, the bitstream can be compressed to a greater extent, improving data compression efficiency and enhancing storage and transmission efficiency.
Smart Images

Figure CN119583820B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of artificial intelligence, and in particular to a joint compression method and device for visible light images and point cloud data, an electronic device and a storage medium. BACKGROUND
[0002] With the rapid development of the Internet of Things, autonomous driving, virtual reality, augmented reality, unmanned aerial vehicles and other fields, the amount of point cloud data and image data generated is growing exponentially. For example, autonomous vehicles generate a large amount of point cloud data during driving to perceive the surrounding environment in real time; virtual reality and augmented reality applications require a large amount of image data to provide immersive visual experiences. The massive storage and efficient transmission of these data have become urgent problems to be solved.
[0003] Current compression techniques for image data and point cloud data have low accuracy. SUMMARY
[0004] In view of the above, the purpose of the present disclosure is to propose a joint compression method and device for visible light images and point cloud data, which can solve the existing problems.
[0005] To achieve the above purpose, in a first aspect, the present disclosure provides a joint compression method for visible light images and point cloud data, comprising: in response to obtaining a visible light image and three-dimensional point cloud data of the same scene, converting the three-dimensional point cloud data into a two-dimensional coordinate system to obtain two-dimensional point cloud data, wherein the depth information of the three-dimensional point cloud data is the value of the point cloud in the two-dimensional coordinate system; calling the joint compression model to analyze and transform the visible light image and the two-dimensional point cloud data respectively to obtain visible light image features and two-dimensional point cloud features; calling the joint compression model to fuse the visible light image features and the two-dimensional point cloud features to eliminate redundant features, and performing a preset comprehensive transformation on the fusion result; calling the joint compression model to separate the image and the point cloud after the preset comprehensive transformation of the fusion result to obtain a reconstructed image and a reconstructed point cloud.
[0006] In a second aspect, a joint compression device of a visible light image and point cloud data is also provided, comprising: an acquisition unit configured to, in response to acquisition of a visible light image and three-dimensional point cloud data of a same scene, convert the three-dimensional point cloud data into a two-dimensional coordinate system to obtain two-dimensional point cloud data, wherein depth information of the three-dimensional point cloud data is a value of the point cloud in the two-dimensional coordinate system; an analysis unit configured to call the joint compression model, and perform analysis transformation on the visible light image and the two-dimensional point cloud data respectively to obtain visible light image features and two-dimensional point cloud features; a fusion unit configured to call the joint compression model, and perform fusion on the visible light image features and the two-dimensional point cloud features to eliminate redundant features, and perform a preset comprehensive transformation on a fusion result; and a separation unit configured to call the joint compression model, and perform image and point cloud separation on the fusion result after the preset comprehensive transformation to obtain a reconstructed image and a reconstructed point cloud.
[0007] In a third aspect, an electronic device is also provided, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the method of the first aspect.
[0008] In a fourth aspect, a computer readable storage medium is also provided, which stores a computer program, wherein the computer program is executed by a processor to implement the method of any one of the first aspect.
[0009] In a fifth aspect, a computer program product is also provided, comprising a computer program, wherein the computer program is executed by a processor to implement the method of any one of the first aspect.
[0010] In general, the present disclosure has at least the following beneficial effects: the embodiments can jointly compress the visible light image and the point cloud data in the same scene through the joint compression model, so as to play the similarity and complementarity between the two kinds of data and achieve better compression effect. By eliminating the redundant features, the code stream can be compressed to a greater extent to achieve better compression effect. BRIEF DESCRIPTION OF DRAWINGS
[0011] In the drawings, like reference numerals refer to like elements throughout the various drawings. The drawings are not necessarily to scale, the emphasis instead being placed on illustrating principles of the present disclosure. It should be understood that the drawings are merely depictions of some embodiments disclosed herein and should not be construed as limiting the scope of the present disclosure.
[0012] Figure 1 A flowchart of a joint compression method of a visible light image and point cloud data according to an embodiment of the present disclosure is shown;
[0013] Figure 2A schematic diagram of a fusion transformation module according to an embodiment of the present disclosure is shown.
[0014] Figure 3 A schematic diagram of a separation module according to an embodiment of the present disclosure is shown.
[0015] Figure 4 A schematic diagram of a joint compression model according to an embodiment of the present disclosure is shown.
[0016] Figure 5 A schematic diagram of a joint compression device for visible light images and point cloud data according to an embodiment of the present disclosure is shown.
[0017] Figure 6 A structural schematic diagram of an electronic device provided by an embodiment of the present disclosure is shown.
[0018] Figure 7 A schematic diagram of a storage medium provided by an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0019] The present disclosure will be further described below in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related application, and not to limit the application. In addition, it should be noted that only parts related to the application are shown in the drawings for ease of description.
[0020] It should be noted that the embodiments in the present disclosure and the features in the embodiments can be combined with each other without conflict. The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.
[0021] Figure 1 A joint compression method for visible light images and point cloud data is shown in the present disclosure. In an embodiment of the present disclosure, the method comprises:
[0022] In step S101, in response to obtaining a visible light image and three-dimensional point cloud data of the same scene, the three-dimensional point cloud data is converted into a two-dimensional coordinate system to obtain two-dimensional point cloud data, wherein the depth information of the three-dimensional point cloud data is the value of the point cloud in the two-dimensional coordinate system.
[0023] In the present embodiment, the execution subject of the joint compression method for visible light images and point cloud data can convert the three-dimensional point cloud data into a two-dimensional coordinate system in response to obtaining a visible light image and three-dimensional point cloud data of the same scene. The mapping result is two-dimensional point cloud data. The execution subject can be a server or a terminal or various electronic devices.
[0024] The same scene can be, for example, the same location or the same area.
[0025] The visible light image is based on the data obtained after uniform sampling in two-dimensional space by photoelectric conversion. If expressed in the form of an array, it is a three-dimensional array of CxHxW, where C is the channel, usually a constant of 3, and H and W represent the height and width of the image, respectively. The point cloud data is usually calculated by millimeter wave radar and laser radar by emitting laser pulses and measuring the time of the reflected echo to calculate the distance of the target, thereby generating three-dimensional point cloud data. The point cloud data can be expressed in the form of an NXM array, where N represents the number of points in the point cloud data, and M represents the characteristics of the point, including geometric characteristics and attribute characteristics. The geometric characteristics are the spatial positions x, y, and z of the point. The attribute characteristics of the point cloud include reflection intensity, color information, etc. We choose to convert the point cloud data to an image. A formula is provided as follows:
[0026]
[0027] where x is the input point cloud, which is a 3xN array, Tvelocam is the rotation and translation matrix R3x4 from the radar to the camera. In actual calculation, the 3x4 matrix needs to be expanded to a 4x4 matrix by adding a fourth row vector [0, 0, 0, 1]. Rrect(0) is the camera rotation correction matrix R3x3. In actual calculation, the 3x3 matrix needs to be expanded to a 4x4 vector by adding a fourth row vector [0, 0, 0] and then adding a fourth column vector [0, 0, 0, 1]T. Prect(i) is the corrected camera projection matrix R3x4, where 0, 1, 2, and 3 represent the number of cameras. 0 represents the left gray camera, 1 represents the right gray camera, 2 represents the left color camera, and 3 represents the right color camera. The point x in the laser radar coordinate system is projected into the left color image y using the following formula:
[0028]
[0029] The final 3xN data y is obtained, and then the first row data u is divided by the second row data to obtain the final y, where the first row data u corresponds to the width W of the image, and the second row data v corresponds to the height H of the image. In this way, the three-dimensional point cloud data is successfully converted to the pixel coordinate system of the two-dimensional image. Then we take the depth information of the point cloud (the first row data of x represents the depth information) as the value of the point cloud in the pixel coordinate system. In this way, we obtain a point cloud data in the form of 1xHxW.
[0030] In step S102, the joint compression model is called to analyze and transform the visible light image and the two-dimensional point cloud data respectively to obtain visible light image features and two-dimensional point cloud features.
[0031] In this embodiment, the execution subject can call the joint compression model to analyze and transform the visible light image and the two-dimensional point cloud data respectively to obtain visible light image features corresponding to the visible light image and two-dimensional point cloud features corresponding to the two-dimensional point cloud data. The execution subject can obtain the features in various ways. For example, in the case where the joint compression model includes a convolution layer, the execution subject can use the convolution layer to convolve the visible light image and the two-dimensional point cloud data to obtain the corresponding features.
[0032] The feature generation process can achieve a certain degree of data compression.
[0033] In step S103, the execution subject calls the joint compression model to fuse the visible light image features and the two-dimensional point cloud features to eliminate redundant features between channels and performs a preset comprehensive transformation on the fusion result.
[0034] In this embodiment, the execution subject can call the joint compression model to fuse the visible light image features and the two-dimensional point cloud features and perform a preset comprehensive transformation on the fusion result. For example, the fusion can be splicing. In this way, the execution subject can uniformly digitally represent the image and the point cloud data, map the three-dimensional point cloud to two-dimensional data in the form of 1xHxW, and splice the image data in the channel direction to form data in the form of 4xHxW.
[0035] The fusion process includes a step of eliminating redundant features between channels. The preset comprehensive transformation can be a transformation of the number of channels.
[0036] In step S104, the execution subject calls the joint compression model to separate the image and the point cloud from the fusion result after the preset comprehensive transformation to obtain a reconstructed image and a reconstructed point cloud.
[0037] In this embodiment, the joint compression model has a separation function and can determine the image content and the point cloud content from the fusion result after the preset comprehensive transformation.
[0038] This embodiment can use the joint compression model to jointly compress the visible light image and the point cloud data in the same scene, thereby taking advantage of the similarity and complementarity between the two types of data to achieve better compression effect. By eliminating redundant features, the code stream can be compressed to a greater extent to achieve better compression effect.
[0039] In some optional implementations of any of the embodiments of the present disclosure, the joint compression model comprises an analysis transformation module for images and an analysis transformation module for point clouds; and the calling the joint compression model to perform analysis transformation on the visible light image and the two-dimensional point cloud data respectively comprises: calling the analysis transformation module for images to process the visible light image to perform image analysis transformation on the visible light image to obtain an image latent representation, the image latent representation being a feature of the visible light image; and calling the analysis transformation module for point clouds to process the two-dimensional point cloud data to perform point cloud analysis transformation on the two-dimensional point cloud data to obtain a point cloud latent representation, the point cloud latent representation being a feature of the two-dimensional point cloud.
[0040] In some optional implementations of any of the embodiments of the present disclosure, the joint compression model further comprises a fusion transformation module; and the calling the joint compression model to fuse the feature of the visible light image and the feature of the two-dimensional point cloud comprises: using the fusion transformation module to perform channel dimension splicing on the image latent representation and the point cloud latent representation to obtain a feature splicing result; and processing the feature splicing result through an attention mechanism to obtain a fusion result.
[0041] In these optional implementations, the depth information of the point cloud (the first row of data of x represents the depth information) is taken as the value of the point cloud in the pixel coordinate system. In this way, a point cloud data in the form of 1xHxW is obtained. The image data is 3xHxW. Splicing the image data and the point cloud data in the channel dimension will obtain a joint digital representation form of an image and a point cloud, i.e., 4xHxW.
[0042] Optionally, the processing the splicing result through the attention mechanism comprises: performing convolution on the feature splicing result through a preset convolution kernel to perform spatial attention mechanism processing on the feature splicing result, and further finding the repeated features between different channels in the feature splicing result as first repeated features; assigning a feature value to each channel in the feature splicing result, and finding the repeated features between different channels in the feature splicing result as second repeated features according to the feature value; and removing the repeated features between different channels in the feature splicing result according to the first repeated features and the second repeated features.
[0043] The attention mechanism can comprise a spatial attention mechanism for convolution and a channel attention mechanism for value assignment. Both of the two attention mechanisms can find the repeated features between different channels in the feature splicing result and remove them, achieving the effect of eliminating redundancy.
[0044] As Figure 2As shown, the feature fusion transformation module of the image and the point cloud in the figure includes a stitching layer, a convolution layer and a CBAM (Convolutional Block Attention Module) module. The fusion transformation module is to splice the latent representations of the image and the point cloud in the channel dimension, then perform 1x1 convolution, and then send the result to the CBAM module.
[0045] The innovation of this part is to enhance the attention of the network to the common features of the image and the point cloud through the channel attention mechanism, and eliminate redundancy.
[0046] In some optional implementations of any of the embodiments of the present disclosure, the joint compression model further includes a separation module; and the calling the joint compression model to separate the preset fusion result after the comprehensive transformation includes: inputting the preset fusion result after the comprehensive transformation into an image branch of the separation module to obtain the reconstructed image; and inputting the preset fusion result after the comprehensive transformation into a point cloud branch of the separation module to obtain the reconstructed point cloud.
[0047] The separation module of the image and the point cloud has two branches. The first branch is the image branch, the input is the result after the comprehensive transformation, the size is Bx4xHxW, and the reconstructed image of Bx3xHxW is obtained through the image division module. The second branch is the point cloud branch, and the result after the comprehensive transformation is also input into the point cloud division module to obtain the reconstructed point cloud of Bx1xHxW.
[0048] The placeholder in the point cloud division module is obtained from the input point cloud information. The input point cloud is an array of Bx1xHxW shape. Since the point cloud is sparse, only the element values of the positions of the points are not zero in HxW, and the element values of the remaining positions are all zero. The placeholder is to set the element values of the positions of the points to 1. Multiplying the placeholder and the reconstructed point cloud element by element can reduce the distortion of the reconstructed point cloud.
[0049] Optionally, the three-dimensional point cloud data is sparse point cloud data, the point cloud branch includes an attention block, a first residual block, a second residual block, a placeholder and a convolution layer, the placeholder is an image corresponding to the positions of the points in the three-dimensional point cloud data; and the inputting the preset fusion result after the comprehensive transformation into the point cloud branch of the separation module to obtain the reconstructed point cloud includes: inputting the preset fusion result after the comprehensive transformation into the attention block, and inputting the output result of the attention block into the first residual block; performing a preset processing on the processing result of the first residual block by using the placeholder, and performing convolution on the result of the preset processing by using the convolution layer; inputting the convolution result into the second residual block, and performing the preset processing on the result of the second residual block by using the placeholder to obtain the reconstructed point cloud.
[0050] A division module of image and point cloud information is proposed, and a placeholder is introduced to improve the reconstruction effect.
[0051] As shown in Figure 3 , a schematic diagram of the separation module is shown.
[0052] In some optional implementations of any embodiment of the present disclosure, the number of channels of the visible light image is 3, the number of channels of the two-dimensional point cloud data is 1, and the number of channels of both the visible light image feature and the two-dimensional point cloud feature is greater than 3; the preset comprehensive transformation on the fusion result includes: reconstructing the fusion result into data with a channel number of 4 to obtain the fusion result after the preset comprehensive transformation.
[0053] In these optional implementations, the number of channels of both the visible light image feature and the two-dimensional point cloud feature is 192, and the data form of both is 192x16x16.
[0054] In some optional implementations of any embodiment of the present disclosure, the training step of the joint compression model includes: adjusting the total code rate of both the reconstructed image and the reconstructed point cloud by adjusting the weight of distortion to train the joint compression model.
[0055] According to rate-distortion optimization, high code rate R will result in small distortion D, and the selection of code rate and distortion is a trade-off. The Lagrange factor λ adjusts the size of the code rate by adjusting the weight of distortion. The optimization objective is as follows: L=R+λD, but in the model we propose, data of two modalities of image and point cloud are involved, so our optimization objective is adjusted as follows: L=R total +λ img D img +λ pc D pc , where R total represents the total code rate of image and point cloud, D img and D pc represent the distortion of image and point cloud respectively, λ img and λ pc represent the distortion weight of image and point cloud respectively.
[0056] In some optional implementations of any embodiment of the present disclosure, the overall learning-based image and point cloud joint compression framework is as shown in Figure 4As shown, the analysis transformation module of the image, the analysis transformation module of the point cloud, the feature fusion transformation module of the image and the point cloud, the comprehensive transformation module, the separation module of the image and the point cloud, the super-prior analysis module, the super-prior comprehensive module, the quantization module, the entropy model module, and the entropy encoding function are shown. The analysis module of the image analysis module, the analysis module of the point cloud changes the channel number of the analysis module to 1, the comprehensive transformation module changes the output channel of the comprehensive transformation module to 4, the super-prior analysis transformation module and the super-prior comprehensive module use the super-prior analysis transformation module and the super-prior comprehensive module. The entropy model uses the entropy model, and the entropy estimation uses the Gaussian mean scale distribution.
[0057] Among them, the analysis transformation module of the image, the analysis transformation module of the point cloud, and the entropy encoding function can realize the compression process of the image.
[0058] The feature fusion transformation module of the image and the point cloud is to splice the image latent representation obtained by image analysis transformation and the point cloud latent representation obtained by point cloud analysis transformation in the channel dimension, and then pass through the spatial and channel attention mechanism to eliminate the redundancy between the channels and obtain the fused image and point cloud latent representation. The separation module of the image and the point cloud is to separate the reconstructed image and point cloud fusion representation obtained after the comprehensive transformation into a reconstructed image and a reconstructed point cloud.
[0059] The embodiment of the present disclosure provides a joint compression device for visible light image and point cloud data, which is used to execute the joint compression method for visible light image and point cloud data as described in the above embodiment, such as Figure 5 As shown, the device comprises: an acquisition unit 501 configured to, in response to acquiring a visible light image and three-dimensional point cloud data of the same scene, convert the three-dimensional point cloud data into a two-dimensional coordinate system to obtain two-dimensional point cloud data, wherein the depth information of the three-dimensional point cloud data is the value of the point cloud in the two-dimensional coordinate system; an analysis unit 502 configured to call the joint compression model, and perform analysis transformation on the visible light image and the two-dimensional point cloud data respectively to obtain visible light image features and two-dimensional point cloud features; a fusion unit 503 configured to call the joint compression model, and fuse the visible light image features and the two-dimensional point cloud features to eliminate redundant features, and perform a preset comprehensive transformation on the fusion result; and a separation unit 504 configured to call the joint compression model, and separate the fusion result after the preset comprehensive transformation into an image and a point cloud to obtain a reconstructed image and a reconstructed point cloud.
[0060] The joint compression device for visible light image and point cloud data provided by the above embodiment of the present disclosure has the same beneficial effects as the method adopted, run or implemented by the application program stored therein, based on the same inventive concept as the joint compression method for visible light image and point cloud data provided by the embodiment of the present disclosure.
[0061] The embodiments of the present disclosure further provide an electronic device corresponding to the joint compression method of the visible light image and the point cloud data provided by the foregoing embodiments, to perform the joint compression method of the visible light image and the point cloud data. The embodiments of the present disclosure are not limited in this regard.
[0062] Reference is made to Figure 6 which shows a schematic diagram of an electronic device provided by some embodiments of the present disclosure. As Figure 6 shown, the electronic device 60 includes a processor 600, a memory 601, a bus 602 and a communication interface 603, the processor 600, the communication interface 603 and the memory 601 are connected through the bus 602; the memory 601 stores a computer program executable on the processor 600, and the processor 600 executes the computer program to perform the method provided by any one of the foregoing embodiments of the present disclosure.
[0063] The memory 601 can include a high-speed random access memory (RAM: Random Access Memory) and can also include a non-volatile memory such as at least one disk memory. The communication between the system network element and at least one other network element is realized through at least one communication interface 603 (which can be wired or wireless), and the Internet, a wide area network, a local network, a metropolitan area network, etc. can be used.
[0064] The bus 602 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. The memory 601 is used to store programs, and the processor 600 executes the programs after receiving execution instructions. The joint compression method of the visible light image and the point cloud data disclosed in any one of the foregoing embodiments of the present disclosure can be applied in the processor 600 or realized by the processor 600.
[0065] The processor 600 can be an integrated circuit chip with a processing capability of signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of hardware in the processor 600 or the instruction in the form of software. The processor 600 described above can be a general processor, including a central processing unit (CPU), a network processor (NP), etc.; can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a ready programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. Each method, step and logic block diagram disclosed in the embodiments of the present disclosure can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiments of the present disclosure can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read only memory, a programmable read only memory or an electrically erasable programmable memory, a register or other mature storage medium in the art. The storage medium is located in the memory 601, and the processor 600 reads the information in the memory 601 and combines the hardware to complete the steps of the above method.
[0066] The electronic device provided by the embodiments of the present disclosure and the joint compression method of the visible light image and the point cloud data provided by the embodiments of the present disclosure have the same beneficial effects as the method adopted, run or implemented.
[0067] The embodiments of the present disclosure also provide a computer readable storage medium corresponding to the joint compression method of the visible light image and the point cloud data provided by the preceding embodiments. Please refer to Figure 7 The computer readable storage medium shown in the figure is an optical disc 70, and a computer program (i.e. program product) is stored on the optical disc 70. When the computer program is run by a processor, the joint compression method of the visible light image and the point cloud data provided by any of the preceding embodiments is executed.
[0068] It should be noted that examples of the computer readable storage medium can also include, but are not limited to, a phase change memory (PRAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), other types of random access memory (RAM), a read only memory (ROM), an electrically erasable programmable read only memory (EEPROM), a flash memory or other optical, magnetic storage medium, which will not be described one by one here.
[0069] The computer readable storage medium provided by the above embodiments of the present disclosure has the same beneficial effects as the method for joint compression of the visible light image and the point cloud data provided by the embodiments of the present disclosure, and has the same beneficial effects as the method adopted, run or implemented by the application program stored therein.
[0070] It should be noted that:
[0071] It should be noted that:
[0072] From the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment method can be realized by means of software and the necessary general hardware platform, of course, it can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present disclosure can be embodied in the form of a software product in essence or in the form of a contribution to the prior art. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a plurality of instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the method described in each embodiment of the present disclosure.
[0073] The embodiments of the present disclosure are described above in conjunction with the accompanying drawings, which are merely specific embodiments of the present disclosure, but the present disclosure is not limited to the above specific embodiments. The above specific embodiments are merely illustrative, not restrictive, and those skilled in the art can make many forms without departing from the scope of the present disclosure and the scope of protection of the claims under the inspiration of the present disclosure.
Claims
1. A method for joint compression of visible light images and point cloud data, characterized in that, The method comprises the steps of: in response to obtaining a visible light image and three-dimensional point cloud data of the same scene, converting the three-dimensional point cloud data into a two-dimensional coordinate system to obtain two-dimensional point cloud data, wherein the depth information of the three-dimensional point cloud data is the value of the point cloud in the two-dimensional coordinate system; calling a joint compression model to analyze and transform the visible light image and the two-dimensional point cloud data respectively to obtain visible light image features and two-dimensional point cloud features, wherein the joint compression model comprises an image analysis and transformation module and a point cloud analysis and transformation module; calling the joint compression model to fuse the visible light image features and the two-dimensional point cloud features to eliminate redundant features between channels and to perform a preset comprehensive transformation on the fusion result; calling the joint compression model to separate the fusion result after the preset comprehensive transformation into an image and a point cloud to obtain a reconstructed image and a reconstructed point cloud; the joint compression model further comprises a separation module; and the calling of the joint compression model to separate the fusion result after the preset comprehensive transformation comprises: inputting the fusion result after the preset comprehensive transformation into an image branch of the separation module to obtain the reconstructed image; and inputting the fusion result after the preset comprehensive transformation into a point cloud branch of the separation module to obtain the reconstructed point cloud; the number of channels of the visible light image is 3, the number of channels of the two-dimensional point cloud data is 1, and the number of channels of both the visible light image features and the two-dimensional point cloud features is greater than 3; the preset comprehensive transformation on the fusion result comprises: reconstructing the fusion result into data with a channel number of 4 to obtain the fusion result after the preset comprehensive transformation; the joint compression model comprises a fusion transformation module; and the calling of the joint compression model to fuse the visible light image features and the two-dimensional point cloud features comprises: using the fusion transformation module to perform channel dimension splicing on the image latent representation and the point cloud latent representation to obtain a feature splicing result; and processing the feature splicing result through an attention mechanism to obtain a fusion result; the processing of the splicing result through the attention mechanism comprises: performing convolution on the feature splicing result through a preset convolution kernel to find repeated features between different channels in the feature splicing result as first repeated features; assigning a feature value to each channel in the feature splicing result, finding repeated features between different channels in the feature splicing result as second repeated features according to the feature value, and removing the repeated features between different channels in the feature splicing result according to the first repeated features and the second repeated features.
2. The method of claim 1, wherein, the calling of the joint compression model to analyze and transform the visible light image and the two-dimensional point cloud data respectively comprises: calling the image analysis and transformation module to process the visible light image to perform image analysis and transformation on the visible light image to obtain an image latent representation, wherein the image latent representation is the visible light image feature; calling the point cloud analysis and transformation module to process the two-dimensional point cloud data to perform point cloud analysis and transformation on the two-dimensional point cloud data to obtain a point cloud latent representation, wherein the point cloud latent representation is the two-dimensional point cloud feature.
3. The method of claim 1, wherein, The three-dimensional point cloud data is sparse point cloud data, the point cloud branch includes an attention block, a first residual block, a second residual block, an occupancy map, and a convolution layer, and the occupancy map is an image corresponding to a position where a point exists in the three-dimensional point cloud data; The step of inputting the fusion result after the preset comprehensive transformation into the point cloud branch of the separation module to obtain the reconstructed point cloud includes: inputting the fusion result after the preset comprehensive transformation into the attention block, and inputting a result output by the attention block into the first residual block; performing preset processing on a processing result of the first residual block by using the occupancy map, and performing convolution on a result of the preset processing by using a convolution layer; inputting a convolution result into the second residual block, and performing the preset processing on a result of the second residual block by using the occupancy map to obtain the reconstructed point cloud.
4. The method of claim 1, wherein, The image branch includes an attention block, a third residual block, and a convolution layer; The step of inputting the fusion result after the preset comprehensive transformation into the image branch of the separation module to obtain the reconstructed image includes: inputting the fusion result after the preset comprehensive transformation into the attention block in the image branch, and inputting a result output by the attention block into the third residual block; inputting a result of the third residual block into the convolution layer in the image branch to obtain the reconstructed image.
5. The method of claim 1, wherein, The training step of the joint compression model includes: adjusting a total code rate of both the reconstructed image and the reconstructed point cloud by adjusting a weight of distortion to train the joint compression model.
6. An apparatus for joint compression of visible light images and point cloud data, characterized in that, includes: an acquisition unit configured to, in response to acquiring a visible light image and three-dimensional point cloud data of a same scene, convert the three-dimensional point cloud data into a two-dimensional coordinate system to obtain two-dimensional point cloud data, wherein depth information of the three-dimensional point cloud data is a value of point cloud in the two-dimensional coordinate system; an analysis unit configured to call a joint compression model to perform analysis transformation on the visible light image and the two-dimensional point cloud data respectively to obtain visible light image features and two-dimensional point cloud features, the joint compression model including an analysis transformation module of an image and an analysis transformation module of point cloud; a fusion unit configured to call the joint compression model to fuse the visible light image features and the two-dimensional point cloud features to eliminate redundant features, and perform preset comprehensive transformation on a fusion result; a separation unit configured to call the joint compression model to separate the fusion result after the preset comprehensive transformation into an image and point cloud to obtain a reconstructed image and a reconstructed point cloud; the joint compression model further includes a separation module; and the separation unit is further configured to perform the calling of the joint compression model to separate the fusion result after the preset comprehensive transformation in the following manner: input the fusion result after the preset comprehensive transformation into an image branch of the separation module to obtain the reconstructed image, and input the fusion result after the preset comprehensive transformation into a point cloud branch of the separation module to obtain the reconstructed point cloud; a channel number of the visible light image is 3, a channel number of the two-dimensional point cloud data is 1, and a channel number of both the visible light image features and the two-dimensional point cloud features is greater than 3. The fusion unit is further configured to perform the preset comprehensive transformation on the fusion result in the following manner: reconstruct the fusion result into data with a channel number of 4 to obtain the fusion result after the preset comprehensive transformation; The joint compression model comprises a fusion transformation module; the fusion unit is further configured to perform the calling of the joint compression model to fuse the visible light image feature and the two-dimensional point cloud feature in the following manner: perform channel dimension splicing on the image latent representation and the point cloud latent representation by using the fusion transformation module to obtain a feature splicing result; and process the feature splicing result by using an attention mechanism to obtain a fusion result; The fusion unit is further configured to perform the processing of the splicing result by using the attention mechanism in the following manner: perform convolution on the feature splicing result by using a preset convolution kernel to find repeated features between different channels in the feature splicing result as first repeated features; assign a feature value to each channel in the feature splicing result, find repeated features between different channels in the feature splicing result as second repeated features according to the feature value, and remove the repeated features between different channels in the feature splicing result according to the first repeated features and the second repeated features.
Citation Information
Patent Citations
Multi-modal fusion target detection system and method
CN116778288A
Video compression coding method based on intelligent feature clustering
CN117528085A