Remote inspection method and device for customs items, medium and equipment

Through multi-directional photography of depth cameras and color cameras, three-dimensional reconstruction is carried out in combination with multi-view and cost-based methods, and the text information on the surface of the item is recognized using a multi-modal large language model, solving the problems of three-dimensional reconstruction and multi-lingual text recognition under sparse view conditions, and achieving efficient remote inspection of customs items.

CN120032076APending Publication Date: 2025-05-23AISINO CORPORATION
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411915890.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art is difficult to perform effective three-dimensional reconstruction in scenarios with limited resources or conditions, especially under sparse view conditions, and it is difficult to identify surface information of items with multilingual and non-fixed text locations.

Method used

Depth camera and color camera are used for multi-directional photography, and point cloud data fusion is combined with multi-view and cost-based methods to generate a high-quality three-dimensional reconstruction model. At the same time, multimodal large language model is used to identify and analyze text information on the surface of the item.

Benefits of technology

Generate high-quality three-dimensional reconstructions under the condition of fewer input images, improve the resolution of tiny text parts, realize accurate identification and analysis of item information in multilingual and text locations, and improve the efficiency of remote inspection of customs items.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032076A_ABST
    Figure CN120032076A_ABST
Patent Text Reader

Abstract

The invention discloses a remote inspection method and device for customs items, a medium and equipment. The method comprises the following steps: performing multidirectional photographing on a target object to be subjected to three-dimensional reconstruction to obtain point cloud image data and image data of the target object in multiple directions; removing an image background of the point cloud image data, obtaining an object image of the target object, and generating high-quality point cloud image data of the target object according to the pose of the depth camera relative to the target object and the object images in multiple directions; fusing the high-quality point cloud image data of the plurality of orientations to obtain a three-dimensional Gaussian initial parameter, and performing object reconstruction on the three-dimensional Gaussian initial parameter to obtain an initial three-dimensional reconstruction model of the target object; performing enhancement and refinement on the initial three-dimensional reconstruction model to obtain a three-dimensional reconstruction model of the target object; identifying character information of the target object on the image data; and checking the target article according to the character information of the target article and the three-dimensional reconstruction model, and determining a checking result of the target article.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of customs article inspection, and more specifically, to a method, device, medium and equipment for remote inspection of customs articles. Background Art

[0002] Customs has a need to inspect special items. These items are generally at the enterprise end, and inspection requires staff to visit the enterprise for inspection. A remote inspection method is needed. Currently, some remote inspections can only display a horizontal 360° display of a two-dimensional image. It is necessary to take multiple images of the item and remove the background, which is inefficient. There is an urgent need for an efficient and accurate method to use a camera to reconstruct the item in three dimensions, and to recognize the text information on the item for semantic analysis, and automatically identify the specifications and categories of the item to facilitate remote inspection by staff.

[0003] The difficulty of 3D reconstruction is that effective 3D reconstruction from sparse views is very complex and difficult, especially in practical application scenarios with limited resources or conditions. First, data capture is very cumbersome, and usually requires capturing a large number of multi-view images for effective 3D reconstruction, which is cumbersome and impractical for non-professionals. Second, when only a few images are available for reconstruction (for example, only 4 images within a 360° range), constructing a multi-view image is very complex and difficult. Figure 1 Consistency is very difficult. The 3D representation may overfit the input image, causing the result to degenerate into fragmented pixel blocks in the training view, lacking overall structure. In addition, for objects captured from sparse views within a 360° range, some parts may be omitted or severely compressed when observed due to extreme viewing angles. This omitted or compressed information is difficult to reconstruct in 3D from the input image alone. Even though a variety of methods have been proposed in recent years to reduce the reliance on dense capture, it remains a challenge to generate high-quality 3D objects when the views are extremely sparse.

[0004] In addition, the existing technology for identifying the surface information of an object is to first perform OCR text recognition, then perform semantic analysis or correspond key information based on the text position. However, it is difficult to analyze semantic information for multiple languages ​​and when the text position is not fixed, and it is difficult to identify the key information. This method uses a multimodal large language model to recognize text information and analyze and extract key information. Summary of the invention

[0005] In view of the deficiencies of the prior art, the present invention provides a method, device, medium and equipment for remote inspection of customs articles.

[0006] According to one aspect of the present invention, a method for remote inspection of customs articles is provided, comprising:

[0007] Use a depth camera and a color camera to take photos of the target object for 3D reconstruction from multiple directions to obtain point cloud image data and image data of the target object from multiple directions;

[0008] Remove the image background of the point cloud image data, obtain the object image of the target object, and generate high-quality point cloud image data of the target object according to the position of the depth camera relative to the target object and the object images in multiple orientations;

[0009] Use multi-view and cost-volume-based fusion of high-quality point cloud image data from multiple orientations to obtain the initial parameters of the three-dimensional Gaussian image, and use the Gaussian reconstruction model to reconstruct the object with the initial parameters of the three-dimensional Gaussian image to obtain the initial three-dimensional reconstruction model of the target object; use the video diffusion model to enhance and refine the initial three-dimensional reconstruction model to obtain the three-dimensional reconstruction model of the target object;

[0010] Use a multimodal large language model to recognize text information of target objects in image data;

[0011] The target object is inspected according to the text information of the target object and the three-dimensional reconstructed model to determine the inspection result of the target object.

[0012] Optionally, high-quality point cloud image data in multiple orientations are fused using multi-view and cost volume-based methods to obtain initial parameters of a three-dimensional Gaussian, including:

[0013] Use CNN and Transformer architecture to extract multi-view image features of high-quality point cloud image data in multiple orientations;

[0014] Calculate the cost volume of high-quality point cloud image data in multiple orientations based on multi-view image features:

[0015] The depth estimation algorithm is used to predict the depth of the cost volume and obtain the initial parameters of the three-dimensional Gaussian.

[0016] Optionally, a Gaussian reconstruction model is used to reconstruct an object using the three-dimensional Gaussian initial parameters to obtain an initial three-dimensional reconstruction model of the target object, including:

[0017] A Gaussian reconstruction model is used to learn and predict the initial parameters of the three-dimensional Gaussian to obtain the three-dimensional Gaussian parameters;

[0018] Based on the position and 3D Gaussian parameters of the target object, an initial 3D reconstruction model of the target object is constructed.

[0019] Optionally, the initial 3D reconstruction model is enhanced and refined using a video diffusion model to obtain a 3D reconstruction model of the target object, including:

[0020] The video diffusion model is used to denoise the initial 3D reconstruction model to obtain the 3D reconstruction model of the target object.

[0021] Optionally, the video diffusion model adopts the latent space loss:

[0022]

[0023] where g refers to the geometric backbone with trainable parameters θ, E is the frozen video diffusion model encoder, and I t is the real result image, F t is the latent space feature of the generated image.

[0024] Optionally, the multimodal large language model includes: a high compression rate encoder and a long context length decoder, and,

[0025] Use a multimodal large language model to recognize text information of target objects in image data, including:

[0026] A high compression rate encoder is used to convert the image data into tokens to obtain token data;

[0027] A long context length decoder is used to decode the token data and output the corresponding text information.

[0028] Optionally, the input image size of the high compression rate encoder is 1024×1024, and each input image will be compressed into a token of size 256×1024; the long context length decoder supports a token of maximum length 8K.

[0029] According to another aspect of the present invention, there is provided a device for remote inspection of customs articles, comprising:

[0030] An acquisition module is used to use a depth camera and a color camera to take multi-directional photos of the target object to be three-dimensionally reconstructed, and obtain point cloud image data and image data of the target object in multiple directions;

[0031] A generation module, used to remove the image background of the point cloud image data, obtain the object image of the target object, and generate high-quality point cloud image data of the target object according to the posture of the depth camera relative to the target object and the object images in multiple orientations;

[0032] The reconstruction module is used to fuse high-quality point cloud image data from multiple directions using multiple views and cost volumes to obtain initial parameters of the three-dimensional Gaussian image, and to reconstruct the object using the initial parameters of the three-dimensional Gaussian image using the Gaussian reconstruction model to obtain an initial three-dimensional reconstruction model of the target object; the video diffusion model is used to enhance and refine the initial three-dimensional reconstruction model to obtain a three-dimensional reconstruction model of the target object;

[0033] A recognition module, used to recognize text information of target objects in image data using a multimodal large language model;

[0034] The inspection module is used to inspect the target object according to the text information of the target object and the three-dimensional reconstruction model, and determine the inspection result of the target object.

[0035] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the storage medium stores a computer program, and the computer program is used to execute the method described in any one of the above aspects of the present invention.

[0036] According to another aspect of the present invention, an electronic device is provided, comprising: a processor; a memory for storing instructions executable by the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the method described in any one of the above aspects of the present invention.

[0037] Beneficial effects of the present invention:

[0038] 1. It provides a 3D reconstruction method for remote customs inspection. It combines an innovative method with large model prior knowledge to generate high-quality reconstruction with fewer input images. This technology can generate high-quality, realistic 360° new perspective images from very few input views, which is particularly suitable for remote inspection. Using the power of various large model priors to improve 3D reconstruction capabilities can achieve 1) robust initialization; 2) prevent overfitting; 3) retain details.

[0039] 2. In view of the small size and scattered distribution of text on special customs items, a multi-scale photo synthesis method is used to improve the resolution of tiny text parts, making the recognition of text parts clearer and more accurate.

[0040] 3. In response to the semantic recognition needs of customs special items, a multimodal large language model is used to recognize the text on the surface of the items, and analyze information such as category, brand, and specifications to automatically identify and compare item information, thereby improving inspection efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] A more complete understanding of exemplary embodiments of the present invention may be obtained by referring to the following drawings:

[0042] Figure 1 It is a flowchart of a method for remote inspection of customs articles provided by an exemplary embodiment of the present invention;

[0043] Figure 2 It is a schematic structural diagram of a device for remote inspection of customs articles provided by an exemplary embodiment of the present invention;

[0044] Figure 3This is a structure of an electronic device provided by an exemplary embodiment of the present invention. DETAILED DESCRIPTION

[0045] Below, the exemplary embodiments according to the present invention will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments of the present invention, and it should be understood that the present invention is not limited to the exemplary embodiments described here.

[0046] It should be noted that the relative arrangement of components and steps, the numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present invention unless specifically stated otherwise.

[0047] Those skilled in the art can understand that the terms "first" and "second" in the embodiments of the present invention are only used to distinguish different steps, devices or modules, etc., and neither represent any specific technical meaning nor indicate the necessary logical order between them.

[0048] It should also be understood that, in the embodiments of the present invention, “plurality” may refer to two or more than two, and “at least one” may refer to one, two or more than two.

[0049] It should also be understood that any component, data or structure mentioned in the embodiments of the present invention can generally be understood as one or more, unless explicitly limited or otherwise indicated in the context.

[0050] In addition, the term "and / or" in the present invention is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in the present invention generally indicates that the associated objects before and after are in an "or" relationship.

[0051] It should also be understood that the description of the various embodiments of the present invention focuses on the differences between the various embodiments, and the same or similar aspects thereof can be referenced to each other, and for the sake of brevity, they will not be described one by one.

[0052] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.

[0053] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the invention, its application, or uses.

[0054] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.

[0055] It should be noted that like reference numerals and letters refer to similar items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.

[0056] Embodiments of the present invention can be applied to electronic devices such as terminal devices, computer systems, servers, etc., which can operate with many other general or special computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, servers, etc. include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, small computer systems, large computer systems, and distributed cloud computing technology environments including any of the above systems, etc.

[0057] Electronic devices such as terminal devices, computer systems, servers, etc. can be described in the general context of computer system executable instructions (such as program modules) executed by computer systems. Generally, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in a distributed cloud computing environment, where tasks are performed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media including storage devices.

[0058] Exemplary Methods

[0059] Figure 1 FIG. 1 is a flow chart of a method for remote inspection of customs articles provided by an exemplary embodiment of the present invention. This embodiment can be applied to electronic devices, such as Figure 1 As shown, the method 100 for remote inspection of customs articles includes the following steps:

[0060] Step 101,

[0061] The 3D reconstruction system consists of six modules: camera module, point cloud generation module, multimodal 3D reconstruction module, iterative Gaussian optimization module, scene enhancement module, and text recognition module. The reconstruction process is as follows:

[0062] 1) Use the depth camera and RGB camera to take pictures of the object and remove the background.

[0063] 2) Perform feature extraction, construct cost volumes and depth estimation on multiple view photos, and use multi-view Transformer and cost volume-based encoder to match and fuse multi-view information.

[0064] 3) Use the Gaussian reconstruction model to reconstruct the rough geometric shape of the object and perform rough 3D reconstruction.

[0065] 4) Further use the pre-trained video diffusion model to refine the object reconstruction model to obtain realistic visual effects. This technology also enhances the ability of 360° scene synthesis by processing the interaction between views.

[0066] 5) Use a large multimodal model to identify textual information on the surface of objects, including categories and specifications, and analyze and extract key information.

[0067] In one embodiment of the present invention, a three-dimensional reconstruction process of an object;

[0068] Step 1: The present invention uses an image as input. Take photos of the object at no less than 4 angles. The object can be placed on a disc, the disc is rotated, a camera takes photos on the side of the disc, another camera takes photos on the top of the object, and the rotation angle of the object is detected. When it turns to 0°, 90°, 180°, and 270°, the box-shaped object is photographed respectively, and the bottle-shaped object is photographed at 0°, 60°, 120°, 180°, 240°, and 300°; and the top of the object is photographed, and the focus, brightness, white balance and other parameters of the image are analyzed during the photo shooting, and the focus, aperture, shutter, white balance and other parameters of the camera are dynamically adjusted. By calculating the focal length, a command is sent to drive the motor to rotate to adjust the focus of the SLR lens, or the focal length parameters are set by software.

[0069] Image background removal: Use a SLR camera to take photos of the items to be inspected on the top and side of the remote inspection table. The items to be inspected are placed on a rotating table and rotated. The SLR camera takes photos of the items at regular intervals to complete 360° photography of the items. An algorithm is required to remove the background of the items on the inspection table to obtain photos of individual items after background removal.

[0070] Image background removal method: obtain an image photo, convert it into a grayscale image, perform edge detection, extract the outline of the object, select n points evenly distributed within the object outline as point prompt words, input the segmentation model SAM (which can be a variant model such as SAM2, efficient-SAM, etc.), segment the object image, and thus remove the background image.

[0071] Step 2: (a) First, extract features from multiple view photos, construct cost volumes and depth estimates, and use multi-view Transformer and cost volume-based encoders (used to calculate parameter sets from a pair (or multiple pairs) of view images for 3D reconstruction) to match and fuse multi-view information. (b) Next, use a Gaussian reconstruction model to reconstruct the rough geometry of the object for rough 3D reconstruction. (c) Further use the pre-trained video diffusion model to refine the object reconstruction model to obtain realistic visual effects. The gradients returned from the diffusion model optimize the geometric backbone network to improve visual quality. This technology also enhances the ability of 360° scene synthesis by processing interactions between views.

[0072] (a) includes:

[0073] Perform feature extraction, cost volume construction, and depth estimation on multiple view photos.

[0074] 1. Feature extraction:

[0075] Given N multi-angle posed images, i.e., multi-view photos, as input, we first use CNN and Transformer architectures to extract multi-view image features. Specifically, we use a residual network to extract 4x downsampled per-view image features. Then, we use a multi-view Transformer with self-attention layers and cross-attention layers to exchange information between different views. The cross-attention layer is performed on each view relative to all other views, which has exactly the same learnable parameters as the two-view scenario. After this operation, the Transformer features of the cross-view are obtained, where C represents the channel dimension. N is the number of views, and H and W represent the image height and width, respectively.

[0076]

[0077] Given N multi-view images And their corresponding camera pose information By camera parameter K i , rotate R i and pan t i Matrix calculation. The goal is to learn the mapping f from the image to the parameters of a 3D Gaussian θ :

[0078] The 3D Gaussian parameters are

[0079] Among them, f θ Parameterized as a feed-forward network, θ is a learnable parameter optimized from a large-scale training dataset. Gaussian parameters are predicted in a pixel-aligned manner, including the position μ j , opacity α j, covariance∑ j and color C j (expressed as spherical harmonics), so the total number of 3D Gaussians for N input images of shape H×W is H×W×N.

[0080] 2. Construct the cost body;

[0081] The method for calculating the cost volume, taking the calculation of the i-th cost volume as an example, given the near and far depth range, the depth is evenly divided into L layers within the near depth and far depth ranges. First, L depth candidates are uniformly sampled in the inverse depth domain. Then use the camera projection matrix P i , P j And each depth candidate d m The feature f for view j j Transform to view i and obtain L transformed features:

[0082]

[0083] Where W represents the conversion operation. Then calculate F i and to obtain the correlation coefficient

[0084]

[0085] When there are more than two views as input, similarly transform the features of another view to view i and calculate their correlation coefficients Finally, all correlation coefficients are pixel-averaged, enabling the model to accept an arbitrary number of views as input.

[0086] Collect all the correlation coefficients to obtain the cost volume of view i:

[0087]

[0088] In this way, we obtain N cost volumes for K input views.

[0089] 3. Depth Estimation

[0090] The cost volume models the cross-view feature matching information of different depth candidates via a plane sweeping stereo approach. K cost volumes are constructed for K input views to predict N depth maps.

[0091] Finally, the estimated depth d for each view is obtained by applying a softmax operation to the cost volume in the depth dimension. After that, the Gaussian mean is calculated as μ = K -1 ud+Δ, where K is the internal parameter value of the camera, u=(u x ,uy , 1) represents each pixel, Δ∈R 3 is the predicted offset, opacity a∈[0, 1], covariance Σ∈R 3×3 ,color

[0092] where S is the order of the spherical harmonics. Once the model predicts a set of three-dimensional Gaussian parameters You can render the target view I by rasterizing it t .

[0093] (b) includes:

[0094] The Gaussian reconstruction model is used to reconstruct the rough geometric shape of the object and perform a rough 3D reconstruction.

[0095] Initial 3D reconstruction model

[0096] The specific processing process of the Gaussian reconstruction model is as follows:

[0097] The model predicts additional Gaussian features that can be directly rendered into the latent space, providing pose and visual cues for subsequent diffusion models.

[0098] The model learns to predict the three-dimensional Gaussian parameters and then uses the pose P of the target camera t Stitching to get a set of RGB images I t To ensure better integration with the subsequent diffusion module, the Gaussian feature f is calculated along with other parameters. i , which can be rasterized into the corresponding latent space feature F t In addition, the view selection strategy is improved to increase the robustness of the model in handling view inputs with a wide range of displacement angles.

[0099] The Gaussian reconstruction model selectively freezes all parameters in the network except the LoRA weights to fine-tune the control network. This process injects scene features into the model and enhances the model's ability to correct blurry scene images. The Lora parameters are integrated into the text conditional encoder and the UNet of ControlNet. The Lora level is set to 16 and integrated in each converter block, linear layer, and convolutional layer. By training the image pairs in the coarse stage and finally using a multi-frame diffusion model, namely video diffusion, to refine the image reconstructed in the coarse stage, the model can generate detailed and realistic images from blurry Gaussian rendered images.

[0100] Specifically, N images are generated from a virtual view using a Gaussian kernel. Subsequently, these images are uniformly enhanced using a Gaussian reconstruction model to generate a set of inpainted images. In order to maintain the coherence of the scene and reduce potential conflicts, the denoising intensity is set low, and limited details are gradually reintroduced into the Gaussian rendered images during each inpainting process. These inpainted virtual view images, together with the monocular depth map and normal map from the training view, will be used to adjust the Gaussian optimization process.

[0101] Since the input view angles are widely spaced, resulting in minimal overlap between specific view pairs, we use a cross-view attention mechanism to mitigate this problem and build the cost volume only within a local group of input views based on the camera position, thereby reducing memory consumption and ensuring stable model convergence.

[0102] (c) includes: using a pre-trained video diffusion model on the initial 3D reconstruction model in (b) to refine the object reconstruction model to obtain the final enhanced model.

[0103] The model effectively utilizes a pre-trained stable video diffusion model to render features directly into the latent space. These features serve as pose and visual cues to guide the denoising process and produce realistic and 3D consistent new views. This method not only improves the synthesis quality, but also reduces the dependence on a large amount of image data, making it possible to achieve high-quality new view synthesis with a limited number of images.

[0104] The diffusion model exploits these cues to refine the coarse reconstruction, generating visually new perspectives that are consistent across multiple views and geometrically accurate.

[0105] The gradients passed back from the diffusion model optimize the geometric backbone network to improve visual quality. In addition, this technology enhances the ability of 360° scene synthesis by selecting the correct camera viewpoint and handling the interaction between views. This technology can generate high-quality, realistic 360° new perspective images from a small number of input views.

[0106] Video Diffusion Model: A multi-frame diffusion model is used to refine the above coarsely reconstructed visual appearance. The diffusion model is pre-trained on a large-scale video dataset, has strong temporal consistency prior knowledge, and performs denoising in the latent space. In particular, given a target sequence x containing M images 1:M , is embedded into the latent space by the encoder E which is initially frozen, In a Markov process:

[0107]

[0108] Add Gaussian noise ε∈N(0,I);

[0109] while α 0 ,..., α t are predefined noises within T steps. Given the noise input then the denoising function ε θ is trained by optimizing the following objective function:

[0110]

[0111] where x 1:M is the target image and y is the conditional input. After training ∈ θ the model can generate videos by iteratively denoising pure Gaussian values conditioned on y .

[0112] The final loss function is calculated using the mean squared error (MSE) between the ground truth and its prediction in the latent space:

[0113] where is obtained by transforming the velocity into the latent space, i.e.:

[0114]

[0115] Multi-view fusion: To ensure an accurate understanding of the scene, the model needs to integrate low-level information (such as depth and texture) and high-level information (such as semantics and geometry). A hybrid conditioning mechanism is adopted to fine-tune the diffusion model with multi-angle view inputs. The CLIP image embedding token information of the original visible view I is used as the text prompt. On each UNet block, a cross-attention module operation is applied to capture the high-level semantic information of the input image. Due to the multi-angle views, these tokens are averaged into a global token. In another process, the spatial conditions from the coarse geometric reconstruction features are concatenated with the latent space noise along the channels. These spatial condition features help the model capture view information and learn the texture information of the object.

[0116] Since the video diffusion model is conditioned on image encoding features and the diffusion module is conditioned on 3D Gaussian rendering features, a latent space alignment loss is used:

[0117]

[0118] to align these two spaces, where g refers to the geometric backbone with trainable parameters θ, E is the frozen diffusion model encoder, I t is the ground truth image, and F t is the latent space feature of the generated image.

[0119] In one embodiment of the present invention, a multimodal large language model is used to recognize text information on an object;

[0120] Text recognition: The object is placed on a rotating table, and the camera can be raised and lowered. There are adjustable light sources around the object. Cameras are placed directly above and to the side of the object. The camera directly above is used to take photos of the top surface of the object and information such as the customs declaration. At the same time, the camera directly above detects the angle of the object in real time when the object rotates. When the side camera takes pictures, the text detector is used to detect the direction and position of the text. When the direction is correct and there are a lot of texts in this view, the side camera is notified to take pictures when the angle of the box-shaped object is close to facing the side camera. For objects with small text, the focal length will be automatically adjusted to pull into the lens, enlarge the position of the text area, and take pictures of the object at multiple scales. Select photos of the object at a suitable angle, perform edge detection on the object image, and measure the edges of the object box in the horizontal and vertical directions in the picture, that is, when the long and wide sides of the object are facing the camera at the angle of rotation, select images at these two angles for text recognition, which can reduce the deformation of the picture and make the text in the picture clearest.

[0121] Use algorithms to recognize text on each side of the packaging of the items to be inspected and on the outer packaging of bottles. Currently, text recognition in multiple languages ​​is supported.

[0122] For boxed objects, photos of four sides and the top are required. For bottle-shaped objects, photos at six angles (0°, 60°, 120°, 180°, 240°, and 300°) are selected and transmitted to the large language model for text recognition and key information extraction.

[0123] Structurally, an encoder-decoder model is adopted. A high compression rate encoder is used to convert optical images into tokens, and a long context length decoder is used to output the corresponding OCR results. The encoder input size is 1024×1024. Each input image will be compressed into a token of 256×1024 size. The decoder supports tokens with a maximum length of 8K to ensure that it can handle long text scenarios. Model training can be divided into three steps, namely, decoupled pre-training of the encoder, joint training of the encoder and the new decoder, and further post-training of the decoder.

[0124] The large language recognition model consists of three modules, namely the image encoder, the linear layer, and the output decoder. The linear layer is a connector between the visual encoder and the language decoder that maps the channel dimension. Three main steps are used to optimize the entire model. First, a plain text recognition task is performed to pre-train the visual encoder. In order to improve training efficiency and save GPU resources, a tiny decoder is selected to pass the gradient to the encoder. In this stage, images containing scene text and manual images containing document-level characters are input into the model to let the encoder collect the encoding capabilities of the two most commonly used characters. In the next stage, the trained visual encoder is connected to the new large decoder to form the model architecture.

[0125] Model training phase 1: Pre-train the visual encoder to effectively adapt it to the OCR task. Phase 2: Build the model by connecting the visual encoder to the multimodal vision model, and use enough customs item classification knowledge to handle more general optical characters at this stage. Phase 3: Customize the model based on the new character recognition function without modifying the visual encoder.

[0126] After the pre-training step of the visual encoder is completed, it is connected to the large language model to build the final architecture of the model. Here, the multimodal large model is used as the decoder because it has a relatively small number of parameters while incorporating prior knowledge of multiple languages. The connector (i.e., the linear embedding layer) is resized to 1024×1024 to be consistent with the input channels of the multimodal model. The model adopts a seamless encoder-decoder paradigm with a total number of parameters of approximately 580M and low computing power requirements. The high compression rate of the encoder (1024×1024 optical pixels to 256 image tokens) saves a lot of token space for the decoder to generate new tokens. At the same time, the decoder's decoding context length (the maximum length used is approximately 8K) ensures that the model can effectively output OCR results in tiny text scenarios.

[0127] Visual encoder generation. The chosen encoder architecture is VitDet, as its local attention can significantly reduce the computational cost for high-resolution images. The last two layers of the encoder are designed following the Vary-tiny setting, converting the 1024×1024×3 input image into 256×1024 image tokens. These image tokens are then projected to the language model (OPT-125M) dimension through a 1024×768 linear layer. In the preprocessing stage, images of each shape are directly resized to a 1024×1024 square, as a square can accommodate images of various aspect ratios in a compromise.

[0128] Merge each text content in order from top to bottom and from left to right. 2) Crop the text area in the original image according to the bounding box and save it as an image slice.

[0129] Multi-crop data engine for very large image OCR. The model supports 1024×1024 input resolution, which is sufficient for common OCR tasks. However, for some scenarios with large images and small text, dynamic resolution is required. Using a high compression rate encoder, dynamic resolution can be achieved under a large sliding window (1024×1024), ensuring that the model can complete OCR tasks with extremely high resolution and tiny text.

[0130] In one embodiment of the present invention, the article inspection process;

[0131] Extract key information of items in images: First, perform text recognition on the input item photos, and then use the text recognition results to guide the large language model to extract key information from the text on the item packaging to avoid typos or incorrect spellings. Key information such as: item category, Chinese name, foreign name, specification, import and export country, ingredients, validity period, unit of use, manufacturer, manufacturer, transportation and storage conditions. Support multi-language reading and extraction of key information. For the key information returned by multiple images, the key includes the following: item category, Chinese name, foreign name, specification, import and export country, ingredients, validity period, unit of use, manufacturer, manufacturer, transportation and storage; the value is the information corresponding to the key. The method of merging information from multiple photos is as follows: "For the same key, compare the value to see if the value is exactly the same, whether it contains a relationship, and then use the information comparison interface to judge the text similarity, and try to retain the keys that are different."

[0132] First, perform text recognition on photos of items from multiple angles and photos of customs declarations, then extract key information respectively, and finally compare the similarity of each key information extracted from the items and customs declarations. If the similarity is higher than the threshold, the information is considered consistent, otherwise the user is prompted with inconsistent relevant information. Assist users to compare whether the item and customs declaration information are consistent.

[0133] 6. The key technical points of the present invention are:

[0134] 1. Use 3D reconstruction methods to provide an efficient remote inspection method for customs inspection.

[0135] 2. The multi-scale photo synthesis method is used to improve the resolution of tiny text parts.

[0136] 3. Use a multimodal large language model to recognize text on the surface of objects and analyze and extract information such as item category, brand, and specifications.

[0137] Beneficial effects of the present invention:

[0138] 1. It provides a 3D reconstruction method for remote customs inspection. It combines an innovative method with large model prior knowledge to generate high-quality reconstruction with fewer input images. This technology can generate high-quality, realistic 360° new perspective images from very few input views, which is particularly suitable for remote inspection. Using the power of various large model priors to improve 3D reconstruction capabilities can achieve 1) robust initialization; 2) prevent overfitting; 3) retain details.

[0139] 2. In view of the small size of the text on the special customs items, a multi-scale photo synthesis method is used to improve the resolution of the tiny text part, making the recognition of the text part clearer and more accurate.

[0140] 3. In response to the semantic recognition needs of customs special items, a multimodal large language model is used to recognize the text on the surface of the items, and analyze information such as category, brand, and specifications to automatically identify and compare item information, thereby improving office efficiency.

[0141] Exemplary Devices

[0142] Figure 2 FIG. 1 is a schematic diagram of a device for remote inspection of customs articles provided by an exemplary embodiment of the present invention. Figure 2 As shown, the device 200 includes:

[0143] An acquisition module 210 is used to use a depth camera and a color camera to take photos of the target object to be 3D reconstructed in multiple directions, and obtain point cloud image data and image data of the target object in multiple directions;

[0144] A generating module 220, configured to remove the image background of the point cloud image data, obtain an object image of the target object, and generate high-quality point cloud image data of the target object according to the position of the depth camera relative to the target object and the object images in multiple orientations;

[0145] The reconstruction module 230 is used to fuse high-quality point cloud image data of multiple orientations using multiple views and cost volumes to obtain initial parameters of a three-dimensional Gaussian image, and to reconstruct an object using the initial parameters of the three-dimensional Gaussian image using a Gaussian reconstruction model to obtain an initial three-dimensional reconstruction model of the target object; and to enhance and refine the initial three-dimensional reconstruction model using a video diffusion model to obtain a three-dimensional reconstruction model of the target object.

[0146] A recognition module 240, for recognizing text information of a target object on the image data using a multimodal large language model;

[0147] The inspection module 250 is used to inspect the target object according to the text information of the target object and the three-dimensional reconstructed model, and determine the inspection result of the target object.

[0148] Optionally, the reconstruction module 230 uses multi-view and cost volume-based fusion of high-quality point cloud image data in multiple orientations to obtain three-dimensional Gaussian initial parameters, including:

[0149] Use CNN and Transformer architecture to extract multi-view image features of high-quality point cloud image data in multiple orientations;

[0150] Calculate the cost volume of high-quality point cloud image data in multiple orientations based on multi-view image features:

[0151] The depth estimation algorithm is used to predict the depth of the cost volume and obtain the initial parameters of the three-dimensional Gaussian.

[0152] Optionally, the reconstruction module 230 uses a Gaussian reconstruction model to reconstruct the object using the three-dimensional Gaussian initial parameters to obtain an initial three-dimensional reconstruction model of the target object, including:

[0153] A Gaussian reconstruction model is used to learn and predict the initial parameters of the three-dimensional Gaussian to obtain the three-dimensional Gaussian parameters;

[0154] Based on the pose and 3D Gaussian parameters of the target object, an initial 3D reconstruction model of the target object is constructed.

[0155] Optionally, the reconstruction module 230 uses a video diffusion model to enhance and refine the initial 3D reconstruction model to obtain a 3D reconstruction model of the target object, including:

[0156] The video diffusion model is used to denoise the initial 3D reconstruction model to obtain the 3D reconstruction model of the target object.

[0157] Optionally, the video diffusion model adopts the latent space loss:

[0158]

[0159] Where g refers to the geometric backbone with trainable parameters θ, E is the frozen diffusion model encoder, and I t is the real result image, F t is the latent space feature of the generated image.

[0160] Optionally, the multimodal large language model includes: a high compression rate encoder and a long context length decoder, and,

[0161] Use a multimodal large language model to recognize text information of target objects in image data, including:

[0162] A high compression rate encoder is used to convert the image data into tokens to obtain token data;

[0163] A long context length decoder is used to decode the token data and output the corresponding text information.

[0164] Optionally, the input image size of the high compression rate encoder is 1024×1024, and each input image will be compressed into a token of size 256×1024; the long context length decoder supports a token of maximum length 8K.

[0165] Exemplary Electronic Devices

[0166] Figure 3 This is a structure of an electronic device provided by an exemplary embodiment of the present invention. Figure 3 As shown, the electronic device 30 includes one or more processors 31 and a memory 32 .

[0167] The processor 31 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.

[0168] The memory 32 may include one or more computer program products, and the computer program product may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 31 may run the program instructions to implement the methods of the software programs of the various embodiments of the present invention described above and / or other desired functions. In one example, the electronic device may also include: an input device 33 and an output device 34, which are interconnected via a bus system and / or other forms of connection mechanisms (not shown).

[0169] In addition, the input device 33 may also include, for example, a keyboard, a mouse, and the like.

[0170] The output device 34 can output various information to the outside. The output device 34 can include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto.

[0171] Of course, to simplify, Figure 3 Only some of the components related to the present invention in the electronic device are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, the electronic device may further include any other appropriate components according to specific application conditions.

[0172] Exemplary computer program products and computer-readable storage media

[0173] In addition to the above-mentioned methods and devices, an embodiment of the present invention may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the method according to various embodiments of the present invention described in the above-mentioned "Exemplary Method" section of this specification.

[0174] The computer program product may be written in any combination of one or more programming languages ​​to write program code for performing the operations of the embodiments of the present invention, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0175] In addition, an embodiment of the present invention may also be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, enable the processor to execute the steps of the method according to various embodiments of the present invention described in the above “Exemplary Method” section of this specification.

[0176] The computer readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can include, for example, but is not limited to, a system, system or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0177] The basic principle of the present invention is described above in conjunction with specific embodiments. However, it should be pointed out that the advantages, strengths, effects, etc. mentioned in the present invention are only examples and not limitations, and it cannot be considered that these advantages, strengths, effects, etc. must be possessed by each embodiment of the present invention. In addition, the specific details disclosed above are only for the purpose of illustration and facilitation of understanding, rather than limitation, and the above details do not limit the present invention to being implemented by adopting the above specific details.

[0178] Each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system embodiment, since it basically corresponds to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0179] The block diagrams of the devices, systems, equipment, and systems involved in the present invention are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagram. As will be appreciated by those skilled in the art, these devices, systems, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open words, referring to "including but not limited to", and can be used interchangeably with them. The words "or" and "and" used here refer to the words "and / or" and can be used interchangeably with them, unless the context clearly indicates otherwise. The word "such as" used here refers to the phrase "such as but not limited to", and can be used interchangeably with it.

[0180] The method and system of the present invention may be implemented in many ways. For example, the method and system of the present invention may be implemented by software, hardware, firmware or any combination of software, hardware, firmware. The above order of steps for the method is only for illustration, and the steps of the method of the present invention are not limited to the order specifically described above, unless otherwise specifically stated. In addition, in some embodiments, the present invention may also be implemented as a program recorded in a recording medium, which includes machine-readable instructions for implementing the method according to the present invention. Thus, the present invention also covers a recording medium storing a program for executing the method according to the present invention.

[0181] It should also be noted that in the system, device and method of the present invention, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent schemes of the present invention. The above description of the disclosed aspects is provided to enable any technician in the field to make or use the present invention. Various modifications to these aspects are very obvious to those skilled in the art, and the general principles defined here can be applied to other aspects without departing from the scope of the present invention. Therefore, the present invention is not intended to be limited to the aspects shown here, but in accordance with the widest range consistent with the principles and novel features disclosed here.

[0182] The above description has been given for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present invention to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions and sub-combinations thereof.

Claims

1. A method for remote inspection of customs articles, characterized in that: include: Using a depth camera and a color camera to take photos of the target object to be three-dimensionally reconstructed in multiple directions to obtain point cloud image data and image data of the target object in multiple directions; Removing the image background of the point cloud image data, acquiring an object image of the target object, and generating high-quality point cloud image data of the target object according to the position of the depth camera relative to the target object and the object images at multiple orientations; Using multi-view and cost volume-based fusion of the high-quality point cloud image data in multiple directions to obtain three-dimensional Gaussian initial parameters, and using a Gaussian reconstruction model to reconstruct the object with the three-dimensional Gaussian initial parameters to obtain an initial three-dimensional reconstruction model of the target object; using a video diffusion model to enhance and refine the initial three-dimensional reconstruction model to obtain a three-dimensional reconstruction model of the target object; Using a multimodal large language model to recognize text information of the target object on the image data; The target object is inspected according to the text information of the target object and the three-dimensional reconstructed model to determine the inspection result of the target object.

2. The method according to claim 1, characterized in that The high-quality point cloud image data in multiple directions are fused using multi-view and cost volume-based methods to obtain initial parameters of a three-dimensional Gaussian, including: Extract multi-view image features of the high-quality point cloud image data in multiple orientations using CNN and Transformer architectures; Calculate the cost volume of the high-quality point cloud image data in multiple orientations based on the multi-view image features: A depth estimation algorithm is used to perform depth prediction on the cost volume to obtain initial parameters of the three-dimensional Gaussian.

3. The method according to claim 1, characterized in that Reconstructing the object using the three-dimensional Gaussian initial parameters using a Gaussian reconstruction model to obtain an initial three-dimensional reconstruction model of the target object, including: Using the Gaussian reconstruction model to learn and predict the three-dimensional Gaussian initial parameters to obtain three-dimensional Gaussian parameters; The initial three-dimensional reconstruction model of the target object is constructed based on the position and posture of the target object and the three-dimensional Gaussian parameters.

4. The method according to claim 1, characterized in that: The initial three-dimensional reconstruction model is enhanced and refined by using a video diffusion model to obtain a three-dimensional reconstruction model of the target object, including: The initial three-dimensional reconstruction model is denoised using a video diffusion model to obtain a three-dimensional reconstruction model of the target object.

5. The method according to claim 4, characterized in that The video diffusion model adopts latent space to its loss: Where g refers to the geometric backbone with trainable parameters θ, E is the frozen diffusion model encoder, and I t is the real result image, F t is the latent space feature of the generated image.

6. The method according to claim 1, characterized in that The multimodal large language model includes: a high compression rate encoder and a long context length decoder, and, Using a multimodal large language model to recognize text information of the target object on the image data includes: Using the high compression rate encoder to convert the image data into tokens to obtain token data; The long context length decoder is used to decode the token data and output corresponding text information.

7. The method according to claim 6, characterized in that The input image size of the high compression rate encoder is 1024×1024, and each input image will be compressed into a token of 256×1024 size; the long context length decoder supports tokens with a maximum length of 8K.

8. A device for remote inspection of customs articles, characterized in that: include: An acquisition module, used to use a depth camera and a color camera to take multi-directional photos of a target object to be three-dimensionally reconstructed, and obtain point cloud image data and image data of the target object in multiple directions; A generating module, configured to remove the image background of the point cloud image data, obtain an object image of a target object, and generate high-quality point cloud image data of the target object according to the position of the depth camera relative to the target object and the object images at multiple orientations; A reconstruction module is used to fuse the high-quality point cloud image data of multiple orientations using multiple views and cost volumes to obtain three-dimensional Gaussian initial parameters, and use a Gaussian reconstruction model to reconstruct the object with the three-dimensional Gaussian initial parameters to obtain an initial three-dimensional reconstruction model of the target object; and use a video diffusion model to enhance and refine the initial three-dimensional reconstruction model to obtain a three-dimensional reconstruction model of the target object; A recognition module, used for recognizing text information of the target object on the image data by using a multimodal large language model; The inspection module is used to inspect the target object according to the text information of the target object and the three-dimensional reconstructed model, and determine the inspection result of the target object.

9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and the computer program is used to execute the method according to any one of claims 1 to 7.

10. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing instructions executable by the processor; The processor is used to read the executable instructions from the memory and execute the instructions to implement the method described in any one of claims 1 to 7.

Citation Information

Cited By

  • Hyperspectral image segmentation method based on fusion point prompt and Markov diffusion

    CN121415064A

  • Hyperspectral image segmentation method based on fusion point prompt and markov diffusion

    CN121415064B