Three-dimensional target object identification method and device, electronic equipment and medium
By fine-tuning and domain migration processing of the three-dimensional target object recognition model, the difficulty of identifying three-dimensional data under different conditions is solved, the applicability and robustness of the model is improved, and the deployment and migration costs are reduced.
Patent Information
- Application Number
- CN202510577930.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-03
AI Technical Summary
The existing three-dimensional target recognition technology results in inconsistent three-dimensional data quality, insufficient sample number and large differences in data distribution when equipment parameters, imaging methods and environmental conditions change, making it difficult to achieve effective model migration and identification.
By obtaining the target object data to be identified and inputting it into a multi-task processing model trained on the source domain data, fine-tuning processing is performed to obtain candidate area information. Then, multiple domain migration submodules are acquired based on the multitasking structure, these submodules are trained to update candidate area information, and finally fine-tune the multitasking model to adapt to the new environment.
It improves the applicability and robustness of the three-dimensional recognition model in new environments, new devices or new data sources, reduces the burden of manual retraining, and realizes the migration of three-dimensional recognition capabilities under low-label or no-label conditions, reducing the cost of model deployment and migration.
Smart Images

Figure CN120088770A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and security inspection, and more particularly to a three-dimensional target recognition method, device, equipment, medium and program product. Background Art
[0002] In three-dimensional target recognition tasks, three-dimensional structure data of a target is obtained by means of CT (Computed Tomography) imaging, lidar scanning, volume reconstruction, etc., and combined with a deep neural network for training to achieve automatic recognition of three-dimensional targets. Currently, most mainstream technologies are based on a supervised learning framework and rely on a large amount of labeled three-dimensional data for model training.
[0003] However, in practical applications, due to differences in device parameters, imaging methods, environmental conditions, etc., the three-dimensional data obtained under different acquisition conditions has problems such as inconsistent quality, insufficient sample quantity, and large differences in data distribution. In addition, the three-dimensional recognition task itself has a high degree of complexity, making the process of fine-tuning and optimizing the model more difficult. The current common practice is to re-collect a large amount of labeled data for a new scenario, train from scratch or adjust an existing model on a large scale, but in practical applications, it is often difficult to obtain the required data volume and annotation quality in a timely manner. Summary of the Invention
[0004] In view of the above problems, the present invention provides a three-dimensional target recognition method, device, equipment, medium and program product.
[0005] A first aspect of the present invention provides a three-dimensional target recognition method, including: obtaining target object data to be recognized; inputting the target object data into a target multi-task processing model, processing the target object data based on the target multi-task processing model to obtain three-dimensional recognition information of the target object; and obtaining a recognition result of the target object based on the three-dimensional recognition information, where the target multi-task processing model is a model obtained by fine-tuning a multi-task processing model trained on source domain data, and the fine-tuning processing includes: processing the source domain data and the target domain data based on the multi-task processing model to obtain candidate region information; obtaining a plurality of domain transfer sub-modules based on a multi-task structure in the multi-task processing model; training at least the plurality of domain transfer sub-modules based on the source domain data, the target domain data, and the candidate region information to obtain updated candidate region information; and fine-tuning the multi-task processing model based on at least the updated candidate region information.
[0006] According to an embodiment of the present invention, obtaining a plurality of domain migration sub-modules based on the multi-task structure in the multi-task processing model specifically includes: decoupling the multi-task structure in the multi-task processing model to obtain a plurality of task sub-paths; constructing a plurality of sub-network structures based on the plurality of task sub-paths; and embedding a domain discriminator in the plurality of sub-network structures to obtain the plurality of domain migration sub-modules.
[0007] According to an embodiment of the present invention, the plurality of domain migration sub-modules include a gradient reversal layer connected to the domain discriminator, and the gradient reversal layer is used to reverse the gradient direction of the domain discriminator during training.
[0008] According to an embodiment of the present invention, the target object data includes three-dimensional image data of the target object obtained by the security inspection device, and the recognition result is used to recognize the category information and position information of the target object.
[0009] According to an embodiment of the present invention, the target multi-task processing model is at least used to perform target detection on the target object.
[0010] According to an embodiment of the present invention, processing the source domain data and the target domain data based on the multi-task processing model to obtain candidate region information specifically includes: respectively inputting the source domain data and the target domain data into the multi-task processing model to obtain corresponding preliminary detection results; and at least extracting a category prediction score and three-dimensional bounding box parameters from the preliminary detection results as the candidate region information.
[0011] According to an embodiment of the present invention, the plurality of domain migration sub-modules at least include a classification migration sub-module and a regression migration sub-module. The classification migration sub-module is used to predict the category information of the target object, and the regression migration sub-module is used to predict the position information of the target object.
[0012] According to an embodiment of the present invention, the input of the regression migration sub-module at least includes the category prediction result output by the classification migration sub-module.
[0013] According to an embodiment of the present invention, the candidate region information includes source domain positive sample candidate region information, source domain negative sample candidate region information, target domain positive sample candidate region information, and target domain negative sample candidate region information.
[0014] According to an embodiment of the present invention, training a plurality of the domain transfer sub-modules based at least on the source domain data, the target domain data, and the candidate region information specifically includes: constructing paired training samples composed of positive samples and negative samples corresponding to the source domain data and the target domain data, where the paired training samples are associated with the candidate region information; inputting the paired training samples into the corresponding domain transfer sub-module to obtain a target prediction result; and updating the candidate region information based on the target prediction result to generate the updated candidate region information.
[0015] A second aspect of the present invention provides a three-dimensional object recognition device, including: a data acquisition module, configured to: acquire object data to be recognized; a processing module, configured to: input the object data into a target multi-task processing model, and process the object data based on the target multi-task processing model to obtain three-dimensional recognition information of the object, where the target multi-task processing model is a model obtained by fine-tuning a multi-task processing model trained on source domain data, and the fine-tuning processing includes: processing the source domain data and the target domain data based on the multi-task processing model to obtain candidate region information; obtaining a plurality of domain transfer sub-modules based on the multi-task structure in the multi-task processing model, training a plurality of the domain transfer sub-modules based at least on the source domain data, the target domain data, and the candidate region information to obtain updated candidate region information; fine-tuning the multi-task processing model based at least on the updated candidate region information; and a recognition result acquisition module, configured to: obtain the recognition result of the object based on the three-dimensional recognition information.
[0016] A third aspect of the present invention provides an electronic device, including: one or more processors; a memory for storing one or more computer programs, where the above one or more processors execute the above one or more computer programs to implement the steps of the above method.
[0017] A fourth aspect of the present invention further provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.
[0018] A fifth aspect of the present invention further provides a computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.
[0019] According to an embodiment of the present invention, through multitasking, multi-path training, and domain transfer mechanisms, the transfer sub-module tunes the task branches, enabling the model to automatically adapt to different data distributions, thereby improving the applicability and robustness of the 3D recognition model in new environments, new devices, or new data sources, and significantly reducing the manual retraining burden. At the same time, by using source domain data and candidate region information, the transfer of 3D recognition capabilities under low-annotation or no-annotation conditions can be achieved, significantly reducing the model deployment and transfer costs, so that under the premise of maintaining the same model structure, more refined cross-task transfer tuning can be realized, and the recognition accuracy of classification and regression tasks can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Through the following description of the embodiments of the present invention with reference to the accompanying drawings, the above content and other objects, features, and advantages of the present invention will become clearer. In the drawings:
[0021] Figure 1 Schematically shows an application scenario diagram of a 3D object recognition method, device, equipment, medium, and program product according to an embodiment of the present invention;
[0022] Figure 2 Schematically shows a flowchart of a 3D object recognition method according to an embodiment of the present invention;
[0023] Figure 3 Schematically shows a model architecture diagram of a 3D object recognition system according to an embodiment of the present invention;
[0024] Figure 4A Schematically shows a structure diagram of a domain transfer sub-module for a classification task according to an embodiment of the present invention;
[0025] Figure 4B Schematically shows a structure diagram of a domain transfer sub-module for a regression task according to an embodiment of the present invention;
[0026] Figure 5 Schematically shows a block diagram of a 3D object recognition device according to an embodiment of the present invention; and
[0027] Figure 6 Schematically shows a block diagram of an electronic device suitable for implementing a 3D object recognition method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, numerous specific details are set forth to provide a comprehensive understanding of the embodiments of the present invention. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.
[0029] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "comprising", "including" and the like used herein indicate the presence of the described features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.
[0030] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0031] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but not be limited to a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0032] First, the technical terms described herein are explained and illustrated as follows.
[0033] The source domain refers to the data distribution environment corresponding to the existing, fully annotated training data, which is used to train the base model or initialize the model parameters. The source domain data usually comes from prior scenarios, standard devices, or controlled experimental conditions, and has large-scale, high-quality annotation information, which can be used as the basis for model learning.
[0034] The target domain refers to the actual data scenario where the model needs to be applied, and its data distribution has different data feature distributions from the source domain.
[0035] An adversarial network is a deep learning method that includes a mutual game structure between two sub-networks, usually consisting of a generator and a discriminator.
[0036] In the current scenario of three-dimensional object analysis and processing, to improve efficiency and intelligence level, computed tomography (CT) devices are widely used in the automatic inspection and analysis of various objects. Through X-ray scanning and image reconstruction technology, CT devices can generate tomographic images containing three-dimensional physical information such as density and material, providing basic data support for various tasks. To achieve accurate analysis and processing of complex three-dimensional structure objects, existing technologies usually adopt deep learning methods, training multi-task processing models through a large number of labeled data, extracting features from CT images and completing tasks such as object localization, classification, segmentation, tracking, and quality assessment.
[0037] However, the applicant has found through research that in the actual application process, there are significant differences in the image acquisition environment of objects. Different models of devices, changes in scanning parameters, differences in imaging algorithms, and different types of specific detection items faced in the application scenario will all lead to significant differences in the data distribution, density characteristics, noise level, etc. of the collected CT images. This data distribution difference will seriously affect the adaptability of the original model in new devices or new scenarios.
[0038] Generally, to meet new recognition requirements, it is necessary to collect a large amount of target data that conforms to the new distribution and perform fine annotation. However, in some fields, there are multiple difficulties in obtaining training samples. On the one hand, the appearance frequency of some specific objects is low, resulting in scarce relevant samples; on the other hand, some items have restrictions such as privacy and control, resulting in high annotation difficulty, long time consumption, and high cost, making it difficult to support the construction of a large-scale labeled dataset. Therefore, the target data not only has a small sample size but also has problems such as class imbalance and imbalance in the proportion of positive and negative samples, making it difficult to support effective training based on deep neural networks.
[0039] The applicant also noticed that existing multi-task processing models are often designed with complex task structures to complete multiple tasks simultaneously (such as object classification and position prediction, three-dimensional segmentation, object tracking, quality detection, etc.). Compared with traditional single-task models, multi-task models have a mechanism of feature coupling and parameter sharing during the training process, which makes their adaptation requirements for different task branches higher and the migration process more complex in cross-domain applications. Most existing transfer learning methods are oriented to single tasks and are difficult to fully adapt to the coordinated adjustment of model parameters and feature spaces under multi-task structures, resulting in unstable training effects and low transfer efficiency.
[0040] Specifically, when directly applying the model trained based on source domain data to the target domain, it often faces the "domain shift" problem, that is, the model that performs well in the source domain has a significant decrease in recognition accuracy in the target domain, which not only affects the detection accuracy of objects but also restricts the wide application of the model in different fields and different tasks.
[0041] Based on this, an embodiment of the present invention provides a three-dimensional object recognition method, including: obtaining object data to be recognized; inputting the object data into a target multi-task processing model, processing the object data based on the target multi-task processing model to obtain three-dimensional recognition information of the object; and obtaining a recognition result of the object based on the three-dimensional recognition information, where the target multi-task processing model is a model obtained by fine-tuning a multi-task processing model trained on source domain data, and the fine-tuning process includes: processing the source domain data and the target domain data based on the multi-task processing model to obtain candidate region information; obtaining a plurality of domain transfer sub-modules based on the multi-task structure in the multi-task processing model; training the plurality of domain transfer sub-modules at least based on the source domain data, the target domain data, and the candidate region information to obtain updated candidate region information; and fine-tuning the multi-task processing model at least based on the updated candidate region information. The three-dimensional object recognition method provided by the embodiment of the present invention, through multi-task processing, multi-path training, and a domain transfer mechanism, transfers sub-modules to optimize task branches, enabling the model to automatically adapt to different data distributions, thereby improving the applicability and robustness of the three-dimensional recognition model in new environments, new devices, or new data sources, and significantly reducing the burden of manual retraining. At the same time, by using the source domain data and the candidate region information, the transfer of three-dimensional recognition capabilities under low-annotation or no-annotation conditions can be achieved, significantly reducing the model deployment and transfer costs, and thus achieving more refined cross-task transfer tuning while maintaining the consistency of the model structure, and improving the recognition accuracy of classification and regression tasks.
[0042] Figure 1 Schematically shows an application scenario diagram of a three-dimensional object recognition method, device, equipment, medium, and program product according to an embodiment of the present invention.
[0043] As Figure 1 shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0044] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc.
[0045] The first terminal device 101 and the second terminal device 102 can be devices deployed in different application scenarios, such as industrial inspection devices, medical imaging devices, environmental monitoring devices, etc. For example, the first terminal device 101 and the second terminal device 102 can include baggage CT scanners, millimeter-wave body scanners, X-ray inspection devices, industrial CT detectors, medical imaging scanners, environmental detection instruments, etc. The first terminal device 101 and the second terminal device 102 are capable of collecting various data of the detected target, such as metadata including image information, material attribute information, three-dimensional structure information, detection time, device number, etc. Through the network 104, the collected data can be uploaded to the server 105 in real time or regularly for subsequent data processing.
[0046] The third terminal device 103 can obtain data related to detection records, recognition results, or device status from the server 105 according to the user's query needs and display them. The third terminal device 103 can be various electronic terminal devices with a display function and supporting network browsing, including but not limited to the central control room terminal, duty management terminal, desktop computer, etc. in the centralized supervision system.
[0047] The server 105 can include a model processing module for target recognition, a data storage module, a task scheduling module, etc., and is used to receive the detection data uploaded from multiple terminal devices and perform unified processing and scheduling. For example, the server 105 can perform recognition model inference operations based on the received three-dimensional image data, extract recognition information such as the category, bounding box, and positional relationship of the target object, and return the processing results to the corresponding terminal or upload them to the supervision system.
[0048] It should be noted that the three-dimensional target object recognition method provided by the embodiments of the present invention can generally be executed by the server 105. Correspondingly, the three-dimensional target object recognition device provided by the embodiments of the present invention can generally be set in the server 105. The three-dimensional target object recognition method provided by the embodiments of the present invention can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the three-dimensional target object recognition device provided by the embodiments of the present invention can also be set in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.
[0049] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in
[0050] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 1The described scenario will be used to Figures 2 to 4B describe in detail the three-dimensional object recognition method of the invention embodiment.
[0051] Figure 2 FIG. schematically shows a flowchart of the three-dimensional object recognition method according to an embodiment of the present invention.
[0052] As Figure 2 shown, the three-dimensional object recognition method of this embodiment includes operations S210 to S230, and this three-dimensional object recognition method can be executed by the server 105.
[0053] In operation S210, object data to be recognized is acquired.
[0054] In an embodiment of the present invention, the object data can be sourced from various types of three-dimensional data acquisition or reconstruction devices, such as three-dimensional lidar, structured light scanners, depth cameras, medical imaging devices (such as Magnetic Resonance Imaging (MRI) devices, CT devices, etc.), industrial non-destructive testing devices, computational photography devices, or virtual modeling systems, etc. Depending on the specific application scenario, the representation form of the object data can also be diverse, such as point cloud data, regular or irregular voxel grid data, three-dimensional image data (such as slice sequences), volume density map representations, or sparse three-dimensional feature representations extracted from the original data, etc.
[0055] In operation S220, the object data is input into the object multi-task processing model, and the object data is processed based on the object multi-task processing model to obtain three-dimensional recognition information of the object. The object multi-task processing model can have multi-task capabilities, such as being able to simultaneously perform detection tasks such as object category discrimination and three-dimensional position estimation. During the processing, the object multi-task processing model will first extract features from the object data, for example, combining the classification branch and the regression branch to respectively output the category prediction result and three-dimensional bounding box parameters of the object, forming an intermediate layer representation of the recognition containing object information.
[0056] In operation S230, the recognition result of the object is obtained based on the three-dimensional recognition information. For example, the three-dimensional recognition information can specifically include but is not limited to the category information, position information, three-dimensional position and scale information of the object bounding box, object confidence score, etc. Based on the three-dimensional recognition information, the embodiments of the present invention can output corresponding recognition results for further business analysis, decision control, automated operation, or human-computer interaction presentation.
[0057] In an embodiment of the present invention, the object multi-task processing model can be a cross-domain adaptation model obtained through fine-tuning processing, and its initial structure can be a multi-task processing model trained on source domain data.
[0058] In an embodiment of the present invention, the multi-task processing model may have a multi-task structure, for example, including a classification branch for object classification and a regression branch for 3D localization, and has the ability to perform comprehensive detection tasks on the input 3D data.
[0059] The object data is sourced from real objects in the scene to be detected, which may itself have diversity, occlusion, deformation, and even be at different noise levels and spatial resolutions, resulting in significant differences between the data and the source domain data used during model training at the semantic and geometric levels. Therefore, directly using the original trained model for recognition on this object data often faces "domain shift" problems such as decreased accuracy and insufficient generalization performance.
[0060] According to an embodiment of the present invention, to enable the model to adapt to the distribution characteristics of the object data and improve the recognition accuracy in the target domain, the multi-task processing model can be fine-tuned. The fine-tuning process can include multiple consecutive training stages to gradually guide the model to generalize from the source domain to the target domain.
[0061] Specifically, first, the original multi-task processing model is used to perform forward inference processing on the source domain data and the target domain data, respectively obtaining candidate region information corresponding to the two data domains for guiding feature alignment and structure adaptation in the subsequent training process.
[0062] Exemplarily, the candidate region information may include the output results of each task branch, such as the predicted category and score of the classification branch, the 3D bounding box information of the regression branch, the region feature vector, the position information of the candidate region, etc., and this information will be used in the subsequent transfer training.
[0063] In some embodiments, the candidate region information may include source domain positive sample candidate region information, source domain negative sample candidate region information, target domain positive sample candidate region information, and target domain negative sample candidate region information. Specifically, the source domain positive sample candidate region information and the source domain negative sample candidate region information can be obtained by manual annotation or automatic screening by an existing model, and have a high label accuracy; the target domain positive sample candidate region information and the target domain negative sample candidate region information can be generated according to the existing weak labels or unlabeled data in the target domain through preliminary inference or a pseudo-label mechanism. This candidate region information will be an important part of the supervision signal in the subsequent training process to guide the model's transfer adaptation between different domains.
[0064] After obtaining the candidate region information, the overall network can be decoupled into multiple task sub-paths based on the multi-task structure of the multi-task processing model, and each task sub-path corresponds to constructing a domain transfer sub-module. The structure of the domain transfer sub-module can include components such as a feature extraction sub-network, a domain discriminator, and an adversarial training mechanism, supporting the decoupled optimization of classification domain transfer and regression domain transfer. By separately feeding the candidate region information of each task in the source domain and the target domain into the corresponding domain transfer sub-module and combining the design of the adversarial loss function, the feature space alignment between the source domain and the target domain at the task level can be achieved.
[0065] In some embodiments, the domain transfer sub-module may further include a gradient reversal layer for reversing the gradient direction of the domain discriminator during the training process to achieve effective alignment of the feature spaces of the source domain and the target domain while maintaining the discriminative ability of each task.
[0066] In the embodiments of the present invention, the data samples of the source domain and the target domain can be constructed as paired training samples and input into each domain transfer sub-module for training in combination with their respective candidate region information. Exemplarily, multiple domain transfer sub-modules can be jointly trained based on the source domain data, the target domain data, and the candidate region information. During the training process, the paired training samples are composed of the source domain positive sample candidate region information, the source domain negative sample candidate region information, the target domain positive sample candidate region information, and the target domain negative sample candidate region information collected from the source domain and the target domain respectively. Each pair of paired training samples is associated with corresponding candidate region information for indicating the key spatial regions or target parts that the model should focus on.
[0067] In the training stage, the constructed paired training samples are input into the corresponding domain transfer sub-modules. Each domain transfer sub-module is responsible for extracting features at different levels and completing the target prediction task, while capturing the semantic differences between the source domain and the target domain. During the process of receiving the paired training samples, the model uses the candidate region information to guide the attention mechanism or the region feature fusion strategy, thereby improving the generalization ability and robustness of the model in the target domain.
[0068] After completing the training of the domain transfer sub-modules, the initial candidate region information can be updated based on the output results of each domain transfer sub-module to obtain updated candidate region information. The updated candidate region information can more accurately reflect the structural characteristics of the target domain data, thereby providing effective support for the subsequent fine-tuning of the model.
[0069] In an embodiment of the present invention, the updated candidate region information can be utilized to fine-tune and optimize the original multi-task processing model by combining source domain data and target domain data. The fine-tuning process can adopt a joint optimization strategy to achieve both the maintenance of source domain recognition performance and the improvement of target domain recognition performance. The target multi-task processing model obtained through the above fine-tuning process has stronger cross-domain adaptation ability and can achieve stable and reliable recognition effects in complex 3D recognition tasks with deviated feature distributions.
[0070] The 3D object recognition method provided by the embodiment of the present invention can be not only used for object recognition and target detection, but also extended to tasks such as 3D scene segmentation, pose estimation, point cloud instance segmentation, and multi-view reconstruction. It is also adapted to various industrial scenarios, such as industrial defect detection, medical image recognition, autonomous driving perception systems, virtual reality modeling, warehouse robot operations, and structural scanning recognition in complex spaces.
[0071] It should be noted that the application fields of the 3D object recognition method provided by the embodiment of the present invention are not limited to the above scenarios, and can also be flexibly extended and adapted according to actual needs to meet the requirements of technological development and market demand.
[0072] Figure 3 Schematically shows a model architecture diagram of a 3D object recognition system according to an embodiment of the present invention.
[0073] As Figure 3 shown, the 3D object recognition system according to an embodiment of the present invention may include multiple functional modules for realizing the effective migration of the source domain model to the target domain. For example, it may include: a 3D feature extraction module, multiple task domain migration sub-modules, a candidate region information module, and a 3D multi-task prediction component.
[0074] It should be noted that each functional module in the 3D object recognition system constitutes a specific architecture expression at the implementation level of the "target multi-task processing model" in the embodiment of the present invention, where the target multi-task processing model is pre-trained based on source domain data and fine-tuned and optimized by combining target domain data and candidate region information.
[0075] Referring to Figure 3 , the source domain and target domain image data are simultaneously input into the 3D feature extraction module, which can extract 3D feature representations with geometric and semantic information based on a 3D convolutional network or a sparse volume encoding structure. This module performs a shared feature encoding operation on the two types of data to ensure the comparability of features in different domains.
[0076] After feature extraction, the model can generate candidate region information for subsequent processing. For example, the candidate region information generated from the source domain data can be represented as Prop s, the candidate region information generated from the target domain data can be represented as Prop t . Task1, Task2... Taskn can represent Task 1, Task 2... Task n respectively. In each candidate region information, structural information components for different tasks can be further subdivided. For example, Prop s_task1 represents the candidate region features related to Task 1 in the source domain; Prop t_task2 represents the candidate region features related to Task 2 in the target domain.
[0077] Refer to Figure 3 , based on the multi-task structure in the multi-task processing model, the three-dimensional object recognition system decouples the overall model into multiple task sub-paths (such as classification tasks, regression tasks, semantic segmentation tasks, etc.), and constructs a corresponding domain transfer sub-module for each sub-path (Task1 domain transfer sub-module, Task2 domain transfer sub-module... Taskn domain transfer sub-module). The design of each domain transfer sub-module can include: a sub-network for current task feature processing; a domain discriminator for discriminating the data domain (source domain / target domain); an optional gradient reversal layer to achieve gradient direction regulation in adversarial training.
[0078] It should be noted that Figure 3 schematically shows that the target domain candidate region information Prop t is input into each task domain transfer sub-module. In fact, in some other embodiments, each domain transfer sub-module can also receive paired training samples from both the source domain and the target domain simultaneously to achieve task-level feature alignment.
[0079] Refer to Figure 3 , each domain transfer sub-module gradually reduces the distribution difference between the source domain and the target domain in the corresponding task features through adversarial training, and improves the expression ability of the model in the execution of target domain tasks. After training is completed, the optimization results of each task path will be used to update the target domain candidate region information Prop t , forming a structured input that is more adaptable to the task distribution.
[0080] The updated target domain candidate region information will be input into the 3D multi-task prediction component, which will fuse the recognition results of multiple task paths and output the comprehensive recognition result of the object. The output result can include task-related information such as the target category, the position and size of the three-dimensional bounding box, the orientation angle, and the confidence level.
[0081] The preferred embodiments of the present invention will be specifically described below by taking the three-dimensional detection task of security inspection CT data objects as an example.
[0082] In this embodiment, the target object data includes three-dimensional tomographic image data obtained by the security inspection CT device, and the image contains feature information such as the spatial shape and material density of the object. The system processes this type of three-dimensional image to achieve automatic recognition of the object category and location, and assist in completing tasks such as dangerous goods identification, luggage sorting, and security control.
[0083] In an embodiment of the present invention, the data of the source domain and the target domain can be respectively processed based on a trained multi-task processing model to obtain preliminary detection results. Among them, the source domain data can be historical collected and labeled security inspection CT image data. For example, the source domain data can be image data obtained based on a slip-ring CT device. This type of device has a stable rotating scanning structure and mature accumulation of labeled samples, and has been used for a long time to train a conventional three-dimensional target object recognition model, with characteristics such as stable distribution, high image quality, and sufficient training sample quantity. The target domain data can be new data from different security inspection sites, different devices, or different setting conditions, which can be unlabeled or have only a small amount of labeling.
[0084] In this embodiment, the candidate region information can be, for each candidate target region, extracting its corresponding class prediction score, prediction label, and three-dimensional bounding box parameters (such as the position of the target center point, size, orientation angle, etc.) to form candidate region information, which will be respectively used in the subsequent training of the transfer sub-module and the model optimization stage. The multiple domain transfer sub-modules can include a classification transfer sub-module and a regression transfer sub-module. Among them, the classification transfer sub-module is used for transfer learning of the class information of the candidate region to reduce the difference in class distribution between the source domain and the target domain; the regression transfer sub-module performs transfer training on the position parameters to improve the three-dimensional positioning accuracy in the target domain.
[0085] Each domain transfer sub-module can include a feature extraction branch, a task discriminator, and a domain discriminator module in its network structure, and can optionally be equipped with structures such as a gradient reversal layer to achieve the adversarial optimization training goal. The training of each domain transfer sub-module can be executed in parallel or alternately, or the overall collaborative effect can be improved through a joint optimization strategy.
[0086] It should be noted that the specific network structure of the domain transfer sub-module can be flexibly configured according to different task types or the characteristics of the target domain data. For example, for some task sub-paths that are not sensitive to domain differences or already have good cross-domain generalization ability, only a feature extraction branch and a task discriminator can be set to simplify the network structure and reduce the model complexity and computational resource overhead. In addition, the domain discriminator can be dynamically enabled or disabled for some sub-modules according to the importance or training progress of different tasks. For example, in the initial stage of multi-task training, the task discriminator can be preferentially trained to quickly converge the basic task capabilities; then the domain discriminator can be gradually introduced to improve the cross-domain generalization performance through the adversarial mechanism, so as to achieve a flexible balance of staged optimization and resource scheduling.
[0087] In some embodiments, the regression migration sub-module may further combine the output result of the classification sub-module as input to enhance its prediction robustness. Specifically, the class prediction results output by the classification task, such as class labels and confidence scores, can be injected as high-semantic information into the regression path to guide the bounding box prediction process to focus more on the real object targets in the candidate regions, thereby reducing the error propagation caused by background noise or pseudo-targets.
[0088] Exemplarily, when the confidence of the candidate region is high and the predicted class is clear, the regression module can assign a higher weight to this region to improve the bounding box localization accuracy; when the class prediction uncertainty is high, the regression module can appropriately suppress the output amplitude of the position regression to reduce the risk of false detection.
[0089] In some embodiments, based on the candidate region information corresponding to the source domain data and the target domain data, positive and negative sample regions corresponding to the source domain data and the target domain data can be divided, and paired training samples can be constructed. The construction of paired training samples not only helps the domain discriminator in adversarial training to learn, but also can make full use of the source domain structure information under the condition of scarce target domain samples, improving the sample efficiency of model training.
[0090] The paired training samples can be input into each domain migration sub-module for training. During the training process, the domain discriminator can perform adversarial training on the samples of the source domain and the target domain to optimize the discriminative ability and feature alignment ability of the module, and finally obtain a more adaptable feature representation for the target domain under each task sub-path.
[0091] In some preferred embodiments, the stability and convergence efficiency of the adversarial training can be further enhanced by introducing domain mixing discriminative loss, multi-scale attention mechanism or inter-layer consistency constraint.
[0092] After the domain migration sub-module is trained, the original candidate region information can be updated based on the classification prediction result and the bounding box prediction result. For example, for some low-confidence candidate regions in the target domain, the class label or bounding box position can be re-evaluated by combining the output of the migration sub-module to improve its representation accuracy under the target domain conditions.
[0093] According to the embodiments of the present invention, the update mechanism is iterative, and the candidate regions can be further refined through staged training and feedback mechanism, providing higher-quality training signals for fine-tuning the backbone model.
[0094] The updated candidate region information can be used to fine-tune the multi-task model, enabling it to enhance the adaptability to target domain data while maintaining the recognition accuracy of the source domain. The optimized target multi-task model can be directly deployed in different security inspection CT systems to achieve automatic 3D recognition of luggage or items in complex environments.
[0095] Figure 4A Schematically shows the structural diagram of the domain migration sub-module for the classification task according to an embodiment of the present invention.
[0096] As Figure 4A shown, in this structure, the input includes the positive sample candidate region information P s,prop and negative sample candidate region information N s,prop of the source domain, as well as the positive sample candidate region information P t,prop and negative sample candidate region information N t,prop of the target domain. Exemplarily, these samples can be constructed according to the prediction scores of the candidate regions in the classification task. Positive samples are usually high-confidence target candidate regions, and negative samples are background regions or low-confidence regions.
[0097] For example, in the actual construction process, the division of positive and negative samples can be combined with heuristic threshold setting. For example, samples with a classification confidence greater than a certain threshold are set as positive samples, those lower than another threshold are set as negative samples, and samples in the intermediate region can be selectively ignored or re-evaluated at a later stage.
[0098] Referring to Figure 4A , the candidate region samples of the source domain and the target domain can be input into the classification domain migration feature extractor together to extract intermediate feature representations F(s) and F(t), corresponding to the feature outputs of the source domain and the target domain respectively. F(s) and F(t) are input into the domain discriminator together. Through the adversarial training mechanism, the feature extractor learns domain-invariant discriminative features, thereby achieving feature alignment between the source domain and the target domain in the classification task. On this basis, F(s) can be further input into the classifier module, combined with the true label information of the source domain to train the prediction ability of the target class to generate the classification prediction result P t,prop,cls .
[0099] In some other embodiments, the candidate region samples of the target domain can also be input into the target domain alternative region information classification prediction module. In this module, pre-classification modeling or pseudo-label prediction of the candidate regions is performed, and based on the generated region classification prediction information as a guiding signal, it is further input into the classifier to generate the classification prediction result P t,prop,cls .
[0100] In this embodiment, the domain discriminator can be based on a binary classification network structure, whose goal is to discriminate whether the input features come from the source domain or the target domain; while the classification domain transfer feature extractor can maximize the domain discriminator error through a gradient reversal layer, so that the features it outputs are insensitive to domain information.
[0101] Figure 4B Schematically shows the structural diagram of the domain transfer sub-module for the regression task according to an embodiment of the present invention.
[0102] As Figure 4B shown, in this structure, the input is the positive sample candidate region information P of the source domain s,prop and the positive sample candidate region information P of the target domain t,prop . Different from the classification task, the regression task can be trained and aligned only for positive samples, so the negative sample information does not participate in the processing in this path. This is because the core goal of the regression task is to accurately predict continuous variables such as the three-dimensional bounding box position and size of the target object, and this kind of supervision is only effective for the real target area, and the background or invalid area does not have an effective regression target.
[0103] Referring to Figure 4B , the positive sample candidate region information P of the source domain s,prop and the positive sample candidate region information P of the target domain t,prop are input into the regression domain transfer feature extractor to generate the intermediate features F(s) and F(t) of the source domain and the target domain respectively. During the training process, F(s) and F(t) are also fed into the regression discriminator for adversarial training to promote the alignment of the regression feature space between the source domain and the target domain, thereby improving the prediction accuracy of the target domain bounding box. On this basis, F(s) can be further input into the classifier module, and the regression prediction result P t,prop,reg is output in combination with the true label information of the source domain.
[0104] In some embodiments, the target domain candidate region information P t, prop can also be input into the regression prediction module for the target domain alternative region information to model or estimate the geometric features of the candidate region, so as to enhance the representation ability of region localization. The output of this module is then input into the regressor module, and jointly trains the regression ability of the three-dimensional bounding box with the intermediate feature F(s) of the feature extractor, and finally outputs the regression prediction result P t,prop,reg , indicating the regression prediction result of the target domain candidate region.
[0105] The regression discriminator structure is similar to the discriminator in the classification path, but its input features may pay more attention to spatial continuity and structural expression ability, such as introducing local geometric coding, voxel statistical features or scale-sensitive representation. The purpose of adversarial training is to guide the regression features to have good cross-domain consistency, so as to improve the localization robustness of the target domain when there is a lack of accurate annotation.
[0106] In addition, in some preferred embodiments, the feature extractor in the regression task can further fuse the output information of the classification sub-module, such as the class prediction result or the class probability distribution, to assist in guiding the selection of the scale, shape, or target perception range of the bounding box, forming a weakly coupled structure between tasks and further improving the overall detection performance.
[0107] Based on the above three-dimensional object recognition method, an embodiment of the present invention also provides a three-dimensional object recognition device. The following will be combined with Figure 5 to describe this device in detail.
[0108] Figure 5 The structural block diagram of the three-dimensional object recognition device according to an embodiment of the present invention is schematically shown.
[0109] As Figure 5 shown, the three-dimensional object recognition device 500 of this embodiment includes a data acquisition module 510, a processing module 520, and a recognition result acquisition module 530.
[0110] The data acquisition module 510 can be used to acquire the data of the object to be recognized. In one embodiment, the data acquisition module 510 can be used to perform the operation S210 described above, which will not be elaborated here.
[0111] The processing module 520 can be used to input the object data into the target multi-task processing model, and process the object data based on the target multi-task processing model to obtain the three-dimensional recognition information of the object. Among them, the target multi-task processing model includes a model obtained by fine-tuning the multi-task processing model trained based on the source domain data. The fine-tuning process includes: processing the source domain data and the target domain data based on the multi-task processing model to obtain candidate region information; obtaining multiple domain transfer sub-modules based on the multi-task structure in the multi-task processing model, and training multiple domain transfer sub-modules at least based on the source domain data, the target domain data, and the candidate region information to obtain updated candidate region information; fine-tuning the multi-task processing model at least based on the updated candidate region information. In one embodiment, the processing module 520 can be used to perform the operation S220 described above, which will not be elaborated here.
[0112] The recognition result acquisition module 530 can be used to obtain the recognition result of the object based on the three-dimensional recognition information. In one embodiment, the recognition result acquisition module 530 can be used to perform the operation S230 described above, which will not be elaborated here.
[0113] According to an embodiment of the present invention, the processing module 520 may further be configured to decouple tasks in the multitasking model to obtain multiple task sub-paths; construct multiple sub-network structures based on the multiple task sub-paths; and embed a domain discriminator in the multiple sub-network structures to obtain multiple domain transfer sub-modules.
[0114] According to an embodiment of the present invention, the multiple domain transfer sub-modules include a gradient reversal layer connected to the domain discriminator, and the gradient reversal layer is configured to reverse the gradient direction of the domain discriminator during training.
[0115] According to an embodiment of the present invention, the processing module 520 may further be configured to input source domain data and target domain data into the multitasking model respectively to obtain corresponding preliminary detection results; and extract at least class prediction scores and three-dimensional bounding box parameters from the preliminary detection results as candidate region information.
[0116] According to an embodiment of the present invention, the multiple domain transfer sub-modules at least include a classification transfer sub-module and a regression transfer sub-module. The classification transfer sub-module is configured to predict the class information of the candidate region, and the regression transfer sub-module is configured to predict the position information of the candidate region. The input of the regression transfer sub-module at least includes the class prediction result output by the classification transfer sub-module.
[0117] According to an embodiment of the present invention, the candidate region information includes source domain positive sample candidate region information, source domain negative sample candidate region information, target domain positive sample candidate region information, and target domain negative sample candidate region information.
[0118] According to an embodiment of the present invention, the processing module 520 may further be configured to construct paired training samples composed of positive and negative samples corresponding to source domain data and target domain data, and the paired training samples are associated with the candidate region information; input the paired training samples into the corresponding domain transfer sub-module to obtain a class prediction result and a bounding box prediction result; and update the candidate region information based on the class prediction result and the bounding box prediction result to generate updated candidate region information.
[0119] According to an embodiment of the present invention, any multiple of the data acquisition module 510, the processing module 520, and the recognition result acquisition module 530 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the data acquisition module 510, the processing module 520, and the recognition result acquisition module 530 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or may be implemented by any other reasonable means such as hardware or firmware through circuit integration or packaging, or may be implemented in any one of the three implementation manners of software, hardware, and firmware, or in any suitable combination of several of them. Alternatively, at least one of the data acquisition module 510, the processing module 520, and the recognition result acquisition module 530 may be at least partially implemented as a computer program module, and when the computer program module is run, corresponding functions may be executed.
[0120] Figure 6 A block diagram of an electronic device suitable for implementing a three-dimensional object recognition method according to an embodiment of the present invention is schematically shown.
[0121] As Figure 6 shown, the electronic device 600 according to an embodiment of the present invention includes a processor 601, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM 602) or a program loaded from a storage section 608 into a random access memory (RAM 603). The processor 601 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 601 may also include on-board memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0122] In the RAM 603, various programs and data required for the operation of the electronic device 600 are stored. The processor 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. The processor 601 performs various operations of the method flow according to the embodiments of the present invention by executing the programs in the ROM 602 and / or the RAM 603. It should be noted that the programs can also be stored in one or more memories other than the ROM 602 and the RAM 603. The processor 601 can also perform various operations of the method flow according to the embodiments of the present invention by executing the programs stored in one or more memories.
[0123] According to an embodiment of the present invention, the electronic device 600 may further include an input / output (I / O) interface 605, and the input / output (I / O) interface 605 is also connected to the bus 604. The electronic device 600 may further include one or more of the following components connected to the input / output (I / O) interface 605: an input portion 606 including a keyboard, a mouse, etc.; an output portion 607 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 608 including a hard disk, etc.; and a communication portion 609 including a network interface card such as a LAN card, a modem, etc. The communication portion 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output (I / O) interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that a computer program read from it can be installed into the storage portion 608 as needed.
[0124] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present invention is implemented.
[0125] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM 603), read-only memory (ROM 602), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the above-described ROM 602 and / or RAM 603 and / or one or more memories other than ROM 602 and RAM 603.
[0126] An embodiment of the present invention further includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the user interaction method provided by the embodiment of the present invention.
[0127] When the computer program is executed by the processor 601, it executes the above functions defined in the system / apparatus of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0128] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 609, and / or be installed from the removable medium 611. The program code contained in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0129] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, and the above-mentioned module, segment of a program, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0130] Those skilled in the art can understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.
Claims
1. A three-dimensional object recognition method, characterized in that: include: Acquire data of the target object to be identified; Inputting the target object data into a target multi-task processing model, processing the target object data based on the target multi-task processing model, and obtaining three-dimensional recognition information of the target object; as well as Acquire the recognition result of the target object based on the three-dimensional recognition information, Among them, the target multi-task processing model is a model obtained by fine-tuning the multi-task processing model trained on source domain data, and the fine-tuning process includes: processing the source domain data and the target domain data based on the multi-task processing model to obtain candidate region information; obtaining multiple domain migration sub-modules based on the multi-task structure in the multi-task processing model; training multiple domain migration sub-modules based on at least the source domain data, the target domain data and the candidate region information to obtain updated candidate region information; and fine-tuning the multi-task processing model based on at least the updated candidate region information.
2. The method according to claim 1, characterized in that The acquiring of multiple domain transfer submodules based on the multi-task structure in the multi-task processing model specifically includes: Decoupling the multi-task structure in the multi-task processing model to obtain multiple task sub-paths; Constructing multiple sub-network structures based on the multiple task sub-paths; and A domain discriminator is embedded in a plurality of the sub-network structures to obtain a plurality of the domain transfer sub-modules.
3. The method according to claim 2, characterized in that The plurality of domain transfer submodules include a gradient reversal layer connected to the domain discriminator, wherein the gradient reversal layer is used to reverse the gradient direction of the domain discriminator during training.
4. The method according to any one of claims 1 to 3, characterized in that: The target object data includes three-dimensional image data of the target object obtained by the security inspection equipment, and the recognition result is used to identify the category information and location information of the target object.
5. The method according to claim 4, characterized in that The target multi-task processing model is at least used to perform target detection on the target object.
6. The method according to claim 5, characterized in that The processing of the source domain data and the target domain data based on the multi-task processing model to obtain candidate region information specifically includes: Inputting the source domain data and the target domain data into the multi-task processing model respectively to obtain corresponding preliminary detection results; and At least a category prediction score and three-dimensional bounding box parameters are extracted from the preliminary detection result as the candidate region information.
7. The method according to claim 5, characterized in that The multiple domain transfer submodules at least include a classification transfer submodule and a regression transfer submodule. The classification transfer submodule is used to predict the category information of the target object, and the regression transfer submodule is used to predict the position information of the target object.
8. The method according to claim 7, characterized in that The input of the regression migration submodule at least includes the category prediction result output by the classification migration submodule.
9. The method according to any one of claims 1 to 3, 5 to 8, characterized in that: The candidate region information includes source domain positive sample candidate region information, source domain negative sample candidate region information, target domain positive sample candidate region information and target domain negative sample candidate region information.
10. The method according to claim 9, characterized in that The training of the plurality of domain transfer submodules based at least on the source domain data, the target domain data and the candidate region information specifically includes: Constructing paired training samples consisting of positive samples and negative samples corresponding to the source domain data and the target domain data, wherein the paired training samples are associated with the candidate region information; Inputting the paired training samples into the corresponding domain transfer submodule to obtain a target prediction result; and Based on the target prediction result, the candidate region information is updated to generate the updated candidate region information.
11. A three-dimensional object recognition device, characterized in that: The device comprises: The data acquisition module is used to: acquire the target object data to be identified; a processing module, used to: input the target object data into a target multi-task processing model, process the target object data based on the target multi-task processing model, and obtain three-dimensional recognition information of the target object, wherein the target multi-task processing model is a model obtained by fine-tuning a multi-task processing model trained with source domain data, and the fine-tuning process includes: processing the source domain data and the target domain data based on the multi-task processing model to obtain candidate area information; obtaining a plurality of domain migration submodules based on a multi-task structure in the multi-task processing model, training a plurality of the domain migration submodules based on at least the source domain data, the target domain data and the candidate area information to obtain updated candidate area information; and fine-tuning the multi-task processing model based on at least the updated candidate area information; and The recognition result acquisition module is used to obtain the recognition result of the target object based on the three-dimensional recognition information.
12. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 10.
13. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
14. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
Citation Information
Patent Citations
Cross-domain target detection method based on domain self-adaption
CN114626461A
Metalearning-based two-stage target intelligent detection algorithm and system
CN114972845A
Cross-domain target detection method based on multi-level domain adaptive weak supervised learning
CN116342942A
Migration learning-driven unexploded ammunition crater positioning and identification method
CN117115678A
High-precision spacecraft domain adaptive detection method
CN117994637A
Cited By
Multi-view X-ray security check image three-dimensional reconstruction management method
CN120747117A
Three-dimensional reconstruction management method based on multi-view X-ray security inspection images
CN120747117B