Deep optimization-based multi-mode workpiece structure detail detection method and device

By combining the deep optimization method of image and text information, the multi-modal workpiece structure detail detection method is adopted to solve the problems of inefficiency and stability of traditional detection methods, and efficient and accurate workpiece structure detection is achieved to adapt to workpiece feature recognition in a variety of scenarios.

CN120510092APending Publication Date: 2025-08-19UNIV OF SCI & TECH BEIJING
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510524501.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The existing workpiece structure detection methods rely on artificial visual inspection or traditional sensors, are inefficient and susceptible to human factors, making it difficult to ensure the stability and accuracy of the detection results. The existing multimodal fusion method fails to effectively improve the detection performance, and it takes a lot of manpower and time to retrain the model when facing unknown defect types.

Method used

Using a multimodal workpiece structure detail detection method based on deep optimization, the workpiece edge segmentation network, deep refinement network, multimodal interactive fusion module and text encoder are combined to obtain and fuse images and text information to generate high-quality workpiece structure description data.

Benefits of technology

It improves the accuracy and robustness of detection, reduces the computational complexity and labor costs, and can automatically identify the workpiece features in new categories or scenarios, meeting the requirements of depth information accuracy and stability in high-demand application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510092A_ABST
    Figure CN120510092A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode workpiece structure detail detection method and device based on deep optimization, and relates to the technical field of workpiece detection. The method comprises the following steps: acquiring training data, and training an initial workpiece structure description network to obtain a trained workpiece structure description network; obtaining a to-be-processed image, a depth image and text description of the target workpiece; inputting a to-be-processed image into the workpiece edge segmentation network to obtain preliminary contour data; inputting the depth image into a workpiece depth refining network to obtain a refined depth image; inputting the text description into a language text conversion module, and inputting obtained features into a text encoder to obtain text encoding features; and inputting the preliminary contour data, the refined depth image and the text coding features into a multi-modal interactive fusion module, and inputting the obtained fusion features into a multi-modal codec to obtain structure description data of the target workpiece. According to the invention, the detection accuracy and robustness of the model in various scenes are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of workpiece detection, and in particular to a multimodal workpiece structural detail detection method and device based on depth optimization. Background Art

[0002] In modern industrial production, the inspection and analysis of workpiece structures are critical to ensuring product quality and production efficiency. With the continuous upgrading of production processes, the complexity and diversity of workpieces are increasing, posing a significant challenge to traditional inspection technologies. Many industries, such as automotive manufacturing, aerospace, and electronics production, require precise inspection of workpiece structure dimensions, shapes, and defects to ensure product safety and reliability. However, existing inspection methods often rely on manual visual inspection or traditional sensor technology. These methods are not only inefficient but also susceptible to human factors, making it difficult to guarantee the stability and accuracy of inspection results.

[0003] In specific inspection tasks, professional image acquisition equipment is used to collect high-definition images of various workpieces from multiple angles and in multiple environmental conditions. The constructed model can perform in-depth analysis of the collected images, accurately parse the depth information of the workpiece in the image and the edge detail information structure, and then accurately identify the key features such as the outline and size of the workpiece. At the same time, the model can also be combined with detailed text descriptions. These texts cover the details of the workpiece in the collected image, such as material properties, design process points, and workpiece structure related information, including internal structural layout, key connection parts, etc. In addition, the model can also integrate and analyze texts for other specific needs, such as the expected application scenarios of the workpiece, performance index requirements, etc., so as to comprehensively and accurately detect and evaluate the workpiece structure.

[0004] With the development of computer vision and artificial intelligence technologies, workpiece structure detection methods based on image analysis have gradually become a research hotspot. By capturing real-time images of workpieces with high-resolution cameras and using deep learning algorithms for feature extraction and defect detection, the efficiency and accuracy of detection can be significantly improved. However, although this method has achieved certain success in many applications, it still faces some technical bottlenecks. First, existing image analysis technologies usually only focus on visual information and fail to effectively integrate text information, which limits a more comprehensive understanding and detection of workpiece structures. In addition, existing systems are mostly single-modal and are prone to insufficient adaptability in different working environments, lighting conditions, and shooting angles, resulting in unsatisfactory detection results.

[0005] To address these issues, researchers have begun exploring multimodal fusion methods, combining image and text information to achieve more accurate workpiece structure detection. However, existing multimodal fusion methods are relatively simple, usually manifested as feature splicing or simple cross-self-attention mechanisms, which fail to achieve fine-grained feature fusion and thus fail to effectively improve detection performance. In addition, with the continuous changes in workpiece design and materials, new structures and their related defect types are emerging in an endless stream, and traditional detection systems often rely on pre-defined categories for training. Faced with unknown new types of defects, the system finds it difficult to respond effectively, requiring a lot of manpower and time to retrain the model. Summary of the Invention

[0006] In order to solve the technical problems existing in the prior art, the embodiments of the present invention provide a multi-modal workpiece structural detail detection method and device based on depth optimization. The technical solution is as follows:

[0007] In one aspect, a multimodal workpiece structural detail detection method based on depth optimization is provided. The method is implemented by a multimodal workpiece structural detail detection device based on depth optimization, and the method includes:

[0008] S1. Acquire training data and train an initial artifact structure description network to obtain a trained artifact structure description network, wherein the artifact structure description network includes an artifact edge segmentation network, an artifact depth refinement network, a multimodal interactive fusion module, a multimodal encoder / decoder, a language-to-text conversion module, and a text encoder;

[0009] S2, obtaining the image to be processed, the depth image and the text description of the target workpiece;

[0010] S3, inputting the image to be processed into the workpiece edge segmentation network to obtain preliminary contour data;

[0011] S4, inputting the depth image into the workpiece depth refinement network to obtain a refined depth image;

[0012] S5. Input the text description into the language-to-text conversion module, and input the obtained features into the text encoder to obtain text encoding features;

[0013] S6. Input the preliminary contour data, refined depth image and text encoding features into the multimodal interactive fusion module, input the obtained fusion features into the multimodal encoder-decoder to obtain the structural description data of the target workpiece.

[0014] On the other hand, a multimodal workpiece structure detail detection device based on depth optimization is provided, which is applied to a multimodal workpiece structure detail detection method based on depth optimization, and the device includes:

[0015] A training unit is used to obtain training data and train the initial artifact structure description network to obtain a trained artifact structure description network, wherein the artifact structure description network includes an artifact edge segmentation network, an artifact depth refinement network, a multimodal interactive fusion module, a multimodal encoder-decoder, a language-to-text conversion module, and a text encoder;

[0016] An acquisition unit, configured to acquire an image to be processed, a depth image, and a text description of a target workpiece;

[0017] A segmentation unit, used to input the image to be processed into the workpiece edge segmentation network to obtain preliminary contour data;

[0018] A refinement unit, configured to input the depth image into a workpiece depth refinement network to obtain a refined depth image;

[0019] The encoding unit is used to input the text description into the language-to-text conversion module, and input the obtained features into the text encoder to obtain text encoding features;

[0020] The fusion unit is used to input the preliminary contour data, the refined depth image and the text encoding features into the multimodal interactive fusion module, and input the obtained fusion features into the multimodal encoder-decoder to obtain the structural description data of the target workpiece.

[0021] On the other hand, a multimodal workpiece structural detail detection device based on depth optimization is provided, and the multimodal workpiece structural detail detection device based on depth optimization includes: a processor; a memory, wherein computer-readable instructions are stored on the memory, and when the computer-readable instructions are executed by the processor, any one of the multimodal workpiece structural detail detection methods based on depth optimization is implemented.

[0022] On the other hand, a computer-readable storage medium is provided, wherein the storage medium stores at least one instruction, and the at least one instruction is loaded and executed by a processor to implement any one of the above-mentioned multimodal workpiece structural detail detection methods based on depth optimization.

[0023] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0024] In embodiments of the present invention, the intelligent boundary detector employed can efficiently and accurately extract object contours from input images. By rapidly processing objects of varying shapes and scales, it significantly reduces computational complexity and improves the stability of the results, addressing the accuracy issues inherent in traditional methods due to computational complexity and low stability. The depth optimization core module redefines the depth image optimization task, achieving high-quality depth image output through Poisson fusion of local difference perturbations and edge distortion noise. This innovative mechanism effectively enhances edge clarity and structural consistency in depth images, meeting the accuracy requirements of depth information in demanding applications. The multimodal interactive fusion module based on mask reconstruction deeply fuses visual and textual features while fully exploiting subtle correlations between image and text, resulting in more complementary multimodal features. This significantly improves information transfer efficiency and resolution, and provides better support for downstream tasks. By incorporating a text analysis mechanism, the present invention provides contextual support for specific artifacts when new categories or scenarios emerge. This innovative design automatically identifies relevant features through natural language processing, eliminating the need for extensive annotated data. This improves the model's detection accuracy and robustness in diverse scenarios, significantly reducing labor costs and maintenance efforts. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0026] Figure 1 This is a flow chart of a multimodal workpiece structural detail detection method based on depth optimization provided by an embodiment of the present invention;

[0027] Figure 2 This is a schematic diagram of the structure of a backbone network provided by an embodiment of the present invention;

[0028] Figure 3 1 is a schematic diagram of the structure of an image edge segmentation network provided by an embodiment of the present invention;

[0029] Figure 4 1 is a schematic structural diagram of a primitive segmentation and primitive sequence decoder provided by an embodiment of the present invention;

[0030] Figure 5 1 is a schematic diagram of the structure of an image depth refinement network provided by an embodiment of the present invention;

[0031] Figure 6 This is a schematic diagram of the structure of a multimodal interactive fusion and decoder provided by an embodiment of the present invention;

[0032] Figure 7 This is a block diagram of a multimodal workpiece structure detail detection device based on depth optimization provided by an embodiment of the present invention;

[0033] Figure 8 It is a structural schematic diagram of a multimodal workpiece structural detail detection device based on depth optimization provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0034] The technical solution of the present invention is described below in conjunction with the accompanying drawings.

[0035] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as an "exemplary" in the present invention should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of the word "exemplary" is intended to present concepts in a concrete manner. Furthermore, in the embodiments of the present invention, "and / or" can mean both or either of the two.

[0036] In the embodiments of the present invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, when the distinction is not emphasized, the meanings they convey are the same. The terms "of," "corresponding," and "corresponding" may sometimes be used interchangeably. It should be noted that, when the distinction is not emphasized, the meanings they convey are the same.

[0037] In the embodiments of the present invention, sometimes a subscript such as W1 may be written as a non-subscript such as W1. When the difference is not emphasized, the meanings to be expressed are the same.

[0038] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0039] The embodiment of the present invention provides a multimodal workpiece structure detail detection method based on depth optimization, which can be implemented by a multimodal workpiece structure detail detection device based on depth optimization, which can be a terminal or a server. Figure 1 The flowchart of the multimodal workpiece structure detail detection method based on deep optimization is shown. The processing flow of the method may include the following steps:

[0040] S1. Obtain training data, train the initial workpiece structure description network, and obtain a trained workpiece structure description network.

[0041] Among them, the artifact structure description network includes an artifact edge segmentation network, an artifact depth refinement network, a multimodal interaction fusion module, a multimodal encoder-decoder, a language-to-text conversion module, and a text encoder.

[0042] In one feasible implementation, a high-resolution depth image camera is used to capture the workpiece structure and simultaneously collect representative textual descriptions of the workpiece structure. The workpiece depth and edge details in the image are manually annotated. The representative textual descriptions are used to build a knowledge database and extract feature knowledge vectors.

[0043] All image-label pairs are divided into training and validation sets in a 4:1 ratio to form a complete dataset. The algorithm is trained on the training set and the model performance is verified on the validation set. The model with the best performance on the validation set is selected as the final model.

[0044] All image-label pairs are augmented by random horizontal mirror flipping, random scale scaling, and random size cropping to expand the training dataset and obtain input images and labels.

[0045] like Figure 2 As shown, for edge network optimization, preprocessed image data is fed into the model. The intelligent boundary detector is trained on a large number of image samples annotated with object bounding boxes. A bounding box detection loss function (such as the Intersection over Union (IoU) loss) is used. Backpropagation is used to continuously adjust the detector parameters to improve bounding box detection accuracy. The primitive segmenter is trained within the object image region where bounding boxes have been detected. Based on the annotated primitive positions and confidence information, it uses a primitive parsing loss function (including position and confidence losses) to optimize parameters, combining techniques such as feature isolation mechanisms, multi-endpoint primitive representation, and dynamic feature adaptation strategies. The primitive sequence decoder is trained using a sequence loss function based on the primitive order annotations in the training data, learning to predict the relative order of primitives.

[0046] For depth refinement, a self-learning distillation algorithm is used to generate pseudo-depth edge labels, supplemented by preprocessed depth images and object contour information. This is then used to train a depth optimization network. Based on the loss function corresponding to the edge-guided gradient correction mechanism and edge fusion smoothing strategy, parameters of the convolutional and pooling layers in the network are adjusted. A coarse-to-fine training strategy is employed, first training the network on low-resolution depth images to learn the global structure, then gradually increasing the resolution to refine details. This allows the network to accurately learn the rules of depth optimization, enhancing both depth edge delineation and the consistency of depth structure.

[0047] S2. Obtain the image to be processed, the depth image, and the text description of the target workpiece.

[0048] S3. Input the image to be processed into the workpiece edge segmentation network to obtain preliminary contour data.

[0049] Optionally, the step S3 of inputting the image to be processed into a workpiece edge segmentation network to obtain preliminary contour data includes:

[0050] S31. Use an intelligent boundary detector to perform a comprehensive scan on the image to be processed and lock the boundary range of a specific object in the image.

[0051] S32. Segmentation is performed based on primitive parser primitives within the boundaries of a specific object, effectively limiting the query to interact only with object features of fixed length.

[0052] S33. By introducing a multi-endpoint primitive representation, the primitive composed of multiple endpoints is parsed to generate multi-endpoint location queries and structure queries.

[0053] S34. Use dynamic feature adaptation strategy to dynamically update query location embedding according to real-time changes in object features.

[0054] S35. Based on the primitive sequence decoder and output results, it accepts queries and position queries, fuses the queries and query position information, obtains the positions and scores of multiple primitive position lines through the feedforward fully connected layer, directly infers the combined output of each primitive, and generates a preliminary contour model of the object.

[0055] In a feasible implementation, Figure 3 and 4As shown in the figure, an intelligent boundary detector (image feature extraction module and detection head) is first used to fully scan the input image. Leveraging its advanced deep learning algorithm architecture, it can precisely locate the boundaries of specific objects in the image. Regardless of the object's shape or scale, it can quickly and accurately delineate its location, building a precise operational domain for subsequent refined processing. Within the identified object boundaries, the primitive parser performs a key role in primitive segmentation. Leveraging a feature isolation mechanism, it effectively restricts queries to interact only with fixed-length object features, avoiding interference between different object features and completely eliminating the risk of incorrectly connecting vertices of different objects. Furthermore, by introducing a multi-endpoint primitive representation, it can more efficiently parse primitives composed of multiple endpoints, generating multi-endpoint location and structure queries to address the need to describe the contours of complex objects. Using a dynamic feature adaptation strategy, query position embeddings can be dynamically updated based on real-time changes in object features, improving the accuracy and completeness of primitive segmentation. During the primitive parsing process, the position and confidence score of each primitive are simultaneously predicted, providing key information for the subsequent primitive sequence decoder. Finally, the primitive sequence decoder accepts queries and position queries based on the output results of the primitive analyzer. After query fusion and query and query position information, the positions and scores of multiple primitive position lines are obtained through the feedforward fully connected layer, and the combined output of each primitive is directly inferred to quickly generate a preliminary contour model of the object.

[0056] S4. Input the depth image into the workpiece depth refinement network to obtain a refined depth image.

[0057] Optionally, the step S4 of inputting the depth image into a workpiece depth refinement network to obtain a refined depth image includes:

[0058] S41. Apply depth optimization enhancement technology to finely optimize the depth image containing specific objects.

[0059] S42. Generate low-noise deep edge pseudo-labels through self-learning distillation technology.

[0060] S43. Use edge-guided gradient correction mechanism and edge fusion smoothing strategy to optimize the deep optimization network.

[0061] S44. Adopt a coarse-to-fine optimization strategy, first control the overall spatial structure and depth balance of the depth image from a macroscopic perspective, and then gradually refine the detailed information of the high-frequency area to obtain a refined depth image.

[0062] In a feasible implementation, Figure 5As shown. Deep optimization enhancement technology is applied to fine-tune depth images containing specific objects: the depth optimization is reshaped into a Poisson fusion problem with local difference perturbations and edge distortion noise, and the core module of deep optimization is built on this basis. Low-noise depth edge pseudo-labels are generated through self-learning distillation technology. With the inherent logical relationship between the image's own depth information and feature information, in the absence of external high-quality annotated data, deep edge features are efficiently extracted and converted into pseudo-labels, providing a reliable source of supervision information for the training of the deep optimization network. Secondly, the edge-guided gradient correction mechanism and edge fusion smoothing strategy are used to optimize the deep optimization network. The edge-guided gradient correction mechanism focuses on the edge areas of the depth image, driving the network to accurately learn and optimize at places where the depth changes sharply, thereby enhancing the clarity of edge depiction; the edge fusion smoothing strategy ensures a natural and smooth transition between different depth areas, ensuring the consistency and coordination of the generated deep structure. A coarse-to-fine optimization strategy is adopted. First, the overall spatial structure and depth balance of the depth image are controlled from a macro perspective to prevent local optimization from damaging the global structure. Then, the detailed information of the high-frequency area is gradually refined. The final output depth image has a reasonable overall depth layout and can clearly present the depth changes of details such as object edges and textures, fully meeting the strict requirements of different application scenarios for depth information accuracy and integrity.

[0063] S5. Input the text description into the language-to-text conversion module, and input the obtained features into the text encoder to obtain text encoding features.

[0064] Optionally, the step S5 inputs the text description into a language-to-text conversion module, inputs the obtained features into a text encoder, and obtains text encoding features, including:

[0065] Using natural language processing technology, text descriptions related to artifacts are extracted and analyzed to identify artifact features, important parameters, and key attributes.

[0066] S6. Input the preliminary contour data, refined depth image and text encoding features into the multimodal interactive fusion module, input the obtained fusion features into the multimodal encoder-decoder to obtain the structural description data of the target workpiece.

[0067] Optionally, the step S6 of inputting the preliminary contour data, the refined depth image, and the text encoding features into a multimodal interactive fusion module, and inputting the obtained fusion features into a multimodal encoder-decoder to obtain the structural description data of the target workpiece includes:

[0068] S61. Input the preliminary contour data, refined depth image and text encoding features into the cross attention layer to generate a new feature representation. The features processed by the cross attention layer enter the Softmax layer.

[0069] S62. Convert the feature into a probability vector through the Softmax layer and find the text description index number corresponding to the position with the highest probability.

[0070] S63. Search and match corresponding character units in the character unit sequence through the index number vector, and splice the found character units together to form a representative description.

[0071] S64: The concatenated representative description is input into the large language model for further processing and generation.

[0072] In a feasible implementation, Figure 6 As shown in the figure, the extracted object contour information and the processed depth image information are closely related and complement each other. Object contour information can serve as a priori guidance in the depth image optimization process, helping to accurately locate the object region in the depth image. This allows for focused optimization of depth regions related to the object, significantly improving the targetedness and accuracy of depth optimization. Furthermore, depth image information provides richer feature information, helping to more accurately discern the object's three-dimensional structure and occlusion relationships during object contour extraction. This further optimizes the primitive parsing and contour generation processes, improving the accuracy and completeness of object contour extraction. The depth and edge features of the workpiece are input into the cross-attention layer to generate a new feature representation. After processing the cross-attention layer, the features are fed into the softmax layer. The softmax layer converts the features into a probability vector and finds the text description index corresponding to the position with the highest probability. The index vector is used to find and match the corresponding character unit in the character unit sequence. The found character units are concatenated to form a representative description. The concatenated representative description is then input into the large language model for further processing and generation.

[0073] In summary, the entire practical application process is as follows: An image containing a specific object and a depth image to be processed are fed into a fully trained fusion framework. The framework then outputs a preliminary object contour model and depth image, which are then subjected to targeted optimization and subsequent processing, resulting in a final output that combines precise object contours with high-quality depth information. Furthermore, the framework combines textual feature descriptions with the text integration capabilities of a large language model to produce a detailed and precise description of the artifact structure, achieving exceptional performance in detailed artifact structure analysis and inspection processing.

[0074] In embodiments of the present invention, the intelligent boundary detector employed can efficiently and accurately extract object contours from input images. By rapidly processing objects of varying shapes and scales, it significantly reduces computational complexity and improves the stability of the results, addressing the accuracy issues inherent in traditional methods due to computational complexity and low stability. The depth optimization core module redefines the depth image optimization task, achieving high-quality depth image output through Poisson fusion of local difference perturbations and edge distortion noise. This innovative mechanism effectively enhances edge clarity and structural consistency in depth images, meeting the accuracy requirements of depth information in demanding applications. The multimodal interactive fusion module based on mask reconstruction deeply fuses visual and textual features while fully exploiting subtle correlations between image and text, resulting in more complementary multimodal features. This significantly improves information transfer efficiency and resolution, and provides better support for downstream tasks. By incorporating a text analysis mechanism, the present invention provides contextual support for specific artifacts when new categories or scenarios emerge. This innovative design automatically identifies relevant features through natural language processing, eliminating the need for extensive annotated data. This improves the model's detection accuracy and robustness in diverse scenarios, significantly reducing labor costs and maintenance efforts.

[0075] Figure 7 This is a block diagram of a multi-modal workpiece structure detail detection device based on depth optimization provided by an embodiment of the present invention, which is used in a multi-modal workpiece structure detail detection method based on depth optimization. Figure 7 , the device comprises:

[0076] A training unit 710 is configured to obtain training data and train an initial artifact structure description network to obtain a trained artifact structure description network, wherein the artifact structure description network includes an artifact edge segmentation network, an artifact depth refinement network, a multimodal interaction fusion module, a multimodal encoder / decoder, a language-to-text conversion module, and a text encoder.

[0077] An acquisition unit 720 is configured to acquire an image to be processed, a depth image, and a text description of a target workpiece;

[0078] The segmentation unit 730 is used to input the image to be processed into the workpiece edge segmentation network to obtain preliminary contour data;

[0079] A refinement unit 740 is configured to input the depth image into a workpiece depth refinement network to obtain a refined depth image;

[0080] The encoding unit 750 is used to input the text description into the language-to-text conversion module, and input the obtained features into the text encoder to obtain text encoding features;

[0081] The fusion unit 760 is used to input the preliminary contour data, the refined depth image and the text encoding features into the multimodal interactive fusion module, and input the obtained fusion features into the multimodal encoder-decoder to obtain the structural description data of the target workpiece.

[0082] Figure 8 is a structural diagram of a multimodal workpiece structure detail detection device based on depth optimization provided by an embodiment of the present invention, such as Figure 8 As shown, the multimodal workpiece structure detail detection device based on depth optimization can include the above Figure 3 The multimodal workpiece structure detail detection device based on depth optimization is shown. Optionally, the multimodal workpiece structure detail detection device 810 based on depth optimization may include a first processor 2001 .

[0083] Optionally, the multimodal workpiece structural detail detection device 810 based on depth optimization may further include a memory 2002 and a transceiver 2003 .

[0084] The first processor 2001, the memory 2002 and the transceiver 2003 may be connected via a communication bus.

[0085] The following combination Figure 8 The components of the multimodal workpiece structural detail detection device 810 based on deep optimization are introduced in detail:

[0086] The first processor 2001 is the control center of the multimodal workpiece structural detail detection device 810 based on depth optimization, and can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 can be one or more central processing units (CPUs), application-specific integrated circuits (ASICs), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more microprocessors (digital signal processors, DSPs) or one or more field programmable gate arrays (FPGAs).

[0087] Optionally, the first processor 2001 may execute various functions of the multimodal workpiece structure detail detection device 810 based on depth optimization by running or executing a software program stored in the memory 2002 and calling data stored in the memory 2002 .

[0088] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 8 CPU0 and CPU1 are shown in FIG.

[0089] In a specific implementation, as an embodiment, the multimodal workpiece structure detail detection device 810 based on depth optimization may also include multiple processors, such as Figure 8 1 and 2. The first processor 2001 and the second processor 2004 are shown in FIG. Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). A processor herein can refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0090] The memory 2002 is used to store the software program for executing the solution of the present invention, and is controlled by the first processor 2001 for execution. The specific implementation method can refer to the above method embodiment and will not be repeated here.

[0091] Alternatively, the memory 2002 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, a random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, an optical disc storage (including a compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and capable of being accessed by a computer, but not limited thereto. The memory 2002 may be integrated with the first processor 2001 or exist independently and be accessed through the interface circuit ( Figure 8 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.

[0092] The transceiver 2003 is used to communicate with a network device or a terminal device.

[0093] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 8The receiver is used to implement a receiving function, and the transmitter is used to implement a sending function.

[0094] Optionally, the transceiver 2003 may be integrated with the first processor 2001 or may exist independently and be connected to the first processor 2001 through the interface circuit ( Figure 8 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.

[0095] It should be noted that Figure 8 The structure of the multimodal workpiece structural detail detection device 810 based on depth optimization shown in the figure does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0096] In addition, the technical effects of the multimodal workpiece structure detail detection device 810 based on depth optimization can refer to the technical effects of the multimodal workpiece structure detail detection method based on depth optimization described in the above method embodiment, and will not be repeated here.

[0097] It should be understood that the first processor 2001 in the embodiment of the present invention may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor, or the processor may be any conventional processor, etc.

[0098] It should also be understood that the memory in the embodiments of the present invention may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0099] The above embodiments can be implemented in whole or in part via software, hardware (e.g., circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product comprises one or more computer instructions or computer programs. When loaded or executed on a computer, the processes or functions described in accordance with the embodiments of the present invention are fully or partially performed. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired means (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0100] It should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " as used herein generally indicates an "or" relationship between the associated objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.

[0101] In this disclosure, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, "at least one of a, b, or c" can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.

[0102] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0103] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0104] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0105] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interface, indirect coupling or communication connection of the device or unit, which can be electrical, mechanical or other forms.

[0106] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0107] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0108] If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical disks.

[0109] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A multimodal workpiece structural detail detection method based on depth optimization, characterized in that: The method comprises: S1. Acquire training data and train an initial artifact structure description network to obtain a trained artifact structure description network, wherein the artifact structure description network includes an artifact edge segmentation network, an artifact depth refinement network, a multimodal interactive fusion module, a multimodal encoder / decoder, a language-to-text conversion module, and a text encoder; S2, obtaining the image to be processed, the depth image and the text description of the target workpiece; S3, inputting the image to be processed into the workpiece edge segmentation network to obtain preliminary contour data; S4, inputting the depth image into the workpiece depth refinement network to obtain a refined depth image; S5. Input the text description into the language-to-text conversion module, and input the obtained features into the text encoder to obtain text encoding features; S6. Input the preliminary contour data, refined depth image and text encoding features into the multimodal interactive fusion module, input the obtained fusion features into the multimodal encoder-decoder to obtain the structural description data of the target workpiece.

2. The multimodal workpiece structural detail detection method based on depth optimization according to claim 1 is characterized in that: The S3 inputs the image to be processed into the workpiece edge segmentation network to obtain preliminary contour data, including: S31. Use an intelligent boundary detector to perform a comprehensive scan of the image to be processed and locate the boundary range of a specific object in the image; S32, within the determined boundary range of a specific object, segmentation is performed based on primitive parser primitives, effectively limiting the query to interact only with object features of a fixed length; S33, by introducing a multi-endpoint primitive representation, parsing the primitives composed of multiple endpoints, and generating multi-endpoint location queries and structure queries; S34, using a dynamic feature adaptation strategy to dynamically update the query location embedding according to the real-time changes in the object features; S35. Based on the primitive sequence decoder and output results, it accepts queries and position queries, fuses the queries and query position information, obtains the positions and scores of multiple primitive position lines through the feedforward fully connected layer, directly infers the combined output of each primitive, and generates a preliminary contour model of the object.

3. The multimodal workpiece structural detail detection method based on depth optimization according to claim 1, characterized in that: The step S4 inputs the depth image into the workpiece depth refinement network to obtain a refined depth image, including: S41, applying depth optimization enhancement technology to finely optimize the depth image containing specific objects; S42. Generate low-noise deep edge pseudo labels through self-learning distillation technology; S43, using edge-guided gradient correction mechanism and edge fusion smoothing strategy to optimize the deep optimization network; S44. Adopt a coarse-to-fine optimization strategy, first control the overall spatial structure and depth balance of the depth image from a macroscopic perspective, and then gradually refine the detailed information of the high-frequency area to obtain a refined depth image.

4. The multimodal workpiece structural detail detection method based on depth optimization according to claim 3 is characterized in that: The step S5 inputs the text description into the language-to-text conversion module, and inputs the obtained features into the text encoder to obtain text encoding features, including: Using natural language processing technology, text descriptions related to artifacts are extracted and analyzed to identify artifact features, important parameters, and key attributes.

5. The multimodal workpiece structural detail detection method based on depth optimization according to claim 4 is characterized in that: The step S6 inputs the preliminary contour data, the refined depth image, and the text encoding features into a multimodal interactive fusion module, and inputs the obtained fusion features into a multimodal encoder / decoder to obtain the structural description data of the target workpiece, including: S61, input the preliminary contour data, the refined depth image and the text encoding features into the cross attention layer to generate a new feature representation, and the features processed by the cross attention layer enter the Softmax layer; S62. Convert the feature into a probability vector through the Softmax layer and find the text description index number corresponding to the position with the highest probability; S63. Search and match corresponding character units in the character unit sequence using the index number vector, and concatenate the found character units to form a representative description. S64: The concatenated representative description is input into the large language model for further processing and generation.

6. A multimodal workpiece structure detail detection device based on depth optimization, wherein the multimodal workpiece structure detail detection device based on depth optimization is used to implement the multimodal workpiece structure detail detection method based on depth optimization according to any one of claims 1 to 5, characterized in that: The device comprises: A training unit is used to obtain training data and train the initial artifact structure description network to obtain a trained artifact structure description network, wherein the artifact structure description network includes an artifact edge segmentation network, an artifact depth refinement network, a multimodal interactive fusion module, a multimodal encoder-decoder, a language-to-text conversion module, and a text encoder; An acquisition unit, configured to acquire an image to be processed, a depth image, and a text description of a target workpiece; A segmentation unit, used to input the image to be processed into the workpiece edge segmentation network to obtain preliminary contour data; A refinement unit, configured to input the depth image into a workpiece depth refinement network to obtain a refined depth image; The encoding unit is used to input the text description into the language-to-text conversion module, and input the obtained features into the text encoder to obtain text encoding features; The fusion unit is used to input the preliminary contour data, the refined depth image and the text encoding features into the multimodal interactive fusion module, and input the obtained fusion features into the multimodal encoder-decoder to obtain the structural description data of the target workpiece.

7. The multimodal workpiece structural detail detection device based on depth optimization according to claim 6, characterized in that: The segmentation unit is used to: S31. Use an intelligent boundary detector to perform a comprehensive scan of the image to be processed and locate the boundary range of a specific object in the image; S32, within the determined boundary range of a specific object, segmentation is performed based on primitive parser primitives, effectively limiting the query to interact only with object features of a fixed length; S33, by introducing a multi-endpoint primitive representation, parsing the primitives composed of multiple endpoints, and generating multi-endpoint location queries and structure queries; S34, using a dynamic feature adaptation strategy to dynamically update the query location embedding according to the real-time changes in the object features; S35. Based on the primitive sequence decoder and output results, it accepts queries and position queries, fuses the queries and query position information, obtains the positions and scores of multiple primitive position lines through the feedforward fully connected layer, directly infers the combined output of each primitive, and generates a preliminary contour model of the object.

8. The multimodal workpiece structural detail detection device based on depth optimization according to claim 6, characterized in that: The refinement unit is used to: S41, applying depth optimization enhancement technology to finely optimize the depth image containing specific objects; S42. Generate low-noise deep edge pseudo labels through self-learning distillation technology; S43, using edge-guided gradient correction mechanism and edge fusion smoothing strategy to optimize the deep optimization network; S44. Adopt a coarse-to-fine optimization strategy, first control the overall spatial structure and depth balance of the depth image from a macroscopic perspective, and then gradually refine the detailed information of the high-frequency area to obtain a refined depth image.

9. A multi-modal workpiece structure detail detection device based on depth optimization, characterized in that: The multimodal workpiece structure detail detection device based on depth optimization includes: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 5 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program code, which can be called by a processor to execute the method according to any one of claims 1 to 5.