Multi-modal weld joint identification method and device based on target large model

By obtaining multimodal data and inputting the target model for processing, the problems of high threshold setting cost and inaccurate identification in weld recognition are solved, and efficient and accurate weld recognition is achieved.

CN120431407AInactive Publication Date: 2025-08-05BEIJING XIAOYU INTELLISYS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510847254.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-08-05
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the existing welding technology, weld recognition requires manual threshold setting, which leads to high cost and inaccurate identification, making it difficult to efficiently identify welds in complex production environments.

Method used

Multimodal data (color information RGB image data and point cloud data) are obtained through the three-dimensional acquisition device, and then data fusion is carried out to the target model for vector mapping and multi-layer perceptron processing to obtain weld information.

Benefits of technology

It reduces the need to set thresholds for each scenario, reduces the cost of manual setup, improves the efficiency and accuracy of weld recognition, has a wider scope of application, and reduces the recognition time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431407A_ABST
    Figure CN120431407A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to a multi-modal weld joint recognition method and device based on a large target model. The method comprises the steps that multi-modal data of a to-be-welded area are obtained through a three-dimensional collecting device, and the multi-modal data comprise color information RGB image data and point cloud data; performing multi-modal data fusion processing on the color information RGB image data and the point cloud data to obtain fused multi-modal data; inputting the fused multi-modal data into a target large model for vector mapping processing and multi-layer perceptron processing, and obtaining classification information and mask information of the to-be-welded area; and first welding seam information of the to-be-welded area is determined according to the classification information and the mask information. By adopting the method, the welding seam identification cost can be reduced, and the welding seam identification accuracy can be improved while the welding seam identification efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technology, and in particular to a multimodal weld recognition method and device based on a large target model. Background Art

[0002] With the advancement of science and technology, robots are gradually appearing in all aspects of production and life. Among them, intelligent robotic welding is a development direction in welding manufacturing, which can reduce labor costs and is attracting increasing attention. Among them, identifying welds before welding to guide the robot for precise welding is a key technology in intelligent robotic welding. For example, weld seam identification can be achieved through processing methods such as point cloud segmentation (finding intersections), registration (aligning with predefined workpiece models), and laser line extraction of angles. However, this process requires setting a large number of thresholds, which need to be manually defined. This can increase welding costs, and manual settings can lead to inaccurate weld seam identification. Summary of the Invention

[0003] The present disclosure provides a multimodal weld recognition method and device based on a large target model, which can reduce the cost of weld recognition, improve the efficiency of weld recognition, and improve the accuracy of weld recognition. The technical solution of the present disclosure is as follows: According to a first aspect of an embodiment of the present disclosure, a multimodal weld recognition method based on a target large model is provided, comprising: Acquire multimodal data of the area to be welded by a three-dimensional acquisition device, wherein the multimodal data includes color information RGB image data and point cloud data; Performing multimodal data fusion processing on the color information RGB image data and the point cloud data to obtain fused multimodal data; Inputting the fused multimodal data into a target large model for vector mapping processing and multi-layer perceptron processing to obtain classification information and mask information of the area to be welded; The first weld seam information of the area to be welded is determined according to the classification information and the mask information.

[0004] According to some embodiments, performing multimodal data fusion processing on the color information RGB image data and the point cloud data to obtain fused multimodal data includes: Performing coordinate conversion processing on the point cloud data to obtain converted point cloud data; Obtaining a tensor of a preset dimension corresponding to the converted point cloud data; Performing an embedding operation on the color information RGB image data to obtain a first embedding vector; Performing an embedding operation on the tensor of the preset dimension to obtain a second embedding vector; The first embedding vector and the second embedding vector are concatenated to obtain fused multimodal data.

[0005] According to some embodiments, performing an embedding operation on the color information RGB image data to obtain a first embedding vector includes: Performing rearrangement processing and segmentation processing on the color information RGB image data to obtain an image block set corresponding to the color information RGB image data; Perform mapping and normalization processing on each image block in the image block set to obtain the first embedding vector.

[0006] According to some embodiments, inputting the fused multimodal data into a target large model for vector mapping processing and multi-layer perceptron processing to obtain classification information and mask information of the area to be welded includes: Using the target large model to encode the fused multimodal data and position embedding vector to obtain an encoded vector; Decoding the encoded vector using the target large model to obtain a decoded vector; Using a multilayer perceptron (MLP) module in the target large model to perform multilayer perceptron processing on the decoded vector to obtain classification information of the area to be welded; Using a pixel embedding module in the target large model to perform vector mapping processing on the decoded vector and the fused multimodal data to obtain a mapped vector; The classification information is processed using the multi-layer perceptron MLP module in the target large model to obtain mask information of the area to be welded.

[0007] According to some embodiments, the method further comprises: Obtaining first loss function information corresponding to the classification information; Obtaining second loss function information corresponding to the mask information; The first weld seam information of the area to be welded is adjusted according to the first loss function information and the second loss function information to obtain the second weld seam information of the area to be welded.

[0008] According to some embodiments, wherein the first weld information of the area to be welded includes a weld length, the method further comprises: When it is determined that the weld length is greater than the length threshold, obtaining third loss function information corresponding to the breakpoint in the first weld information of the area to be welded and fourth loss function information corresponding to the breakpoint distance in the first weld information of the area to be welded; The first weld seam information of the area to be welded is adjusted according to the third loss function information and the fourth loss function information to obtain the third weld seam information of the area to be welded.

[0009] According to some embodiments, the method further comprises: Obtain historical color information RGB image data and historical point cloud data; The historical color information RGB image data and the historical point cloud data are used to train the initial large model, and the target large model is obtained when the weld information output by the initial large model meets the information requirements or the training times of the initial large model reaches a threshold.

[0010] According to a second aspect of an embodiment of the present disclosure, a multimodal weld recognition device based on a target large model is provided, comprising: A data acquisition unit, configured to acquire multimodal data of the area to be welded through a three-dimensional acquisition device, wherein the multimodal data includes color information RGB image data and point cloud data; The data acquisition unit is configured to perform multimodal data fusion processing on the color information RGB image data and the point cloud data to acquire fused multimodal data; An information acquisition unit, configured to input the fused multimodal data into a target large model for vector mapping processing and multi-layer perceptron processing, and obtain classification information and mask information of the area to be welded; The information acquisition unit is further configured to determine first weld seam information of the area to be welded based on the classification information and the mask information.

[0011] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, including: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the multimodal weld recognition method based on the target large model described in any one of the aforementioned aspects.

[0012] According to a fourth aspect of an embodiment of the present disclosure, a storage medium is provided. When instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the multimodal weld recognition method based on a target large model described in any one of the aforementioned aspects.

[0013] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, including a computer program, which implements the method described in any one of the aforementioned aspects when executed by a processor.

[0014] The technical solutions provided by the embodiments of the present disclosure bring at least the following beneficial effects: In some or related embodiments, a three-dimensional acquisition device is used to obtain multimodal data of the area to be welded, wherein the multimodal data includes color information RGB image data and point cloud data; the color information RGB image data and the point cloud data are subjected to multimodal data fusion processing to obtain fused multimodal data; the fused multimodal data is input into a target large model for vector mapping processing and multi-layer perceptron processing to obtain classification information and mask information of the area to be welded; and first weld information of the area to be welded is determined based on the classification information and the mask information. Therefore, by using multimodal data in three-dimensional space, it is possible to reduce the need to set corresponding thresholds for each scenario, reduce the high cost and poor recognition accuracy of manually setting thresholds when the production environment is complex and the tooling is different, and reduce the long time required for point cloud segmentation. By using a large model to obtain weld information, the weld recognition time can be reduced, the restrictions on the use scenario are reduced, the scope of application of weld recognition can be expanded, and the weld recognition efficiency can be improved while improving the weld recognition accuracy.

[0015] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0017] Figure 1 This is a flow chart of a first multimodal weld recognition method based on a target large model provided by an embodiment of the present disclosure; Figure 2 is a flow chart of a second multimodal weld recognition method based on a target large model provided by an embodiment of the present disclosure; Figure 3 This is a schematic diagram showing an example of the structure of a large model provided by an embodiment of the present disclosure; Figure 4is a flow chart of a third multimodal weld recognition method based on a target large model provided by an embodiment of the present disclosure; Figure 5 This is a schematic diagram showing another example structure of a large model provided by an embodiment of the present disclosure; Figure 6 is a block diagram illustrating a multimodal weld recognition based on a target large model according to an exemplary embodiment; Figure 7 The figure is a schematic diagram showing an example of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0018] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0019] The present disclosure provides a multimodal weld recognition method, device, electronic device, and storage medium based on a large target model. In some embodiments, the terms "multimodal weld recognition method based on a large target model" and "information processing method" and "communication method" are interchangeable; the terms "multimodal weld recognition device based on a large target model" and "information processing device" and "communication device" are interchangeable; and the terms "information processing system" and "communication system" are interchangeable.

[0020] The embodiments of the present disclosure are not exhaustive and are merely illustrative of some embodiments, and are not intended to be a specific limitation on the scope of protection of the present disclosure. In the absence of contradiction, each step in a certain embodiment can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a certain embodiment can also be implemented as an independent embodiment, and the order of the steps in a certain embodiment can be arbitrarily exchanged. In addition, the optional implementation methods in a certain embodiment can be arbitrarily combined; in addition, the embodiments can be arbitrarily combined. For example, some or all steps of different embodiments can be arbitrarily combined, and a certain embodiment can be arbitrarily combined with the optional implementation methods of other embodiments.

[0021] In each embodiment of the present disclosure, unless otherwise specified or provided for by logic, the terms and / or descriptions between the embodiments are consistent and can be referenced by each other. The technical features in different embodiments can be combined to form a new embodiment based on their inherent logical relationships.

[0022] The terms used in the embodiments of the present disclosure are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure.

[0023] In the embodiments of the present disclosure, unless otherwise specified, elements expressed in the singular, such as "a", "an", "the", "above", "said", "the", "the", etc., may mean "one and only one", or "one or more", "at least one", etc. For example, when using articles such as "a", "an", "the" in English in translation, the noun following the article may be understood as a singular expression or a plural expression.

[0024] In the embodiments of the present disclosure, “plurality” refers to two or more.

[0025] In some embodiments, the terms “at least one,” “one or more,” “a plurality of,” “multiple,” etc. may be used interchangeably.

[0026] The prefixes such as "first" and "second" in the embodiments of the present disclosure are only used to distinguish different description objects and do not constitute any restriction on the position, order, priority, quantity or content of the description objects. For the statement of the description object, please refer to the description in the context of the claims or embodiments, and no unnecessary restriction should be constituted due to the use of prefixes. For example, if the description object is a "field", the ordinal number before the "field" in the "first field" and the "second field" does not limit the position or order between the "fields". "First" and "second" do not limit whether the "fields" they modify are in the same message, nor do they limit the order of the "first field" and the "second field". For another example, if the description object is a "level", the ordinal number before the "level" in the "first level" and the "second level" does not limit the priority between the "levels". For another example, the number of description objects is not limited by the ordinal number and can be one or more. Taking "first device" as an example, the number of "devices" can be one or more. In addition, the objects modified by different prefixes can be the same or different. For example, if the description object is "device", then the "first device" and the "second device" can be the same device or different devices, and their types can be the same or different; for another example, if the description object is "information", then the "first information" and the "second information" can be the same information or different information, and their contents can be the same or different.

[0027] In some embodiments, "terminal" or "terminal device" may be referred to as "user equipment (UE)", "user terminal", "mobile station (MS)", "mobile terminal (MT)", subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless electronic device (wireless communication device), remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, etc.

[0028] In some embodiments, data, information, etc. may be obtained with the user's consent.

[0029] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.

[0030] Figure 1 This is a flowchart of the first multi-modal weld recognition method based on a target large model provided by the embodiment of the present disclosure, such as Figure 1 As shown, the multimodal weld seam recognition method based on the target large model can be used in the scenario of workpiece weld seam recognition, including the following steps: In step S11, multimodal data of the area to be welded is acquired through a three-dimensional acquisition device, wherein the multimodal data includes color information RGB image data and point cloud data; In some embodiments, the execution subject of the embodiments of the present disclosure may be, for example, an electronic device. This electronic device is not specifically a fixed electronic device. For example, when the device identification changes, the electronic device may also change accordingly. For example, when the structure of the electronic device changes, the electronic device may also change accordingly. Among them, the electronic device that is the execution subject of the embodiments of the present disclosure may be, for example, a robot. For example, when the robot identification changes, the robot may also change accordingly.

[0031] In some embodiments, the area to be welded may be, for example, an area to be welded. The area to be welded is not specifically a fixed area. For example, if the area of the area to be welded changes, the area to be welded may also change accordingly. For example, if the shape of the area to be welded changes, the area to be welded may also change accordingly.

[0032] According to some embodiments, the 3D acquisition device may be, for example, a device that captures images of the area to be welded. The 3D acquisition device is not specifically a fixed device. For example, if the device identification of the 3D acquisition device changes, the 3D acquisition device may also change accordingly. For example, if the device composition of the 3D acquisition device changes, the 3D acquisition device may also change accordingly. The 3D acquisition device may be, for example, a three-dimensional (3D) camera.

[0033] In some embodiments, multimodal data may include data acquired through a variety of different forms or sensory channels. The multimodal data does not specifically refer to a fixed set of data. For example, when the data type corresponding to the multimodal data changes, the multimodal data may also change accordingly. For example, when the specific value corresponding to the multimodal data changes, the multimodal data may also change accordingly.

[0034] According to some embodiments, the color information RGB image data may be, for example, color information RGB image data acquired from the area to be welded. The method for acquiring the color information RGB image data is not limited. The color information RGB image data does not specifically refer to fixed data. For example, if the area to be welded changes, the color information RGB image data may also change accordingly. For example, if the three-dimensional acquisition device used to acquire the color information RGB image data changes, the color information RGB image data may also change accordingly.

[0035] According to some embodiments, point cloud data refers to a set of vectors in a three-dimensional coordinate system. The point cloud data of the embodiment of the present disclosure may be, for example, point cloud data obtained by scanning the area to be welded.

[0036] In some embodiments, multimodal data of the area to be welded is acquired through a three-dimensional acquisition device, wherein the multimodal data includes color information RGB image data and point cloud data.

[0037] In step S12, multimodal data fusion processing is performed on the color information RGB image data and the point cloud data to obtain fused multimodal data; According to some embodiments, multimodal data fusion processing may, for example, refer to integrating information from different types of sensors or data sources to provide more comprehensive and accurate decision-making and analysis capabilities. The multimodal data fusion processing of the disclosed embodiments may, for example, be the process of integrating color information (RGB) image data and point cloud data. This multimodal data fusion processing does not specifically refer to a fixed processing procedure. For example, when the processing module corresponding to the multimodal data fusion processing changes, the multimodal data fusion processing may also change accordingly.

[0038] In some embodiments, the fused multimodal data may be, for example, data obtained by fusing multimodal data. The name of the fused multimodal data is not limited. For example, when the multimodal data is referred to as first multimodal data, the fused multimodal data may be, for example, referred to as second multimodal data. The fused multimodal data does not specifically refer to a fixed data. For example, when the multimodal data fusion processing method changes, the fused multimodal data may also change accordingly. For example, when the multimodal data changes, the fused multimodal data may also change accordingly.

[0039] In some embodiments, multimodal data fusion processing may be performed on the color information RGB image data and the point cloud data to obtain fused multimodal data.

[0040] In step S13, the fused multimodal data is input into the target large model for vector mapping processing and multi-layer perceptron processing to obtain classification information and mask information of the area to be welded; In some embodiments, the target large model can be, for example, a model that has been trained to obtain classification information and mask information. The target large model is not specifically a fixed model. For example, when the model parameters of the target large model change, the target large model can also change accordingly. For example, when the model structure of the target large model changes, the target large model can also change accordingly.

[0041] In some embodiments, the vector mapping process can be, for example, the process of processing vectors by the pixel embedding module in the large model. The vector mapping process can, for example, map each pixel in an image to a low-dimensional vector space. The vector mapping process does not specifically refer to a fixed process. For example, when the mapping method changes, the vector mapping process can also change accordingly.

[0042] According to some embodiments, the multi-layer perceptron processing may be, for example, a process in which an MPL module in a large model processes a vector.

[0043] In some embodiments, the fused multimodal data can be input into a target large model for vector mapping processing and multi-layer perceptron processing to obtain classification information and mask information of the area to be welded.

[0044] The classification information can be used, for example, to indicate the classification to which the area to be welded belongs. This classification information is not fixed information. For example, if the area to be welded changes, the classification information may also change accordingly. For example, if the method for determining the classification information changes, the classification information may also change accordingly.

[0045] The mask information may be used, for example, to indicate information on operations to be performed on the area to be welded.

[0046] In step S14, first weld seam information of the area to be welded is determined according to the classification information and the mask information.

[0047] According to some embodiments, the first weld information may be, for example, weld information obtained based on the output of the target large model. The "first" in the first weld information is used to distinguish it from the remaining weld information and does not specifically refer to a fixed information. For example, if the area to be welded changes, the first weld information may also change accordingly.

[0048] In some embodiments, first weld seam information of the area to be welded may be determined based on the classification information and the mask information.

[0049] In some or related embodiments, a three-dimensional acquisition device is used to obtain multimodal data of the area to be welded, wherein the multimodal data includes color information RGB image data and point cloud data; the color information RGB image data and the point cloud data are subjected to multimodal data fusion processing to obtain fused multimodal data; the fused multimodal data is input into a target large model for vector mapping processing and multi-layer perceptron processing to obtain classification information and mask information of the area to be welded; and first weld information of the area to be welded is determined based on the classification information and mask information. Therefore, the multimodal data in three-dimensional space can be used to reduce the need to set corresponding thresholds for each scenario, reduce the high cost and poor recognition accuracy of manually setting thresholds when the production environment is complex and the tooling is different, and reduce the time required for point cloud segmentation. Using a large model to obtain weld information can reduce the weld recognition time, reduce the restrictions on the use scenario, and expand the scope of application of weld recognition. Moreover, the multimodal data is three-dimensional data and can reflect the information of the area to be welded in three-dimensional space, thereby reducing and improving the efficiency of weld recognition while improving the accuracy of weld recognition. In addition, multimodal data can be directly input into the large model without the need to identify the welding area or obtain the large model corresponding to the workpiece type, and no manual settings are required, which can reduce the weld identification steps and improve the accuracy of weld identification.

[0050] Figure 2 This is a flow chart of a second multimodal weld recognition method based on a target large model provided by an embodiment of the present disclosure, such as Figure 2 As shown, the multimodal weld seam recognition method based on the target large model can be used in the workpiece weld seam recognition scenario, including the following steps: In step S21, multimodal data of the area to be welded is acquired by a three-dimensional acquisition device, wherein the multimodal data includes color information RGB image data and point cloud data; Among them, the relevant description can be as above, and will not be repeated here.

[0051] According to some embodiments, Figure 3 This is a schematic diagram of an example structure of a large model provided by an embodiment of the present disclosure, such as Figure 3As shown, the multimodal data can be, for example, input data for a large target model. The multimodal data can include, for example, RGB image data and point cloud data. The RGB image data can be used to extract semantic information from the large target model, while the point cloud can be used to obtain spatial information from the large target model, thereby improving the robustness of the large model. For example, weld seam identification can be performed when there is dirt on the workpiece surface, or weld beads at different spatial locations can be distinguished. The large model can be constructed based on a temporal Transformer, and the backbone network can be, for example, a 50-layer residual network (Res50) pretrained on Common Objects in Context (COCO) instance segmentation. Multimodal fusion is achieved through designed patch embedding. Through a 6-layer encoder and a 9-layer decoder, pixel embedding, and an MLP module, the weld location and attributes are simultaneously output.

[0052] In step S22, coordinate conversion processing is performed on the point cloud data to obtain converted point cloud data; Among them, the relevant description can be as above, and will not be repeated here.

[0053] In some embodiments, the point cloud data may be subjected to camera coordinate conversion processing to transform the point cloud data into a camera coordinate system, for example, as shown in formula (1): (1) Where: R is a 3x3 rotation matrix, T is a 3x1 translation vector, P world is the coordinate of the 3D point 3x1 in the world coordinate system, P camera is the coordinate of the 3x1 3D point in the camera coordinate system.

[0054] In step S23, a tensor of a preset dimension corresponding to the converted point cloud data is obtained; Among them, the relevant description can be as above, and will not be repeated here.

[0055] According to some embodiments, for example, the x-axis, y-axis, and z-axis data of the point cloud data are flattened and then spliced to adjust them into a tensor of [b, c, H, W] dimensions, where missing parts can be padded with 0.

[0056] In step S24, an embedding operation is performed on the color information RGB image data to obtain a first embedding vector; Among them, the relevant description can be as above, and will not be repeated here.

[0057] In some embodiments, the first embedding vector may be, for example, an embedding vector associated with the color information RGB image data. The first in the first embedding vector is used to distinguish from the other embedding vectors and does not specifically refer to a fixed embedding vector.

[0058] According to some embodiments, performing an embedding operation on color information RGB image data to obtain a first embedding vector includes: Performing rearrangement processing and segmentation processing on the color information RGB image data to obtain an image block set corresponding to the color information RGB image data; Mapping and normalizing are performed on each image block in the image block set to obtain a first embedding vector. This improves the accuracy of obtaining the first embedding vector and the accuracy of obtaining weld information.

[0059] For example, the input tensor can be rearranged using the re-dimensioning (einops.rearrange) module, transforming it from shape [b, c, H, W] to [b, h, w, c*p1*p2], where h and w are the number of blocks after segmentation. Each image block is mapped to the embedding dimension embed_dim through a linear layer, and the embedding result is normalized.

[0060] In step S25, an embedding operation is performed on the tensor of the preset dimension to obtain a second embedding vector; Among them, the relevant description can be as above, and will not be repeated here.

[0061] According to some embodiments, performing an embedding operation on a tensor of a preset dimension to obtain a second embedding vector includes: Performing rearrangement and segmentation processing on the tensor of the preset dimension to obtain a set of image blocks corresponding to the tensor of the preset dimension; Mapping and normalizing each image block in the set of image blocks corresponding to the tensor of the preset dimension to obtain a second embedding vector. This can improve the accuracy of obtaining the first embedding vector and the accuracy of obtaining weld information.

[0062] In step S26, the first embedding vector and the second embedding vector are concatenated to obtain fused multimodal data; Among them, the relevant description can be as above, and will not be repeated here.

[0063] According to some embodiments, the embedding data may be concatenated (cat) together as hidden states and input into the model.

[0064] In step S27, the fused multimodal data is input into the target large model for vector mapping processing and multi-layer perceptron processing to obtain classification information and mask information of the area to be welded; Among them, the relevant description can be as above, and will not be repeated here.

[0065] According to some embodiments, the fused multimodal data is input into a target large model for vector mapping processing and multi-layer perceptron processing to obtain classification information and mask information of the area to be welded, including: The target large model is used to encode the fused multimodal data and position embedding vector to obtain the encoded vector; Decoding the encoded vector using the target large model to obtain a decoded vector; The multi-layer perceptron (MLP) module in the target large model is used to process the decoded vectors to obtain the classification information of the area to be welded; Use the pixel embedding module in the target large model to perform vector mapping on the decoded vector and the fused multimodal data to obtain the mapped vector; The multi-layer perceptron (MLP) module in the target large model is used to process the classification information and obtain the mask information of the area to be welded. This can improve the recognition effect of welds and can simultaneously output classification information classes and masks, improving the efficiency of weld recognition.

[0066] The encoding process may be performed by the encoding layer, for example, and the number of encoding layers may be 6 layers, for example; the decoding process may be performed by the decoding layer, for example, and the number of decoding layers may be 9 layers, for example.

[0067] According to some embodiments, Figure 4 This is a flowchart of a third multimodal weld recognition method based on a target large model provided by an embodiment of the present disclosure, such as Figure 4As shown in the figure, the pixel embedding module can be used to connect the original data with the memory data output by the encoder, and then the mask information can be obtained through an MLP module.

[0068] According to some embodiments, Figure 5 This is another schematic diagram of a large model structure provided by the embodiment of the present disclosure. Figure 5 As shown, the MLP module can also output class bounding box information, which is used to reduce the convergence time of large models and improve the convergence speed of large models.

[0069] In step S28, first weld seam information of the area to be welded is determined according to the classification information and the mask information.

[0070] Among them, the relevant description can be as above, and will not be repeated here.

[0071] According to some embodiments, the method further comprises: Obtain first loss function information corresponding to the classification information; Obtaining the second loss function information corresponding to the mask information; The first weld seam information of the area to be welded is adjusted according to the first loss function information and the second loss function information to obtain the second weld seam information of the area to be welded.

[0072] According to some embodiments, wherein the first weld seam information of the area to be welded includes the weld seam length, the method further comprises: When it is determined that the weld length is greater than the length threshold, obtaining third loss function information corresponding to the breakpoint in the first weld information of the area to be welded and fourth loss function information corresponding to the breakpoint distance in the first weld information of the area to be welded; The first weld seam information of the area to be welded is adjusted according to the third loss function information and the fourth loss function information to obtain the third weld seam information of the area to be welded.

[0073] According to some embodiments, the method further comprises: Obtain historical color information RGB image data and historical point cloud data; The initial large model is trained using historical color information RGB image data and historical point cloud data. When the weld information output by the initial large model meets the information requirements or the initial large model training times reach a threshold, the target large model is acquired. Therefore, the model can be trained using historical multimodal data, improving the accuracy of large model acquisition and weld recognition.

[0074] Among them, the first loss function corresponding to the classification classes can be, for example, Focal Loss, the loss function corresponding to the bounding boxes can be, for example, IoU Loss, and the second loss function information corresponding to the masks can be, for example, FocalLoss and Dice Loss, wherein specifically it can be: (2) Among them, pt is the predicted probability of the sample, that is, the historical multimodal data, αt is used to balance the weights of positive and negative samples, and γ controls the weights of difficult-to-classify samples.

[0075] (3) Among them, A1 and B1 are the bounding boxes of the predicted value pred and the true value (Ground Truth, gt).

[0076] (4) Among them, A2 and B2 are the masks of pred and gt respectively.

[0077] According to some embodiments, when the weld bead length is greater than the expected length, in order to improve the continuity of the weld trajectory, a path continuity loss can be set: The breakpoint Loss can be expressed as follows: (5) The break point distance Loss can be expressed as formula (6): (6) Where |C| represents the number of connected components, Ci and Ci+1 represent two adjacent connected components, and ||x−y| is the Euclidean distance between two breakpoints.

[0078] In one or related embodiments, a tensor of a preset dimension corresponding to the converted point cloud data is obtained; an embedding operation is performed on the color information RGB image data to obtain a first embedding embedding vector; an embedding operation is performed on the tensor of the preset dimension to obtain a second embedding embedding vector; the first embedding embedding vector and the second embedding embedding vector are concatenated to obtain fused multimodal data; therefore, the fused multimodal data can be obtained through the embedding vector, thereby improving the accuracy of multimodal data acquisition and improving the accuracy of weld recognition.

[0079] A block diagram of a multi-modal weld seam recognition device based on a target large model is shown according to an exemplary embodiment. Figure 6, the apparatus 600 comprises: The data acquisition unit 601 is used to acquire multimodal data of the area to be welded through a three-dimensional acquisition device, wherein the multimodal data includes color information RGB image data and point cloud data; The data acquisition unit 601 is used to perform multimodal data fusion processing on the color information RGB image data and the point cloud data to obtain fused multimodal data; The information acquisition unit 602 is used to input the fused multimodal data into the target large model for vector mapping processing and multi-layer perceptron processing to obtain classification information and mask information of the area to be welded; The information acquisition unit 602 is further configured to determine first weld seam information of the area to be welded according to the classification information and the mask information.

[0080] According to some embodiments, the data acquisition unit 601 is configured to perform multimodal data fusion processing on the color information RGB image data and the point cloud data, and to obtain the fused multimodal data, specifically for: Perform coordinate conversion on the point cloud data to obtain the converted point cloud data; Get the tensor of preset dimensions corresponding to the converted point cloud data; Perform an embedding operation on the color information RGB image data to obtain a first embedding vector; Perform embedding operation on the tensor of preset dimension to obtain the second embedding vector; The first embedding vector and the second embedding vector are concatenated to obtain fused multimodal data.

[0081] According to some embodiments, the data acquisition unit 601 is configured to perform an embedding operation on the color information RGB image data, and to obtain a first embedding vector, specifically to: Performing rearrangement processing and segmentation processing on the color information RGB image data to obtain an image block set corresponding to the color information RGB image data; Mapping and normalization are performed on each image block in the image block set to obtain the first embedding vector.

[0082] According to some embodiments, the information acquisition unit 602 is used to input the fused multimodal data into the target large model for vector mapping processing and multi-layer perceptron processing, and to obtain classification information and mask information of the area to be welded, specifically for: The target large model is used to encode the fused multimodal data and position embedding vector to obtain the encoded vector; Decoding the encoded vector using the target large model to obtain a decoded vector; The multi-layer perceptron (MLP) module in the target large model is used to process the decoded vectors to obtain the classification information of the area to be welded; Use the pixel embedding module in the target large model to perform vector mapping on the decoded vector and the fused multimodal data to obtain the mapped vector; The multi-layer perceptron (MLP) module in the target large model is used to classify information and process it to obtain the mask information of the area to be welded.

[0083] According to some embodiments, the information acquisition unit 602 is further configured to: Obtain first loss function information corresponding to the classification information; Obtaining the second loss function information corresponding to the mask information; The first weld seam information of the area to be welded is adjusted according to the first loss function information and the second loss function information to obtain the second weld seam information of the area to be welded.

[0084] According to some embodiments, the first weld information of the area to be welded includes the weld length, and the information acquisition unit 702 is further configured to: When it is determined that the weld length is greater than the length threshold, obtaining third loss function information corresponding to the breakpoint in the first weld information of the area to be welded and fourth loss function information corresponding to the breakpoint distance in the first weld information of the area to be welded; The first weld seam information of the area to be welded is adjusted according to the third loss function information and the fourth loss function information to obtain the third weld seam information of the area to be welded.

[0085] According to some embodiments, the method information obtaining unit 602 is further configured to: Obtain historical color information RGB image data and historical point cloud data; The initial large model is trained using historical color information RGB image data and historical point cloud data. When the weld information output by the initial large model meets the information requirements or the training times of the initial large model reaches a threshold, the target large model is obtained.

[0086] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0087] In some or related embodiments, a data acquisition unit is used to acquire multimodal data of the area to be welded through a three-dimensional acquisition device, wherein the multimodal data includes color information RGB image data and point cloud data; the data acquisition unit is used to perform multimodal data fusion processing on the color information RGB image data and the point cloud data to obtain fused multimodal data; the information acquisition unit is used to input the fused multimodal data into a target large model for vector mapping processing and multi-layer perceptron processing to obtain classification information and mask information of the area to be welded; the information acquisition unit is further used to determine the first weld information of the area to be welded based on the classification information and mask information. Therefore, by using multimodal data in three-dimensional space, it is possible to reduce the need to set corresponding thresholds for each scenario, reduce the high cost and poor recognition accuracy of manually setting thresholds when the production environment is complex and the tooling is different, and reduce the long time required by point cloud segmentation. By using a large model to acquire weld information, it is possible to reduce the weld recognition time and reduce the restrictions on the use scenario, thereby expanding the scope of application of weld recognition, reducing the efficiency of weld recognition and improving the accuracy of weld recognition.

[0088] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device 700 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0089] like Figure 7 As shown, electronic device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. Various programs and data required for the operation of electronic device 700 may also be stored in RAM 703. Computing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to bus 704.

[0090] Multiple components in the electronic device 700 are connected to the I / O interface 705, including an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the electronic device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0091] The computing unit 701 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above. For example, in some embodiments, the above-described methods may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the above-described methods may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to perform the above-described methods via any other suitable means (e.g., via firmware).

[0092] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0093] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0094] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0095] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0096] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.

[0097] A computer system may include a client and a server. The client and server are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a host product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical servers and VPS services ("Virtual Private Servers" or "VPS"). The server may also be a server in a distributed system or a server integrated with blockchain.

[0098] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0099] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A multimodal weld recognition method based on a large target model, characterized in that: include: Acquire multimodal data of the area to be welded by a three-dimensional acquisition device, wherein the multimodal data includes color information RGB image data and point cloud data; Performing multimodal data fusion processing on the color information RGB image data and the point cloud data to obtain fused multimodal data; Inputting the fused multimodal data into a target large model for vector mapping processing and multi-layer perceptron processing to obtain classification information and mask information of the area to be welded; The first weld seam information of the area to be welded is determined according to the classification information and the mask information.

2. The method according to claim 1, characterized in that The performing multimodal data fusion processing on the color information RGB image data and the point cloud data to obtain fused multimodal data includes: Performing coordinate conversion processing on the point cloud data to obtain converted point cloud data; Obtaining a tensor of a preset dimension corresponding to the converted point cloud data; Performing an embedding operation on the color information RGB image data to obtain a first embedding vector; Performing an embedding operation on the tensor of the preset dimension to obtain a second embedding vector; The first embedding vector and the second embedding vector are concatenated to obtain fused multimodal data.

3. The method according to claim 2, characterized in that The performing an embedding operation on the color information RGB image data to obtain a first embedding vector includes: Performing rearrangement processing and segmentation processing on the color information RGB image data to obtain an image block set corresponding to the color information RGB image data; Perform mapping and normalization processing on each image block in the image block set to obtain the first embedding vector.

4. The method according to claim 1, wherein The fused multimodal data is input into the target large model for vector mapping processing and multi-layer perceptron processing to obtain classification information and mask information of the area to be welded, including: Using the target large model to encode the fused multimodal data and position embedding vector to obtain an encoded vector; Decoding the encoded vector using the target large model to obtain a decoded vector; Using the multi-layer perceptron MLP module in the target large model to perform multi-layer perceptron processing on the decoded vector to obtain classification information of the area to be welded; Using a pixel embedding module in the target large model to perform vector mapping processing on the decoded vector and the fused multimodal data to obtain a mapped vector; The classification information is processed using the multi-layer perceptron MLP module in the target large model to obtain mask information of the area to be welded.

5. The method according to claim 1, wherein The method further comprises: Obtaining first loss function information corresponding to the classification information; Obtaining second loss function information corresponding to the mask information; The first weld seam information of the area to be welded is adjusted according to the first loss function information and the second loss function information to obtain the second weld seam information of the area to be welded.

6. The method according to claim 5, characterized in that in, The first weld seam information of the area to be welded includes a weld seam length, and the method further includes: When it is determined that the weld length is greater than the length threshold, obtaining third loss function information corresponding to the breakpoint in the first weld information of the area to be welded and fourth loss function information corresponding to the breakpoint distance in the first weld information of the area to be welded; The first weld seam information of the area to be welded is adjusted according to the third loss function information and the fourth loss function information to obtain the third weld seam information of the area to be welded.

7. The method according to claim 1, characterized in that The method further comprises: Obtain historical color information RGB image data and historical point cloud data; The historical color information RGB image data and the historical point cloud data are used to train the initial large model, and the target large model is obtained when the weld information output by the initial large model meets the information requirements or the training times of the initial large model reaches a threshold.

8. A multimodal weld recognition device based on a large target model, characterized in that: include: A data acquisition unit, configured to acquire multimodal data of the area to be welded through a three-dimensional acquisition device, wherein the multimodal data includes color information RGB image data and point cloud data; The data acquisition unit is configured to perform multimodal data fusion processing on the color information RGB image data and the point cloud data to acquire fused multimodal data; An information acquisition unit, configured to input the fused multimodal data into a target large model for vector mapping processing and multi-layer perceptron processing, and obtain classification information and mask information of the area to be welded; The information acquisition unit is further configured to determine first weld seam information of the area to be welded based on the classification information and the mask information.

9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the multimodal weld recognition method based on a target large model as claimed in any one of claims 1 to 7.

10. A storage medium storing instructions, characterized in that: When the instructions are executed on an electronic device, the electronic device is enabled to execute the multimodal weld recognition method based on a target large model according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Target detection method and device based on large model and electronic equipment

    CN119339066A

  • Weld joint identification method and device fused with large model, electronic equipment and storage medium

    CN119407392A

  • Weld joint identifying and tracking method based on multi-modal information, automatic welding device and computer equipment

    CN119681385A