Target device detection method, device, equipment, storage medium and program product

CN122530615APending Publication Date: 2026-08-07NUCTECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NUCTECH CO LTD
Filing Date
2026-05-20
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

但是,这种检测方式通常需要由人工对X射线成像得到的伪彩色图像进行审核检测,人工成本高,且检出率难以保证,存在漏检风险

Benefits of technology

[0016]在本公开的实施例中,从伪彩色图像提取轮廓形状信息后,基于提示词工程生成与伪彩色图像适配的动态文本提示词。利用多模态大模型对文本提示词、对象相关信息以及伪彩色图像进行处理,实现了视觉特征与文本引导的多模态协同输入,提高了对目标装置的整体检出率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122530615A_ABST
    Figure CN122530615A_ABST
Patent Text Reader

Abstract

The present disclosure provides a target device detection method, device, equipment, storage medium and program product, belongs to the technical field of image intelligent safety detection, and specifically relates to the technical field of computer vision. The method comprises the following steps: identifying object contour information of a plurality of objects from a pseudo-color image, wherein the pseudo-color image comprises a plurality of colors, the colors are used to represent object materials, and at least one object in the plurality of objects is a component used to constitute a target device; generating a text prompt word used to identify the target device based on the object contour information of the plurality of objects; and inputting the text prompt word, the object contour information of the plurality of objects and the pseudo-color image into a multi-modal large model to output a detection result for the target device, wherein the detection result is used to represent whether the plurality of objects can constitute the target device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of intelligent security detection technology for images, specifically to the field of computer vision technology, and more specifically to a method, apparatus, device, storage medium, and program product for detecting a target device. Background Technology

[0002] In security checks, X-rays are typically used to inspect luggage to determine if it contains dangerous items such as explosives. However, this method usually requires manual review of the pseudo-color images obtained from the X-ray imaging, which is costly and has a low detection rate, posing a risk of missed detections. Summary of the Invention

[0003] In view of this, the present disclosure provides a method, apparatus, device, storage medium, and program product for detecting a target device.

[0004] One aspect of this disclosure provides a method for detecting a target device, comprising: identifying object contour information of multiple objects from a pseudo-color image, wherein the pseudo-color image includes multiple colors, the colors being used to characterize the material of the objects, and at least one of the multiple objects being a component for assembling a target device; generating text prompts for identifying the target device based on the object contour information of the multiple objects; and inputting the text prompts, the object contour information of the multiple objects, and the pseudo-color image into a multimodal large model, and outputting a detection result for the target device, wherein the detection result characterizes whether the multiple objects can assemble into a target device.

[0005] According to embodiments of this disclosure, the text prompts include rewritten prompts for each object, and the object contour information includes contour shape information and contour position information; generating text prompts for identifying a target device based on the respective object contour information of multiple objects includes: acquiring prior knowledge information for a pseudo-color image, wherein the prior knowledge information is used to indicate the mapping relationship between object material and color; determining, for each object, the object color matching each contour position information from the pseudo-color image; and generating rewritten prompts for rewriting contour shape information based on the prior knowledge information and the object color.

[0006] According to embodiments of this disclosure, a text prompt for rewriting contour shape information is generated based on prior knowledge information and object color, including: determining the object material corresponding to the object color based on prior knowledge information; determining a target component that matches the object material and contour shape information; and generating a rewriting prompt for rewriting the contour shape information into the target component, wherein the target component is used to compose a target device.

[0007] According to embodiments of this disclosure, the method for detecting a target device further includes: acquiring multiple target components corresponding to multiple objects respectively; among the multiple target components, there are multiple predetermined components in a predetermined component combination, and the contour position information corresponding to each of the multiple predetermined components satisfies a predetermined positional relationship; generating a guiding prompt word corresponding to the predetermined component combination, wherein the text prompt word includes a guiding prompt word, and the predetermined component combination corresponds to the triggering method of the target device.

[0008] According to embodiments of this disclosure, the text prompts further include task prompts for guiding the multimodal large model to perform classification tasks; inputting the text prompts and the object contour information of each of the multiple objects into the multimodal large model, and outputting detection results for the target device, includes: when the text prompts include rewritten prompts, using the multimodal large model, under the guidance of the rewritten prompts, rewriting the contour shape information of each of the multiple objects into target parts; and under the guidance of the task prompts, generating detection results based on the target parts of each of the multiple objects, contour position information, and pseudo-color images; when the text prompts include both rewritten prompts and guiding prompts, using the multimodal large model, under the guidance of the rewritten prompts, rewriting the contour shape information of each of the multiple objects into target parts; and under the guidance of the task prompts, generating detection results that match the semantic rules represented by the guiding prompts based on the target parts of each of the multiple objects, contour position information, and pseudo-color images.

[0009] According to embodiments of this disclosure, the method for detecting a target device further includes: obtaining object image features of multiple objects from a pseudo-color image, and replacing the pseudo-color image with the multiple object image features to input a multimodal model, so that the multimodal model generates a detection result for the target device based on text prompts, object contour information of multiple objects, and image features.

[0010] According to embodiments of this disclosure, a multimodal large model is fine-tuned in the following manner: multiple fine-tuning datasets are acquired, each fine-tuning dataset including sample pseudo-color images, image labels, outline information of each labeled object in the sample pseudo-color images, and sample text prompts; wherein the sample text prompts are generated based on the outline information of multiple labeled objects; multiple first sample pseudo-color images in the multiple fine-tuning datasets include multiple styles for composing the same component of the target device; the multimodal large model to be fine-tuned is fine-tuned using the multiple fine-tuning datasets, and a multimodal large model is obtained under predetermined conditions.

[0011] According to embodiments of this disclosure, the predetermined condition includes the convergence of the loss function, which includes: a loss term for measuring the classification difference between the sample detection results output by the multimodal large model and the image labels; and a loss term for measuring the difference between the first relative position between multiple labeled objects determined by the multimodal large model from the sample pseudo-color image and the second relative position between the labeled contour position information in the contour information of the multiple labeled objects.

[0012] Another aspect of this disclosure provides a target device detection apparatus, comprising: an identification module for identifying object contour information of multiple objects from a pseudo-color image, wherein the pseudo-color image includes multiple colors, the colors being used to characterize the material of the objects, and at least one of the multiple objects being a component for assembling a target device; a generation module for generating text prompts for identifying the target device based on the object contour information of the multiple objects; and a detection module for inputting the text prompts, the object contour information of the multiple objects, and the pseudo-color image into a multimodal large model, and outputting a detection result for the target device, wherein the detection result characterizes whether the multiple objects can assemble into a target device.

[0013] Another aspect of this disclosure provides an electronic device comprising: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the methods described above.

[0014] Another aspect of this disclosure provides a computer-readable storage medium storing computer-executable instructions that, when executed, are used to implement the methods described above.

[0015] Another aspect of this disclosure provides a computer program product including computer-executable instructions that, when executed, are used to implement the methods described above.

[0016] In the embodiments of this disclosure, after extracting contour shape information from the pseudo-color image, dynamic text prompts adapted to the pseudo-color image are generated based on prompt word engineering. By utilizing a multimodal large model to process the text prompts, object-related information, and the pseudo-color image, multimodal collaborative input of visual features and text guidance is achieved, improving the overall detection rate of the target device. Attached Figure Description

[0017] The above and other objects, features and advantages of this disclosure will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0018] Figure 1 An exemplary system architecture for a detection method of a target device that can be applied according to embodiments of the present disclosure is illustrated.

[0019] Figure 2 A flowchart illustrating a detection method for a target device according to an embodiment of the present disclosure is shown schematically.

[0020] Figure 3 A schematic diagram illustrating a predetermined combination of components determined according to an embodiment of the present disclosure is shown.

[0021] Figure 4 A schematic diagram illustrating the detection results output according to an embodiment of the present disclosure is shown.

[0022] Figure 5 A block diagram of a detection apparatus for a target device according to an embodiment of the present disclosure is shown schematically.

[0023] Figure 6 A block diagram of an electronic device suitable for implementing a detection method for a target device according to an embodiment of the present disclosure is shown schematically. Detailed Implementation

[0024] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.

[0025] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0026] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0027] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0028] In the embodiments disclosed herein, the collection, updating, analysis, processing, use, transmission, provision, disclosure, and storage of data (e.g., including but not limited to user personal information) comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures have been taken to prevent unauthorized access to user personal information data and to safeguard user personal information security, network security, and national security.

[0029] In the embodiments disclosed herein, user authorization or consent is obtained before acquiring or collecting user personal information.

[0030] In related technologies, image-based deep learning object detection models can be used to detect hazardous materials. However, the diverse appearances of hazardous materials such as explosives make it difficult for conventional object detection models to fully learn all types of hazardous materials, thus failing to guarantee a high detection rate.

[0031] In view of the above, embodiments of this disclosure provide a method for detecting a target device, comprising: identifying object contour information of multiple objects from a pseudo-color image, wherein the pseudo-color image includes multiple colors, the colors are used to characterize the material of the objects, and at least one of the multiple objects is a component for assembling a target device; generating text prompts for identifying the target device based on the object contour information of the multiple objects; and inputting the text prompts, the object contour information of the multiple objects, and the pseudo-color image into a multimodal large model, and outputting a detection result for the target device, wherein the detection result is used to characterize whether the multiple objects can assemble into a target device.

[0032] Figure 1 An exemplary system architecture for a detection method of a target device that can be applied according to embodiments of the present disclosure is illustrated.

[0033] It is important to note that Figure 1 The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.

[0034] like Figure 1 As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, an image acquisition unit 104, a server 105, and a network 106. The network 106 serves as a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, the image acquisition unit 104, and the server 105. The network 106 may include various connection types, such as wired and / or wireless communication links, etc.

[0035] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 106 to receive or send messages, etc. The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be equipped with a communication client application for acquiring and displaying image information processed by the image acquisition unit 104 and the server 105.

[0036] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0037] The image acquisition unit 104 can be used to acquire a pseudo-color image of the target to be detected and send the pseudo-color image to the server 105 via the network 106. The image acquisition unit 104 can be an X-ray based imaging device.

[0038] Server 105 can be a server that provides various services, such as analyzing and processing data such as pseudo-color images uploaded by image acquisition unit 104, and feeding back the processing results to terminal devices.

[0039] It should be noted that the target device detection method provided in this embodiment can generally be executed by server 105. Correspondingly, the target device detection device provided in this embodiment can generally be located in server 105. The target device detection method provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the image acquisition unit 104 and / or server 105. Correspondingly, the target device detection device provided in this embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the image acquisition unit 104 and / or server 105. Alternatively, the target device detection method provided in this embodiment can also be executed by the first terminal device 101, the second terminal device 102, and the third terminal device 103, or by other terminal devices different from the first terminal device 101, the second terminal device 102, and the third terminal device 103. Accordingly, the detection device for the target device provided in the embodiments of this disclosure may also be disposed in the first terminal device 101, the second terminal device 102, and the third terminal device 103, or disposed in other terminal devices different from the first terminal device 101, the second terminal device 102, and the third terminal device 103.

[0040] It should be understood that Figure 1 The number of terminal devices, networks, image acquisition units, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0041] Figure 2 A flowchart illustrating a detection method for a target device according to an embodiment of the present disclosure is shown schematically.

[0042] like Figure 2 As shown, the method includes operations S210~S230.

[0043] In operation S210, the object contour information of each of the multiple objects is identified from the pseudo-color image.

[0044] In operation S220, text prompts for identifying the target device are generated based on the object contour information of each of the multiple objects.

[0045] In operation S230, text prompts, object contour information of multiple objects, and pseudo-color images are input into the multimodal large model, and the detection results for the target device are output.

[0046] Pseudo-color images can be acquired using X-ray equipment. For example, an X-ray device sends X-rays from one side towards a target and receives the X-rays passing through the target on the other side, then images the received X-rays to obtain a pseudo-color image. It should be noted that pseudo-color images can also be acquired using other imaging methods that can map different materials to different colors, and are not limited to X-ray imaging.

[0047] A pseudo-color image can include multiple colors, which are used to characterize the material of an object. The pseudo-color image can include multiple objects, at least one of which is a component used to compose a target device. The target device can refer to a device that needs to be detected in a security inspection scenario, and may include explosives, etc.

[0048] A pseudo-color image can be input into an image encoder for recognition. After recognition, multiple objects can be labeled using bounding boxes. A pseudo-color image is not merely about color; it also contains information such as object shape and boundaries. Therefore, the image encoder can perform connected component analysis on the binarized pseudo-color image to identify and label multiple objects from it.

[0049] Object contour information is used to indicate information related to the object contour identified from the pseudo-color image. This information may include contour shape information (representing the contour shape) and contour position information (representing the object contour position). It is understood that the connected component analysis described above can distinguish the object from other background information in the pseudo-color image. Therefore, the object contour information can be directly obtained from the object boundary obtained after connected component analysis. For example, if object P is a rectangle, the contour shape information in the object contour information can be "rectangle," and the contour position information can be determined based on the object's bounding box. For instance, if the center position of the bounding box of object P is marked as (x0, y0), and the length and width are L and W respectively, then the contour position information can be the center position of the bounding box (x0, y0), and the length L and width W.

[0050] It is understandable that multiple objects with the same shape can be located in multiple positions within a pseudocolor image. By using contour shape information and contour position information, an object with a specific contour and its exact location within the pseudocolor image can be uniquely identified.

[0051] After identifying multiple objects in a pseudo-color image, a single object can be uniquely located based on its individual outline shape and position information. Then, the colors at multiple locations in the pseudo-color image can be uniquely associated with the located object to obtain its color, and the object's material can be further determined based on the color. These steps can be understood as mapping the outline shape, outline position, and material (or color) of objects in a pseudo-color image. For example, if a pseudo-color image contains rectangular objects in both the upper left and lower right corners, the object with a blue rectangle in the upper left corner and the object with a red rectangle in the lower right corner can be uniquely identified through the association of their outline shape, outline position, and color (material).

[0052] Based on the object's outline information and material, it is possible to make a preliminary judgment on each object, determine the content or area related to the target device that needs to be focused on in the pseudo-color image, and generate prompt words to prompt the multimodal large model to focus on the aforementioned content or area related to the target device.

[0053] Multimodal large models are artificial intelligence models that can process, understand, and generate multiple types of information. Multimodal large models can project multiple input information from various information modalities into a shared and high-dimensional semantic space. Within this semantic space, multiple input information is fused and then subjected to subsequent semantic processing, and output information of different modalities is obtained as needed.

[0054] The text prompts, the outline information of each of the multiple objects, and pseudo-color images are input into a multimodal large model. Based on the text prompts, the multimodal large model locates the objects of interest from the pseudo-color images using their outline information. Further analysis of these objects yields detection results. These detection results characterize whether the multiple objects can form a target device.

[0055] According to embodiments of this disclosure, after extracting object contour information from a pseudo-color image, dynamic text prompts adapted to the pseudo-color image are generated based on prompt word engineering. By utilizing a multimodal large model to process the text prompts, object-related information, and the pseudo-color image, multimodal collaborative input of visual features and text guidance is achieved, improving the overall detection rate of the target device.

[0056] According to embodiments of this disclosure, the text prompts include rewritten prompts for each object, and the object contour information includes contour shape information and contour position information; generating text prompts for identifying a target device based on the respective object contour information of multiple objects includes: acquiring prior knowledge information for a pseudo-color image, wherein the prior knowledge information is used to indicate the mapping relationship between object material and color; determining, for each object, the object color matching each contour position information from the pseudo-color image; and generating rewritten prompts for rewriting contour shape information based on the prior knowledge information and the object color.

[0057] The mapping relationship between object material and color indicated by prior knowledge information for pseudo-color images can be preset according to the imaging principle of pseudo-color images.

[0058] For example, organic matter appears orange-red in a pseudo-color image, common metals appear blue, liquids appear light green, and heavy metals appear purple. The mapping relationship indicated by the prior knowledge information can be set as follows: orange-red corresponds to organic matter, blue corresponds to common metals, light green corresponds to liquids, and purple corresponds to heavy metals.

[0059] A pseudo-color image can contain multiple colors. These multiple colors can belong to a single object (e.g., an object with multiple materials) or multiple objects (e.g., an object with only one material). Judging solely by the colors in a pseudo-color image (e.g., treating a single color region as an object) can lead to misclassification. For example, an object B might be composed of multiple materials, such as a main material that is orange-red (corresponding to organic matter) and parts that are purple (e.g., parts inlaid with heavy metal components). In this case, if object segmentation is based solely on color, the organic matter and heavy metal might be identified as two separate objects, when in fact they are different components of the same object, resulting in misclassification.

[0060] In this embodiment of the disclosure, the contour location information can locate objects from the dimensions of their positions. Therefore, based on the contour location information determined from the pseudo-color image (such as marking the center position of the object's bounding box), the color covering the contour location information can be determined as the object color matching each contour location information, that is, the color of the object located at that contour location information. For example, the color of each object matching each object can be determined by the overlap between the color and the region of the contour location information.

[0061] For example, taking object B as an example again, the outline position information of object B includes: center position (x b y b The outline is 100 pixels long and 50 pixels wide. Within the area indicated by this outline location information, the main body is orange-red, and some areas are purple. Based on the relationship between color and position, orange-red and purple can be used as the object color to match object B. Alternatively, orange-red can be used as the object color to match object B (i.e., the color with the largest area within the location area). It is understandable that if an object C has only one material, the color falling within the location area can be directly used as the object color to match that object.

[0062] For each object, the color of its contour position information can be determined from the pseudo-color image using the above method, thus obtaining the object's color. Subsequently, based on the object color and the mapping relationship indicated by prior knowledge information, the material of each object can be determined. Furthermore, rewriting prompts can be generated based on the purpose or location of each material in the target device. The rewriting prompts can be used to modify the description information of the object's contour shape, making the multimodal large model pay more attention to key information such as the object's material and location.

[0063] In one example, atomic numbers can be used to represent different colors, thus transforming the mapping between color and object material into a mapping between atomic numbers and object material. After determining the object color for multiple objects, the object color is matched with its atomic number, and then the object material is determined based on the atomic number and the mapping relationship.

[0064] In the process of matching object color with atomic number, the object color and the mapping relationship between multiple atomic numbers and multiple colors can be input into the multimodal large model, and the multimodal large model outputs the atomic number corresponding to the object color.

[0065] According to embodiments of this disclosure, by combining prior knowledge information representing the mapping relationship between color and material, rewritten prompts are generated for each object, and contour information is associated with material properties, making the text prompts more consistent with the semantic information of the pseudo-color image, thus providing more accurate guidance for the multimodal model.

[0066] According to embodiments of this disclosure, a text prompt for rewriting contour shape information is generated based on prior knowledge information and object color, including: determining the object material corresponding to the object color based on prior knowledge information; determining a target component that matches the object material and contour shape information; and generating a rewriting prompt for rewriting the contour shape information into the target component, wherein the target component is used to compose a target device.

[0067] For each object, the object material can be determined based on prior knowledge and the object's color.

[0068] For a target device, a component library can be set up in advance based on expert experience. For example, if it is determined through expert experience that a tubular metal object is usually the outer shell of the target device, then a key-value pair can be set in the component library as: metal, tubular - outer shell of the target device.

[0069] Based on the object's material and outline shape information, a matching process is performed in a pre-set component library to determine the target component that matches the object's material and outline information. Then, the outline shape information can be rewritten based on the target component to obtain rewriting prompts.

[0070] In one example, a yellow tubular object appears in a pseudo-color image. Based on prior knowledge, the object's material can be determined to be metal. Combining the object's material and outline shape information, a match is made in a component library, and the resulting target component indicates that the object is the outer shell of a target device. Therefore, the rewrite prompt can include rewriting the object's color and outline shape information to reflect the target device, i.e., "Rewrite the yellow metal tubular object as the outer shell of an explosive detonator, focusing on its material and location."

[0071] In another example, for a red blocky object in a pseudo-color image, the rewriting prompt could be "Rewrite the red blocky object as a high explosive core for explosives, focusing on matching the material properties corresponding to its pseudo-color".

[0072] In another example, for the blue area in the pseudo-color image, based on prior knowledge, it can be determined that the object is made of plastic. The rewritten prompt could be "Rewrite the light blue plastic block as ordinary non-explosive plastic, which does not require special attention".

[0073] According to embodiments of this disclosure, the material of an object is determined based on its color, and then the target component corresponding to the object is obtained by combining the outline shape information. This can transform low-level outline shape information into high-level semantics for the components constituting the target device, greatly improving the model's ability to understand and recognize component types.

[0074] According to embodiments of this disclosure, the method for detecting a target device further includes: acquiring multiple target components corresponding to multiple objects respectively; among the multiple target components, there are multiple predetermined components in a predetermined component combination, and the contour position information corresponding to each of the multiple predetermined components satisfies a predetermined positional relationship; generating a guiding prompt word corresponding to the predetermined component combination, wherein the text prompt word includes a guiding prompt word, and the predetermined component combination corresponds to the triggering method of the target device.

[0075] After obtaining the target components corresponding to multiple objects by matching them in the component library, the target components corresponding to each of the multiple objects can be checked based on a number of pre-defined component combinations to determine whether the multiple target components can form one or more pre-defined component combinations. The pre-defined component combinations include multiple components arranged according to a predetermined positional relationship, and different pre-defined component combinations represent target devices of different specifications.

[0076] If multiple predetermined components are identified within a predetermined component combination among multiple target components, the contour position information corresponding to each of the predetermined components is checked. If the multiple contour position information satisfies a predetermined positional relationship, a guiding prompt word corresponding to the predetermined component combination is generated. This guiding prompt word informs the multimodal large model to perform detection based on the features of the predetermined component combination. The predetermined positional relationship can be set so that all contour position information is adjacent.

[0077] In one example, if the target to be detected is determined through a pseudo-color image to include an explosive casing, a power supply device, and an electronic triggering device, and these components are interconnected, the guiding prompt can be set to "Detected detonator casing, battery, and electronic triggering device. The three are interconnected, matching the electric detonation method. Focus on confirming the integrity of the connection of each component."

[0078] In another example, if the target to be detected is determined to include an explosive casing and a physical triggering device through a pseudo-color image, and the above components are connected to each other, the guiding prompt can be set to "Detected detonator casing and fuse, which are adjacent to each other, matching the fire ignition explosion mode, focusing on confirming the connection stability between the fuse and the detonator".

[0079] Figure 3 A schematic diagram illustrating a predetermined combination of components determined according to an embodiment of the present disclosure is shown.

[0080] like Figure 3As shown, there are four objects in the target to be detected, including a yellow tubular object 301, a red block-shaped object 302, and two other interfering objects. By matching in the component library, it can be determined that the target component corresponding to the yellow tubular object 301 is the explosive casing, and the target component corresponding to the red block-shaped object 302 is the explosive contents, and the yellow tubular object 301 and the red block-shaped object 302 are connected to each other.

[0081] The target component and multiple predetermined components in the predetermined component combination are identical, and the contour position information corresponding to each of the target components satisfies the predetermined positional relationship. Therefore, it is determined that the yellow tubular object 301 and the red block-shaped object 302 constitute the predetermined component combination. In this case, the guiding prompt can be set to "Only detonator and explosive core are detected, no triggering component is found, the explosion mode cannot be determined at present, it is necessary to determine whether the core triggering component is missing, and a complete explosive cannot be formed."

[0082] According to embodiments of this disclosure, if the target component satisfies a predetermined component combination and the positional relationship satisfies a predetermined positional relationship, a corresponding guiding prompt word is generated to guide the multimodal large model to perform detection according to the triggering or assembly rules of the target device, thereby improving the detection rate and detection accuracy of the target device.

[0083] According to embodiments of this disclosure, the text prompts also include task prompts for guiding the multimodal large model to perform a classification task. The task prompts can guide the multimodal large model to perform the classification task according to the submitted text prompts and output a classification result indicating whether a target device exists.

[0084] The multimodal large model inputs text prompts and the contour information of multiple objects, and outputs detection results for the target device. This includes: when the text prompts include rewritten prompts, using the multimodal large model, guided by the rewritten prompts, rewriting the contour shape information of multiple objects into target components; and, guided by task prompts, generating detection results based on the target components, contour position information, and pseudo-color images of multiple objects; when the text prompts include both rewritten prompts and guiding prompts, using the multimodal large model, guided by the rewritten prompts, rewriting the contour shape information of multiple objects into target components; and, guided by task prompts, generating detection results matching the semantic rules represented by the guiding prompts based on the target components, contour position information, and pseudo-color images of multiple objects.

[0085] In cases where text prompts include rewritten prompts, a multimodal large model is used to rewrite the contour shape information of multiple objects into the target part indicated by the rewritten prompt, so as to convert low-level image features into high-level semantic features of the actual part.

[0086] Guided by task prompts, and based on multiple target components, contour position information, and pseudo-color images, it is possible to determine whether multiple target components constitute a target device and obtain detection results.

[0087] When the text prompt also includes a guiding prompt, based on multiple target components, contour position information, and a pseudo-color image, it can be determined whether the multiple target components form a predetermined component combination corresponding to the guiding prompt, thus obtaining a detection result. Since the predetermined component combination can constitute a target device, if the detection result indicates that multiple target components form a predetermined component combination, it can be determined that the multiple target components constitute the target device.

[0088] The detection results can include whether the target device is present in the pseudo-color image. If the target device is present, the detection results can also include the type and location of the target device. Since the detection results are obtained from a multimodal large model, when outputting the location of the target device, the multimodal large model can output a textual description of the location or a pseudo-color image, and mark the location of the target device with a bounding box on the pseudo-color image.

[0089] Figure 4 A schematic diagram illustrating the detection results output according to an embodiment of the present disclosure is shown.

[0090] like Figure 4 As shown, there are five objects in the target to be detected, including a yellow tubular object 301 and a brown linear object 401, as well as three other interfering objects. Using the above-described target device detection method, it can be determined that the yellow tubular object 301 and the brown linear object 401 conform to the composition and positional relationship of an explosive casing and a physical triggering device.

[0091] Therefore, the detection results include text content, such as: "Target device detected! Location marked by bounding box." The detection results also include... Figure 4 The image content shown, Figure 4 The bounding box 402 in the diagram marks the location of the target device to facilitate subsequent inspections by relevant personnel.

[0092] According to embodiments of this disclosure, by rewriting the hierarchical guidance of prompt words and task prompt words, the model first identifies components and then verifies combination rules, resulting in clear logic and more standardized detection results that better meet task requirements. Furthermore, when the text prompt words also include guiding prompt words, the semantic rules represented by the guiding prompt words can be used to specifically monitor predetermined component combinations corresponding to the guiding prompt words, thereby improving detection efficiency and targeting.

[0093] According to embodiments of this disclosure, the method for detecting a target device further includes: obtaining object image features of multiple objects from a pseudo-color image, and replacing the pseudo-color image with the multiple object image features to input a multimodal model, so that the multimodal model generates a detection result for the target device based on text prompts, object contour information of multiple objects, and image features.

[0094] For each object, a pseudo-color image can be input into an image encoder, and a feature extraction algorithm can be used to extract the object image features from the pseudo-color image. The object image features of each object, text prompts, and object contour information of each object are then input into a multimodal large model, so that the multimodal large model can generate detection results for the target device according to the above input information.

[0095] According to embodiments of this disclosure, introducing object image features and replacing the original image can reduce interference from redundant information unrelated to the object in the pseudo-color image, making the multimodal large model more focused on the features of the component itself, and further improving the accuracy of the detection results.

[0096] Typically, a target device consists of multiple components, each of which can be designed in different ways, resulting in multiple styles of components that make up the target device. Fine-tuning of the multimodal large model allows it to learn the different styles of multiple components, thereby improving its generalization ability.

[0097] Specifically, the multimodal large model is fine-tuned in the following way: multiple fine-tuning datasets are acquired, which include sample pseudo-color images, image labels, outline information of each labeled object in the sample pseudo-color images, and sample text prompts; wherein, the sample text prompts are generated based on the outline information of multiple labeled objects; multiple first sample pseudo-color images in the multiple fine-tuning datasets include multiple styles of the same component used to compose the target device; the multimodal large model to be fine-tuned is fine-tuned using the multiple fine-tuning datasets, and the multimodal large model is obtained under predetermined conditions.

[0098] The sample pseudo-color images in the fine-tuning dataset can be obtained in the same way as the pseudo-color images, which will not be repeated here.

[0099] Image labels in the fine-tuning dataset indicate whether the sample pseudo-color image includes the target device. If the target device is included, the image label also includes the target device's location information within the sample pseudo-color image. Image labels can be obtained manually or determined based on the detection results of validated historical data output by the multimodal large model.

[0100] In fine-tuning the dataset, the outline information of each labeled object in the sample pseudo-color image can be obtained through manual annotation or determined based on the recognition results of the verified pseudo-color image.

[0101] By using multiple fine-tuning datasets, the multimodal large model to be fine-tuned is fine-tuned. Under predetermined conditions, it can be determined that the multimodal large model to be fine-tuned has learned multiple styles of the parts corresponding to the fine-tuning datasets. Based on the model parameters of the current multimodal large model to be fine-tuned, the multimodal large model can be determined.

[0102] According to embodiments of this disclosure, a model is fine-tuned using a dataset containing multiple styles of the same component to enhance the generalization ability of the multimodal large model for components of different shapes, enabling the multimodal large model to adapt to diverse real-world detection scenarios.

[0103] According to embodiments of this disclosure, the predetermined condition includes the convergence of the loss function, which includes: a loss term for measuring the classification difference between the sample detection results output by the multimodal large model and the image labels; and a loss term for measuring the difference between the first relative position between multiple labeled objects determined by the multimodal large model from the sample pseudo-color image and the second relative position between the labeled contour position information in the contour information of the multiple labeled objects.

[0104] The outline information of the labeled object can include the outline shape information and the outline position information. The outline shape information is similar to the outline shape information mentioned above, and the outline position information is similar to the outline position information mentioned above.

[0105] A multimodal large model can be used to process the sample pseudo-color images, the outline information of each labeled object, and the sample text prompts included in the fine-tuning dataset to obtain sample detection results indicating whether the multiple labeled objects in the sample pseudo-color images can form a target device. Based on the sample detection results and image labels, a loss term representing classification differences can be determined, such as the common KL loss or other classification losses.

[0106] Besides the multimodal large model's ability to detect target devices, its ability to determine the positional relationships between labeled objects also affects the accuracy of the detection results. Therefore, the multimodal large model can be used to determine the first relative position between multiple labeled objects in a sample pseudo-color image. The second relative position is then determined based on the positional information of the labeled contours among the multiple labeled contours. For example, the Euclidean distance and / or Mahalanobis distance between the multiple labeled contour positions can be used as the second relative position. After identifying the positional information of multiple labeled objects from the sample pseudo-color image using the multimodal large model, the Euclidean distance and / or Mahalanobis distance between these multiple positional information are used as the first relative position. Then, based on the difference between the first and second relative positions, a loss term representing the positional difference is determined.

[0107] A loss function is constructed based on the loss terms representing classification differences and the loss terms representing position differences. The multimodal large model is determined when the loss function converges. This allows the fine-tuned multimodal large model to meet the requirements for both the detection capability of the target device and the judgment capability of the positional relationship, thereby improving the performance of the multimodal large model.

[0108] According to embodiments of this disclosure, introducing classification loss and relative position loss during fine-tuning enables the multimodal large model to simultaneously learn component categories and spatial positional relationships, significantly improving the accuracy of target device combination judgment.

[0109] Figure 5 A block diagram of a detection apparatus for a target device according to an embodiment of the present disclosure is shown schematically.

[0110] like Figure 5 As shown, the detection device 500 for the target device includes an identification module 510, a generation module 520, and a detection module 530.

[0111] The recognition module 510 is used to recognize the object outline information of each of multiple objects from a pseudo-color image, wherein the pseudo-color image includes multiple colors, the colors are used to characterize the material of the objects, and at least one of the multiple objects is a component used to compose the target device.

[0112] The generation module 520 is used to generate text prompts for identifying the target device based on the object outline information of multiple objects;

[0113] The detection module 530 is used to input text prompts, object outline information of multiple objects, and pseudo-color images into a multimodal large model, and output detection results for the target device. The detection results are used to characterize whether multiple objects can form the target device.

[0114] According to embodiments of this disclosure, the generation module 520 includes an information acquisition submodule, a color determination submodule, and a prompt word generation submodule. The object contour information includes contour shape information and contour position information.

[0115] The information acquisition submodule is used to acquire prior knowledge information for pseudo-color images, where prior knowledge information is used to indicate the mapping relationship between object material and color.

[0116] The color determination submodule is used to determine the object color from the pseudo-color image that matches the contour position information for each object.

[0117] The prompt word generation submodule is used to generate rewrite prompt words for rewriting outline shape information based on prior knowledge information and object color.

[0118] According to embodiments of this disclosure, the prompt word generation submodule includes a material determination unit and a prompt word generation unit.

[0119] The material determination unit is used to determine the object material corresponding to the object's color based on prior knowledge information.

[0120] The prompt word generation unit is used to determine the target component that matches the object's material and outline shape information, and to generate a rewriting prompt word for rewriting the outline shape information into the target component, wherein the target component is used to compose the target device.

[0121] According to embodiments of this disclosure, the target device detection device 500 further includes a component acquisition module and a prompt word generation module.

[0122] The component acquisition module is used to acquire multiple target components that correspond to multiple objects.

[0123] The prompt word generation module is used to generate a guiding prompt word corresponding to the predetermined component combination when there are multiple predetermined components in a predetermined component combination among multiple target components, and the contour position information of each of the multiple predetermined components satisfies a predetermined positional relationship. The text prompt word includes a guiding prompt word, and the predetermined component combination corresponds to the triggering method of the target device.

[0124] According to embodiments of this disclosure, the detection module 530 includes a first detection submodule and a second detection submodule.

[0125] The first detection submodule is used to rewrite the contour shape information of multiple objects into target parts under the guidance of the rewriting prompts, when the text prompts include rewriting prompts; and to generate detection results based on the target parts, contour position information and pseudo-color images of multiple objects under the guidance of the task prompts.

[0126] The second detection submodule is used to rewrite the contour shape information of multiple objects into target parts under the guidance of the rewriting prompts, when the text prompts include rewriting prompts and guiding prompts. Under the guidance of the task prompts, it generates detection results that match the semantic rules represented by the guiding prompts based on the target parts, contour position information and pseudo-color images of multiple objects.

[0127] According to embodiments of this disclosure, the detection device 500 for the target device further includes a feature input module.

[0128] The feature input module is used to obtain the object image features of multiple objects from the pseudo-color image, and use the multiple object image features to replace the pseudo-color image input to the multimodal model, so that the multimodal model can generate detection results for the target device based on the text prompt words, the object contour information of multiple objects and the image features.

[0129] According to embodiments of this disclosure, the target device detection device 500 further includes a dataset acquisition module and a model fine-tuning module.

[0130] The dataset acquisition module is used to acquire multiple fine-tuning datasets. The fine-tuning datasets include sample pseudo-color images, image labels, outline information of each labeled object in the sample pseudo-color images, and sample text prompts. The sample text prompts are generated based on the outline information of multiple labeled objects. The multiple first sample pseudo-color images in the multiple fine-tuning datasets include multiple styles for the same component that makes up the target device.

[0131] The model fine-tuning module is used to fine-tune the large multimodal model to be fine-tuned using multiple fine-tuning datasets, and obtain the large multimodal model under predetermined conditions.

[0132] Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure, or at least part of the functions of any one or more of them, can be implemented in one module. Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be implemented by dividing them into multiple modules. Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be at least partially implemented as hardware circuitry, such as a Field-Programmable Gate Array (FPGA), a Programmable Logic Array (PLA), a System-on-Chip, a System-on-a-Substrate, a System-on-Package, an Application-Specific Integrated Circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be at least partially implemented as computer program modules, which, when run, can perform corresponding functions.

[0133] For example, any plurality of the identification module 510, generation module 520, and detection module 530 may be combined into one module / unit / subunit, or any one of these modules / units / subunits may be split into multiple modules / units / subunits. Alternatively, at least part of the functionality of one or more of these modules / units / subunits may be combined with at least part of the functionality of other modules / units / subunits and implemented in one module / unit / subunit. According to embodiments of this disclosure, at least one of the identification module 510, generation module 520, and detection module 530 may be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging the circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the identification module 510, generation module 520, and detection module 530 may be implemented at least partially as a computer program module, which can perform corresponding functions when the computer program module is run.

[0134] It should be noted that the apparatus portion in the embodiments of this disclosure corresponds to the method portion in the embodiments of this disclosure. The description of the apparatus portion is specifically referred to in the method portion, and will not be repeated here.

[0135] Figure 6 A block diagram of an electronic device suitable for implementing a detection method for a target device according to an embodiment of the present disclosure is shown schematically.

[0136] Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0137] like Figure 6 As shown, an electronic device 600 according to an embodiment of this disclosure includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage portion 608 into a random access memory (RAM) 603. The processor 601 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include onboard memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this disclosure.

[0138] RAM 603 stores various programs and data required for the operation of electronic device 600. Processor 601, ROM 602, and RAM 603 are interconnected via bus 604. Processor 601 performs various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 602 and / or RAM 603. It should be noted that the programs may also be stored in one or more memories other than ROM 602 and RAM 603. Processor 601 may also perform various operations of the method flow according to embodiments of the present disclosure by executing programs stored in said one or more memories.

[0139] According to embodiments of this disclosure, the electronic device 600 may further include an input / output (I / O) interface 605, which is also connected to a bus 604. The electronic device 600 may also include one or more of the following components connected to the input / output (I / O) interface 605: an input section 606 including a keyboard, mouse, etc.; an output section 607 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output (I / O) interface 605 as needed. A removable medium 611, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 610 as needed so that computer programs read from it can be installed into the storage section 608 as needed.

[0140] According to embodiments of this disclosure, the method flow according to embodiments of this disclosure can be implemented as a computer software program. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via communication section 609, and / or installed from removable medium 611. When the computer program is executed by processor 601, it performs the functions defined in the system of embodiments of this disclosure. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0141] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.

[0142] According to embodiments of this disclosure, the computer-readable storage medium can be a non-volatile computer-readable storage medium. Examples include, but are not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0143] For example, according to embodiments of this disclosure, a computer-readable storage medium may include the ROM 602 and / or RAM 603 described above and / or one or more memories other than ROM 602 and RAM 603.

[0144] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods provided in the embodiments of this disclosure. When the computer program product is run on an electronic device, the program code is used to enable the electronic device to implement the methods provided in the embodiments of this disclosure.

[0145] When the computer program is executed by the processor 601, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0146] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and downloaded and installed via the communication section 609, and / or installed from the removable medium 611. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0147] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can execute entirely on a user's computing device, partially on a user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0148] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions. Those skilled in the art will understand that the features described in the various embodiments can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this disclosure. In particular, the features described in the various embodiments of this disclosure may be combined and / or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.

[0149] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.

Claims

1. A method for detecting a target device, comprising: Identify the object outline information of multiple objects from a pseudo-color image, wherein the pseudo-color image includes multiple colors, the colors are used to characterize the object material, and at least one of the multiple objects is a component for assembling a target device; Generate text prompts for identifying the target device based on the object contour information of each of the multiple objects; and The text prompt, the object contour information of each of the multiple objects, and the pseudo-color image are input into a multimodal large model, and the detection result for the target device is output. The detection result is used to characterize whether the multiple objects can form the target device.

2. The method according to claim 1, wherein, The text prompts include rewritten prompts for each of the objects, and the object outline information includes outline shape information and outline position information; The step of generating text prompts for identifying the target device based on the object contour information of each of the multiple objects includes: Obtain prior knowledge information for the pseudo-color image, wherein the prior knowledge information is used to indicate the mapping relationship between object material and color; For each of the aforementioned objects, determine the object color from the pseudo-color image that matches the contour position information of each of the aforementioned objects; and Based on the prior knowledge information and the object color, a rewriting prompt word is generated for rewriting the outline shape information.

3. The method according to claim 2, wherein, The step of generating the rewriting prompt word for rewriting the outline shape information based on the prior knowledge information and the object color includes: Based on the prior knowledge information, determine the object material corresponding to the object color; A target component matching the object material and the outline shape information is identified, and a rewriting prompt is generated to rewrite the outline shape information into the target component, wherein the target component is used to compose the target device.

4. The method according to claim 3, wherein, The method further includes: Obtain the multiple target components that correspond to the multiple objects respectively; Among the multiple target components, there are multiple predetermined components in a predetermined component combination, and the contour position information corresponding to each of the multiple predetermined components satisfies a predetermined positional relationship. A guiding prompt word corresponding to the predetermined component combination is generated, wherein the text prompt word includes the guiding prompt word, and the predetermined component combination corresponds to the triggering method of the target device.

5. The method according to claim 3 or 4, wherein, The text prompts also include task prompts used to guide the multimodal large model in performing classification tasks; The text prompts and the object contour information of each of the multiple objects are input into a multimodal large model, and the detection results for the target device are output, including: When the text prompt includes the rewriting prompt, the multimodal large model is used to rewrite the contour shape information of each of the multiple objects into the target component under the guidance of the rewriting prompt; and under the guidance of the task prompt, the detection result is generated based on the target component of each of the multiple objects, the contour position information, and the pseudo-color image. When the text prompt includes the rewriting prompt and the guiding prompt, the multimodal large model is used to rewrite the contour shape information of each of the multiple objects into the target component under the guidance of the rewriting prompt; and under the guidance of the task prompt, the detection result matching the semantic rules represented by the guiding prompt is generated based on the target component of each of the multiple objects, the contour position information, and the pseudo-color image.

6. The method according to claim 1, wherein, The method further includes: The object image features of each of the multiple objects are obtained from the pseudo-color image, and the pseudo-color image is replaced with the multiple object image features and input into the multimodal model, so that the multimodal model generates the detection result for the target device based on the text prompt words, the object contour information of each of the multiple objects, and the image features.

7. The method according to any one of claims 1 to 4, wherein, The multimodal large model is fine-tuned using the following method: Multiple fine-tuning datasets are acquired, each fine-tuning dataset including sample pseudo-color images, image labels, outline information of each labeled object in the sample pseudo-color images, and sample text prompts; wherein, the sample text prompts are generated based on the outline information of the multiple labeled objects; the multiple first sample pseudo-color images in the multiple fine-tuning datasets include multiple styles for composing the same component of the target device; The multimodal large model to be fine-tuned is obtained by using multiple fine-tuning datasets to fine-tune the model under predetermined conditions.

8. The method according to claim 7, wherein, The predetermined conditions include the convergence of the loss function, which includes: A loss term used to measure the classification difference between the sample detection results output by the multimodal large model and the image labels; A loss term used to measure the difference between the first relative position between the multiple labeled objects determined by the multimodal large model from the sample pseudo-color image and the second relative position between the labeled contour position information in the contour information of the multiple labeled objects.

9. A detection device for a target device, characterized in that, The device includes: The recognition module is used to identify the object outline information of multiple objects from a pseudo-color image, wherein the pseudo-color image includes multiple colors, the colors are used to characterize the material of the objects, and at least one of the multiple objects is a component for assembling a target device. The generation module is used to generate text prompts for identifying the target device based on the object contour information of each of the multiple objects; The detection module is used to input the text prompt, the object contour information of each of the multiple objects, and the pseudo-color image into a multimodal large model, and output the detection result for the target device, wherein the detection result is used to characterize whether the multiple objects can form the target device.

10. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 8.

11. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 8.

12. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 8.