An engineering drawing information desensitization method based on a multi-modal large model
By using a multimodal large model to perform semantic screening on engineering drawings, sensitive information in images and text is identified and processed. This solves the problems of poor positioning robustness and lack of semantic understanding in existing technologies, and achieves accurate positioning and removal of sensitive information, while preserving the professional usability of the drawings.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- WUXI XUELANG DIGITAL TECH CO LTD
- Filing Date
- 2026-04-27
- Publication Date
- 2026-07-24
Smart Images

Figure CN122113061B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of drawing processing technology, and more specifically, to a method for desensitizing engineering drawing information based on a multimodal large model. Background Technology
[0002] In the engineering design industry, drawings, as core deliverables, typically contain a large amount of confidential information, such as project contract numbers, client company names, handwritten signatures of designers / reviewers, company logos, internal approval stamps, and watermarks. When such drawings are used for artificial intelligence model training, third-party collaboration, or public demonstrations, there is a serious risk of data leakage if sensitive information is not effectively anonymized.
[0003] Existing technologies typically rely on regular expressions or optical character recognition (OCR) results to perform string matching within a pre-defined keyword library (e.g., "Party A:", "Contract No.:") and then perform rectangular cropping on the matching area to achieve desensitization of engineering drawings.
[0004] However, the above technologies have problems such as poor positioning robustness, lack of semantic understanding and weak positioning ability in practical engineering applications, which reduce the professional usability of the anonymized drawings. Summary of the Invention
[0005] The purpose of this application is to provide a method for desensitizing engineering drawing information based on a multimodal large model, in order to address the shortcomings of the prior art mentioned above, and to solve the problems of poor positioning robustness, lack of semantic understanding and weak positioning ability in the prior art.
[0006] To achieve the above objectives, the technical solutions adopted in the embodiments of this application are as follows: In a first aspect, one embodiment of this application provides a method for desensitizing engineering drawing information based on a multimodal large model, the method comprising: Obtain the engineering drawings to be desensitized and the desensitization method, wherein the desensitization method includes at least one of the following: pixel overlay, watermark desensitization, and Gaussian blurring. The engineering drawings to be desensitized and the preset semantic screening prompts are input into a pre-trained multimodal large model. The multimodal large model performs semantic screening to obtain multiple candidate image sensitive information and multiple candidate text sensitive information corresponding to the engineering drawings to be desensitized. The candidate image sensitive information includes: image keywords and image bounding boxes. The candidate text sensitive information includes: text keywords, text values and text bounding boxes. Based on the multiple candidate image sensitive information and the multiple candidate text sensitive information, at least one target sensitive information is determined; Based on the desensitization method and the at least one target sensitive information, the engineering drawing to be desensitized is processed to obtain the desensitized engineering drawing.
[0007] Secondly, another embodiment of this application provides an engineering drawing information desensitization device based on a multimodal large model, the device comprising: The acquisition module is used to acquire the engineering drawings to be desensitized and the desensitization method, wherein the desensitization method includes at least one of the following: pixel overlay, watermark desensitization, and Gaussian blur coding; The semantic screening module is used to input the engineering drawings to be desensitized and the preset semantic screening prompts into a pre-trained multimodal large model. The multimodal large model performs semantic screening to obtain multiple candidate image sensitive information and multiple candidate text sensitive information corresponding to the engineering drawings to be desensitized. The candidate image sensitive information includes: image keywords and image bounding boxes. The candidate text sensitive information includes: text keywords, text values and text bounding boxes. The determining module is used to determine at least one target sensitive information based on the multiple candidate image sensitive information and the multiple candidate text sensitive information; The desensitization processing module is used to desensitize the engineering drawing to be desensitized according to the desensitization method and the at least one target sensitive information, so as to obtain the desensitized engineering drawing.
[0008] Thirdly, another embodiment of this application provides an electronic device, including: a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of any of the methods described in the first aspect above.
[0009] Fourthly, another embodiment of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of any of the methods described in the first aspect above.
[0010] The beneficial effects of this application are as follows: by acquiring the engineering drawings to be desensitized and the desensitization method, and performing semantic screening using a multimodal large model, multiple candidate image sensitive information and multiple candidate text sensitive information are obtained. It is possible to simultaneously understand the visual elements in the images and the contextual semantics in the text, and determine at least one target sensitive information based on the multiple candidate image sensitive information and multiple candidate text sensitive information. This enables efficient determination of sensitive information, and thus, based on the desensitization method and at least one target sensitive information, the engineering drawings to be desensitized are processed to obtain desensitized engineering drawings. It is possible to accurately locate and remove text-type and image-type sensitive information in the engineering drawings, while preserving the original usable information of the drawings to the maximum extent. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 A flowchart illustrating a method for desensitizing engineering drawing information based on a multimodal large model, as provided in an embodiment of this application; Figure 2 This is a schematic diagram of a process for determining at least one target sensitive information in the engineering drawing information desensitization method based on a multimodal large model provided in the embodiments of this application; Figure 3 This is a flowchart illustrating the process of determining at least one target image sensitive information from multiple candidate image sensitive information in the engineering drawing information desensitization method based on a multimodal large model provided in this application embodiment; Figure 4 This is a flowchart illustrating the process of determining at least one target text sensitive information from multiple candidate text sensitive information in the engineering drawing information desensitization method based on a multimodal large model provided in this application embodiment; Figure 5 This is a schematic diagram of the process of obtaining desensitized engineering drawings in the engineering drawing information desensitization method based on a multimodal large model provided in the embodiments of this application; Figure 6 This is a flowchart illustrating the process of determining the foreground outline and background region corresponding to the target bounding box in the engineering drawing information desensitization method based on a multimodal large model provided in this application embodiment. Figure 7 This is another flowchart illustrating the process of obtaining desensitized engineering drawings in the engineering drawing information desensitization method based on a multimodal large model provided in the embodiments of this application. Figure 8 A schematic diagram of an engineering drawing information desensitization device based on a multimodal large model provided in this application embodiment; Figure 9 This is a schematic diagram of the electronic device structure provided in an embodiment of this application. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.
[0014] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0015] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.
[0016] Existing technologies typically rely on regular expressions or optical character recognition (OCR) results to perform string matching within a pre-defined keyword library (e.g., "Party A:", "Contract No.:") and then perform rectangular cropping on the matching area to achieve desensitization of engineering drawings.
[0017] However, the above technologies have problems such as poor positioning robustness, lack of semantic understanding and weak positioning ability in practical engineering applications, which reduce the professional usability of the anonymized drawings.
[0018] Based on the aforementioned problems, this application proposes a method for desensitizing engineering drawing information based on a multimodal large model. The method involves acquiring the engineering drawing to be desensitized and the desensitization method, performing initial semantic screening using a multimodal large model to obtain multiple candidate image sensitive information and multiple candidate text sensitive information, and identifying at least one target sensitive information. Then, based on the desensitization method and at least one target sensitive information, the engineering drawing to be desensitized is processed to obtain the desensitized engineering drawing. This method can accurately locate and remove text and image sensitive information in the engineering drawing, while preserving the original usable information of the drawing to the maximum extent.
[0019] The following describes in detail the method for desensitizing engineering drawing information based on a multimodal large model provided in this application, with reference to several embodiments.
[0020] Figure 1 This is a flowchart illustrating a method for desensitizing engineering drawing information based on a multimodal large model, as provided in an embodiment of this application. Figure 1 As shown, the subject executing this method can be any electronic device with processing capabilities, and the method includes: S101. Obtain the drawings of the project to be desensitized and the desensitization method.
[0021] Optionally, obtain the engineering drawings to be de-identified and the de-identification method.
[0022] The engineering drawings to be anonymized must contain at least one piece of sensitive information. Specifically, the sensitive information can be visual elements such as logos, signatures, and seals, or semantic elements such as text. The anonymization methods include at least one of the following: pixel overlay, watermark anonymization, and Gaussian blur masking.
[0023] In one example, the desensitization method can be determined in response to the user's selection of a desensitization method.
[0024] In another example, the user-inputted desensitization strength information can be obtained, and the desensitization method can be determined based on this information. For instance, based on the user-inputted desensitization strength information, the desensitization method can be determined as follows: watermark desensitization for the logo element and Gaussian blur masking for the text element.
[0025] S102. Input the engineering drawings to be desensitized and the preset semantic screening prompts into the pre-trained multimodal large model. The multimodal large model performs semantic screening to obtain multiple candidate image sensitive information and multiple candidate text sensitive information corresponding to the engineering drawings to be desensitized.
[0026] Optionally, the engineering drawings to be desensitized and the preset semantic screening prompts are input into a pre-trained multimodal large model. The multimodal large model performs semantic screening on the engineering drawings to be desensitized based on the semantic screening prompts, and infers multiple candidate image sensitive information and multiple candidate text sensitive information corresponding to the engineering drawings to be desensitized.
[0027] Among them, the multimodal large model can be the Qwen3-VL visual language model.
[0028] Among them, the sensitive information of the candidate image and the sensitive information of the candidate text are both general sensitive information.
[0029] The sensitive information of the candidate images includes: image keyword key and image bounding box (bbox); the sensitive information of the candidate texts includes: text keyword key, text value, and text bounding box (bbox).
[0030] Specifically, image keywords refer to the summary and classification of sensitive information in the selected image, while text keywords refer to the summary and classification of sensitive information in the selected text.
[0031] In one example, the sensitive information for the candidate image could be: {"key": "Company Logo", "bbox":[x1, y1, x2, y2]}. The sensitive information for the candidate text could be: {"key": "Auditor", "value": "Zhang San", "bbox": [x3, y3, x4, y4]}.
[0032] In another example, the candidate image sensitivity information may also include a value field, which can be a fixed value "image_region".
[0033] S103. Based on multiple candidate image sensitive information and multiple candidate text sensitive information, determine at least one target sensitive information.
[0034] Optionally, after obtaining multiple candidate image sensitive information and multiple candidate text sensitive information, the candidate image sensitive information and the candidate text sensitive information can be judged to determine at least one target sensitive information.
[0035] That is, after obtaining the candidate image sensitive information and candidate text sensitive information that represent general sensitive information, the candidate image sensitive information and candidate text sensitive information can be filtered and matched, which can more accurately match the target sensitive information that needs to be desensitized and retain more original usable information in the engineering drawings to be desensitized.
[0036] Among them, target sensitive information includes: target image sensitive information and / or target text sensitive information.
[0037] In one example, the sensitive information of each candidate image and each candidate text can be verified using a pre-constructed sensitive information database. If the verification passes, the corresponding candidate image and candidate text sensitive information are used as the target sensitive information.
[0038] In another example, based on a pre-defined knowledge graph of engineering drawing structures, the standard spatial location regions corresponding to each candidate text sensitive information can be obtained. Then, through a multimodal large model, the spatial overlap between the actual bounding box of each candidate text sensitive information and the standard spatial location region, as well as the semantic matching degree between the text value of the candidate text sensitive information and the semantic tag associated with the standard spatial location region, can be calculated. Based on the weighted fusion result of spatial overlap and semantic matching degree, a semantic consistency score is generated. Candidate text sensitive information with a semantic consistency score higher than a pre-defined semantic consistency threshold is taken as target text sensitive information.
[0039] In another example, adversarial perturbation samples are generated for the regions corresponding to the sensitive information of each candidate image. The adversarial perturbation samples and semantic initial screening prompts are then input into the multimodal large model to obtain secondary recognition results. If the image keywords in the secondary recognition results are consistent with the original recognition results, and the intersection-union ratio (IoU) of the image bounding boxes is greater than the preset IoU threshold, then the sensitive information of the candidate image is taken as the sensitive information of the target image.
[0040] S104. Based on the desensitization method and at least one target sensitive information, perform desensitization processing on the engineering drawings to be desensitized to obtain the desensitized engineering drawings.
[0041] Optionally, after obtaining the sensitive information of each target, the sensitive information of each target can be desensitized according to the desensitization method to obtain the desensitized engineering drawings.
[0042] For example, it is possible to iterate through each target sensitive information, and for the current target sensitive information that has been iterated through, the current target sensitive information is found from the engineering drawing to be desensitized according to the current target sensitive information and the desensitization method corresponding to the current target sensitive information, and desensitization processing is performed according to the desensitization method corresponding to the current target sensitive information. After all target sensitive information has been processed, the desensitized engineering drawing is obtained.
[0043] In one example, if the desensitization method is Gaussian blurring, the target bounding box coordinates are obtained based on the determined target sensitive information. Then, based on the obtained bounding box coordinates, the corresponding local area is cropped from the engineering drawing to be desensitized. The cropped local area is then Gaussian blurred, and the blurred local area is pasted back to the corresponding position on the engineering drawing to be desensitized according to the original bounding box coordinates, thus achieving the desensitization process.
[0044] In this embodiment, by acquiring the engineering drawings to be desensitized and the desensitization method, and performing semantic screening using a multimodal large model, multiple candidate image sensitive information and multiple candidate text sensitive information are obtained. This allows for the simultaneous understanding of visual elements in the images and contextual semantics in the text. Based on the multiple candidate image sensitive information and multiple candidate text sensitive information, at least one target sensitive information is determined, enabling efficient determination of sensitive information. Thus, based on the desensitization method and at least one target sensitive information, the engineering drawings to be desensitized are processed to obtain desensitized engineering drawings. This allows for precise location and removal of text-based and image-based sensitive information in the engineering drawings, while preserving the original usable information of the drawings to the maximum extent.
[0045] In one possible implementation, Figure 2 This is a flowchart illustrating the process of determining at least one target sensitive information in the engineering drawing information desensitization method based on a multimodal large model provided in this application embodiment, with reference to... Figure 2 As shown, in step S103 above, at least one target sensitive information is determined based on multiple candidate image sensitive information and multiple candidate text sensitive information, including: S201. Based on a preset sensitive information database, determine at least one target image sensitive information from multiple candidate image sensitive information, and / or, based on a preset sensitive information database, determine at least one target text sensitive information from multiple candidate text sensitive information.
[0046] Optionally, based on a preset sensitive information database, multiple candidate image sensitive information and multiple candidate text sensitive information can be retrieved and recalled to determine at least one target image sensitive information from the multiple candidate image sensitive information, and / or to determine at least one target text sensitive information from the multiple candidate text sensitive information.
[0047] Specifically, the sensitive information database can be a multimodal sensitive information vector database. The sensitive information database can be dynamically constructed.
[0048] For example, the specific construction process of the sensitive information database includes: text embedding, image embedding, and vector storage.
[0049] Specifically, the text embedding includes: using the Qwen3-Embedding model to encode predefined sensitive keywords (such as “client”, “auditor”, “confidentiality level”, etc.) and their semantic variants into 768-dimensional vectors.
[0050] Specifically, image embedding includes: constructing a self-supervised image embedding model Image-Embedding based on DINOv3 (Vision Transformer architecture), and encoding images using the self-supervised image embedding model Image-Embedding.
[0051] Among them, the self-supervised image embedding model Image-Embedding, based on DINOv3, removes the original classification head and replaces it with a learnable fully connected layer, outputting a 768-dimensional embedding vector.
[0052] Meanwhile, when training the self-supervised image embedding model, Image-Embedding, we first used images including a large number of engineering drawing logos, signatures, and seals as samples. We then paired these images to form data pairs, and manually labeled these data pairs with similarity scores ranging from 0 to 1. Finally, we employed a dual-tower structure (parameter sharing), cosine similarity calculation connecting the two towers, and the MSE loss function for end-to-end training, enabling the model to possess semantic retrieval capabilities based on image search.
[0053] Specifically, the dual-tower structure refers to inputting sample pairs into two identical models, with parameters shared between the two models. The model outputs (two vectors) are then compared using cosine similarity to obtain a similarity score. The loss between the model outputs and the labels is then calculated using the MSE loss function. Finally, backpropagation is performed using this loss to update the model parameters. This allows the model's output vectors to reasonably express their semantic features, adapt to the cosine similarity calculation method, and yield more accurate similarity results.
[0054] For example, cosine similarity The calculation can be performed using the following formula:
[0055] in, The input consists of two vectors.
[0056] For example, the mean squared error loss of MSE You can refer to the following formula:
[0057] in, This is the target value.
[0058] Specifically, vector storage includes: building an efficient vector index library using Facebook AI SimilaritySearch (FAISS) to store text and image embedding vectors separately, supporting fast nearest neighbor search in high-dimensional space.
[0059] For example, the vectors embedded with text / images, along with the corresponding original text and image indexes, are stored in the FAISS vector database. When a target vector enters the search, the target vector is compared with the vectors in the database using cosine similarity calculation. Vectors with similarity values above a threshold are obtained, and the corresponding original data is retrieved through the index.
[0060] S202. Treat the sensitive information of each target image and the sensitive information of each target text as a separate target sensitive information.
[0061] Optionally, the sensitive information of each target image and the sensitive information of each target text can be treated as a single target sensitive information.
[0062] By using a pre-defined sensitive information database, multiple candidate image and text sensitive information are retrieved and recalled to identify the target image and text sensitive information. Each target image and text sensitive information is then treated as a separate target sensitive information, enabling joint semantic and visual discrimination and precise location of sensitive information, thus ensuring the usability of the anonymized drawings. Furthermore, in practical applications, when users need to change sensitive information, they only need to update the sensitive information database to confirm the target sensitive information, improving operational simplicity and the reusability of the sensitive information database.
[0063] In one possible implementation, Figure 3 This is a flowchart illustrating the process of determining at least one target image sensitive information from multiple candidate image sensitive information in the engineering drawing information desensitization method based on a multimodal large model provided in this application embodiment. (Refer to...) Figure 3 As shown, in step S201 above, determining at least one target image sensitive information from multiple candidate image sensitive information based on a preset sensitive information database includes: S301. According to the preset text encoding module, perform text encoding processing on the image keywords of the sensitive information of the selected image to obtain the image keyword encoding vector.
[0064] Optionally, a preset text encoding module is used to perform text encoding processing on the image keywords of the sensitive information of the selected image to obtain the image keyword encoding vector. The text encoding module can be a Qwen3-Embedding encoding module.
[0065] By performing text encoding on the image keywords of the sensitive information in the selected image, an image keyword encoding vector is obtained. This vector accurately expresses the true semantic category of the sensitive information, allowing users to customize desensitized content based on the relevant category.
[0066] S302. Based on the image bounding box of the sensitive information of the candidate image, cut out the region image of the sensitive information of the candidate image from the engineering drawing to be desensitized.
[0067] Optionally, the sensitive information of the selected image can be cropped from the engineering drawing to be desensitized according to the image bounding box of the sensitive information of the selected image to obtain the region image of the sensitive information of the selected image.
[0068] S303. According to the preset image encoding module, perform image encoding processing on the region image of the sensitive information of the selected image to obtain the image encoding vector.
[0069] Optionally, an image encoding process is performed on the region of the selected image containing sensitive information using a preset image encoding module to obtain an image encoding vector. The image encoding module is an Image-Embedding model.
[0070] S304. Based on the image keyword encoding vector and / or image encoding vector, perform a search from the sensitive information database to obtain the search results.
[0071] Optionally, after obtaining the image keyword encoding vector and the image encoding vector, a vector search can be performed from the sensitive information database based on the image keyword encoding vector and / or the image encoding vector to obtain the search results.
[0072] In one example, a vector cosine similarity search is performed from a sensitive information database based on the image keyword encoding vector to obtain search results. These results include multiple candidate vectors similar to the image keyword encoding vector and their corresponding similarity scores.
[0073] In another example, a vector cosine similarity search is performed from a sensitive information database based on the image encoding vector to obtain the search results. These results include multiple candidate vectors similar to the image encoding vector and their corresponding similarity scores.
[0074] In another example, a vector cosine similarity search is performed from a sensitive information database based on the image keyword encoding vector and the image encoding vector to obtain search results. These results include multiple candidate vectors similar to the image keyword encoding vector and the image encoding vector, along with their corresponding similarity scores.
[0075] S305. Based on the search results, determine whether to use the sensitive information of the candidate image as the sensitive information of the target image. If so, use the sensitive information of the candidate image as the sensitive information of the target image.
[0076] In one example, the search results are judged sequentially according to a preset similarity threshold. If the similarity result is greater than the preset similarity threshold, the corresponding candidate vector is regarded as sensitive information of the target image.
[0077] In another example, the maximum similarity result in the search results is determined. If the maximum similarity result is greater than a preset similarity threshold, the corresponding candidate vector is used as sensitive information of the target image.
[0078] Image keyword encoding vectors and image encoding vectors are obtained through encoding. Based on the image keyword encoding vectors and / or image encoding vectors, a search is performed in the sensitive information database to obtain search results. Based on the search results, it is determined whether to use the candidate image sensitive information as a target image sensitive information. If so, the candidate image sensitive information is used as a target image sensitive information. This enables multimodal joint judgment, reduces the false judgment rate of target image sensitive information, and improves search recall efficiency.
[0079] In one possible implementation, Figure 4 This application provides a flowchart illustrating the process of determining at least one target text sensitive information from multiple candidate text sensitive information in a method for desensitizing engineering drawing information based on a multimodal large model. (Refer to...) Figure 4 As shown, in step S201 above, at least one target text sensitive information is determined from multiple candidate text sensitive information based on a preset sensitive information database, including: S401. According to the preset text encoding module, the text keywords of the selected sensitive text information are processed by text encoding to obtain the text keyword encoding vector.
[0080] Optionally, a preset text encoding module is used to perform text encoding processing on the text keywords of the selected text sensitive information to obtain a text keyword encoding vector. The text encoding module can be a Qwen3-Embedding encoding module.
[0081] By encoding the text keywords of the selected sensitive information, a text keyword encoding vector is obtained, which can accurately express the true semantic category of the sensitive information, thus enabling users to customize desensitized content according to the relevant category.
[0082] S402. According to the preset text encoding module, the text value of the selected sensitive information is processed by text encoding to obtain the text value encoding vector.
[0083] Optionally, a preset text encoding module is used to perform text encoding processing on the text values of the sensitive information to be selected, to obtain a text value encoding vector. The text encoding module can be a Qwen3-Embedding encoding module.
[0084] S403. Based on the text keyword encoding vector and / or text value encoding vector, perform a search from the sensitive information database to obtain the search results.
[0085] Optionally, after obtaining the text keyword encoding vector and the text value encoding vector, a vector search can be performed from the sensitive information database based on the text keyword encoding vector and / or the text value encoding vector to obtain the search results.
[0086] In one example, a vector cosine similarity search is performed from a sensitive information database based on the text keyword encoding vector to obtain the search results. The search results include multiple candidate vectors similar to the text keyword encoding vector and their corresponding similarity scores.
[0087] In another example, a vector cosine similarity search is performed from a sensitive information database based on the text value encoded vector to obtain the search results. The search results include multiple candidate vectors similar to the text value encoded vector and their corresponding similarity scores.
[0088] In another example, a vector cosine similarity search is performed from a sensitive information database based on the text keyword encoding vector and the text value encoding vector to obtain the search results. The search results include multiple candidate vectors similar to the text keyword encoding vector and the text value encoding vector, along with their corresponding similarity results.
[0089] S404. Based on the search results, determine whether to treat the candidate text sensitive information as a target text sensitive information. If so, treat the candidate text sensitive information as a target text sensitive information.
[0090] In one example, the search results are judged sequentially according to a preset similarity threshold. If the similarity result is greater than the preset similarity threshold, the corresponding candidate vector is regarded as sensitive information of the target text.
[0091] In another example, the maximum similarity result in the search results is determined. If the maximum similarity result is greater than a preset similarity threshold, the corresponding candidate vector is taken as sensitive information of the target text.
[0092] By encoding text keywords into encoding vectors and text encoding vectors, and then retrieving them from a sensitive information database based on the text keyword encoding vectors and / or text encoding vectors, the search results can be obtained. Based on the search results, the target text sensitive information can be identified, which can reduce the false positive rate of target text sensitive information and improve the search recall efficiency.
[0093] In one possible implementation, Figure 5 This is a schematic diagram illustrating the process of obtaining desensitized engineering drawings in the engineering drawing information desensitization method based on a multimodal large model provided in this application embodiment, with reference to... Figure 5 As shown, in step S104 above, the engineering drawings to be desensitized are processed according to the desensitization method and at least one target sensitive information to obtain the desensitized engineering drawings, including: S501. If the desensitization method is pixel coverage, then based on the target bounding box in the target sensitive information and the preset edge operator, determine the foreground outline and background area corresponding to the target bounding box from the engineering drawing to be desensitized.
[0094] Optionally, if the desensitization method is pixel coverage, then based on the target bounding box in the target sensitive information and the preset edge operator, the foreground outline and background area corresponding to the target bounding box are determined from the engineering drawing to be desensitized, so as to achieve accurate positioning of the sensitive content boundary.
[0095] The preset edge operator can be the Canny edge detection operator.
[0096] For example, the foreground contour and background area corresponding to the target bounding box can be identified from the engineering drawing to be desensitized by using an edge operator according to the target bounding box in the target sensitive information.
[0097] S502. Cluster the background area to determine the dominant background color.
[0098] Optionally, the background area is clustered according to color values, and the dominant background color is determined based on the number of pixels of each color after clustering, thereby achieving adaptive color selection.
[0099] S503. Fill the foreground outline corresponding to the target bounding box with the dominant background color to obtain the desensitized engineering drawing of the target sensitive information. After all the target sensitive information has been processed, the desensitized engineering drawing is obtained.
[0100] Optionally, the foreground outline corresponding to the target bounding box can be filled with the dominant background color to obtain an anonymized engineering drawing of the target's sensitive information, thereby achieving a natural transition and reducing visual abruptness. Furthermore, it can adapt to diverse background environments, avoiding color overflow or color difference issues, while preserving background integrity.
[0101] Optionally, after all sensitive information has been processed, desensitized engineering drawings are obtained.
[0102] Optionally, after filling the foreground outline corresponding to the target bounding box with the dominant background color, the extracted foreground outline can be superimposed in a non-intrusive manner to record or identify the location of the original sensitive information.
[0103] In one possible implementation, Figure 6 This is a flowchart illustrating the process of determining the foreground contour and background region corresponding to the target bounding box in the engineering drawing information desensitization method based on a multimodal large model provided in this application embodiment. (Refer to...) Figure 6 As shown, in step S501 above, based on the target bounding box in the target sensitive information and the preset edge operator, the foreground contour and background region corresponding to the target bounding box are determined from the engineering drawing to be desensitized, including: S601. Based on the target bounding box in the target sensitive information, obtain the local area from the engineering drawing to be desensitized.
[0104] Optionally, a local area can be obtained from the engineering drawing to be desensitized based on the target bounding box in the target sensitive information, thereby narrowing the processing scope, avoiding full-image processing of the entire drawing, and improving computational efficiency. Here, the local area refers to the local image indicated by the target bounding box.
[0105] S602. Within a local region, edge detection is performed using edge operators to generate a binary edge map.
[0106] Optionally, within a local area, edge operators are used to calculate the gradient magnitude and direction of each pixel in the image, and edge points are filtered out according to a preset double threshold to generate a binary edge map, thereby detecting finer and continuous edges. Furthermore, the double threshold mechanism can effectively distinguish between strong and weak edges and reduce noise interference.
[0107] S603. Perform morphological closing operation on the binary edge map to obtain the foreground contour, and determine the background region based on the foreground contour.
[0108] Optionally, a morphological closing operation of first dilation and then erosion is performed on the binary edge map to connect the originally broken edges into closed contours. In the closed contour map, all closed contours are extracted by a contour search algorithm to obtain the foreground contour. The foreground contour is then inverted to obtain the background region.
[0109] In one possible implementation, the clustering of the background region in S502 above to determine the dominant background color includes: Cluster the background region to determine at least one cluster and the number of pixels in each cluster. Use the color of the cluster center of the cluster with the most pixels as the dominant background color.
[0110] Optionally, the background area is clustered according to the RGB color values of the pixels to obtain at least one cluster, and the number of pixels in each cluster is determined. The color of the cluster center of the cluster with the most pixels is used as the dominant background color, thereby maximizing the integration of the covered area with the original background and reducing visual abruptness.
[0111] For example, the color of the cluster center can be the average color of all pixels in the cluster.
[0112] In one possible implementation, Figure 7 This is another flowchart illustrating the process of obtaining desensitized engineering drawings in the engineering drawing information desensitization method based on a multimodal large model provided in this application embodiment, referring to... Figure 7 As shown, in step S104 above, the engineering drawings to be desensitized are processed according to the desensitization method and at least one target sensitive information to obtain the desensitized engineering drawings, including: S701. If the desensitization method is watermark desensitization, then based on the target bounding box in the target sensitive information, determine the target area corresponding to the target bounding box from the engineering drawing to be desensitized.
[0113] Optionally, if the desensitization method is watermark desensitization, then based on the target bounding box in the target sensitive information, the target area corresponding to the target bounding box is determined from the engineering drawing to be desensitized. The target area is the location of the target watermark image.
[0114] S702, Obtain the target watermark image.
[0115] The target watermark image includes a transparency channel and a watermark color channel.
[0116] Specifically, the alpha channel records the opacity information of each pixel in the watermark, which is used to determine the overlay strength of the watermark during subsequent reverse decoding. The RGB channel records the color information of the watermark itself.
[0117] S703. Spatial registration is performed on the target watermark image according to the target bounding box to obtain the registered target watermark image.
[0118] Optionally, the target watermark image is geometrically transformed according to the coordinates and size of the target bounding box to obtain the registered target watermark image, thereby ensuring that the target watermark image is precisely aligned with the actual watermark position.
[0119] Specifically, geometric transformations include scaling, translation, rotation, and affine or perspective transformations. Scaling ensures the watermark image size matches the target bounding box size. Translation places the watermark image in the correct coordinate position. Rotation corrects the angle of tilted watermarks. Affine or perspective transformations accommodate more complex deformations.
[0120] S704. Based on the registered target watermark image and the target bounding box, obtain the desensitized engineering drawings.
[0121] Optionally, based on a preset linear overlay model and the registered target watermark image, the target bounding box is inversely solved to obtain the original image content, and the original image content is edge-blending processed to obtain the desensitized engineering drawing.
[0122] For example, the reverse engineering calculation can be performed pixel by pixel according to the reverse engineering formula to restore the original image, and the reverse engineering result can be numerically cropped to ensure that the pixel value is within the effective range, thereby obtaining the desensitized engineering drawing, which can effectively restore the original image content.
[0123] Optionally, the boundary between the reverse-resolution region and the surrounding background can be feathered or smoothed to eliminate splicing marks, and the processed target region can be pasted back to the corresponding position in the original engineering drawing to generate an engineering drawing with desensitized target sensitive information.
[0124] Optionally, after all sensitive information has been processed, desensitized engineering drawings are obtained.
[0125] In one possible implementation, determining at least one target sensitive information in step S103 based on multiple candidate image sensitive information and multiple candidate text sensitive information includes: In response to a user's selection operation for at least one candidate image sensitive information and at least one candidate text sensitive information among a plurality of candidate image sensitive information, each selected candidate image sensitive information and each selected candidate text sensitive information is treated as a target sensitive information.
[0126] Optionally, after obtaining multiple candidate image sensitive information and multiple candidate text sensitive information, the multiple candidate image sensitive information and multiple candidate text sensitive information can be displayed in a graphical user interface, and in response to the user's selection operation on at least one candidate image sensitive information and at least one candidate text sensitive information among the multiple candidate image sensitive information and multiple candidate text sensitive information, the selected candidate image sensitive information and the selected candidate text sensitive information are respectively regarded as target sensitive information, thereby obtaining target sensitive information.
[0127] In one example, in response to a user's selection operation for at least one of a plurality of candidate image sensitive information, each selected candidate image sensitive information is treated as a target sensitive information.
[0128] In another example, in response to a user's selection operation for at least one of a plurality of candidate text sensitive information, each selected candidate text sensitive information is treated as a target sensitive information.
[0129] By responding to a user's selection of at least one of a plurality of candidate image sensitive information and at least one of a plurality of candidate text sensitive information, the target sensitive information can be determined, thereby improving the user experience and increasing the flexibility of the desensitization process.
[0130] Based on the same inventive concept, this application also provides an engineering drawing information desensitization device based on a multimodal large model, which corresponds to the engineering drawing information desensitization method based on a multimodal large model. Since the principle of the device in this application is similar to the engineering drawing information desensitization method based on a multimodal large model described above, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0131] Reference Figure 8 As shown, Figure 8 This is a schematic diagram of an engineering drawing information desensitization device based on a multimodal large model provided in an embodiment of this application. The device includes: an acquisition module 801, a semantic screening module 802, a determination module 803, and a desensitization processing module 804. The acquisition module 801 is used to acquire the engineering drawings to be desensitized and the desensitization method. The desensitization method includes at least one of the following: pixel overlay, watermark desensitization, and Gaussian blur coding. The semantic screening module 802 is used to input the engineering drawings to be desensitized and the preset semantic screening prompts into the pre-trained multimodal large model. The multimodal large model performs semantic screening to obtain multiple candidate image sensitive information and multiple candidate text sensitive information corresponding to the engineering drawings to be desensitized. The candidate image sensitive information includes: image keywords and image bounding boxes. The candidate text sensitive information includes: text keywords, text values and text bounding boxes. The determining module 803 is used to determine at least one target sensitive information based on multiple candidate image sensitive information and multiple candidate text sensitive information; The desensitization processing module 804 is used to desensitize the engineering drawings to be desensitized according to the desensitization method and at least one target sensitive information, so as to obtain the desensitized engineering drawings.
[0132] Optionally, module 803 is specifically used for: Based on a preset sensitive information database, at least one target image sensitive information is determined from multiple candidate image sensitive information, and / or, based on a preset sensitive information database, at least one target text sensitive information is determined from multiple candidate text sensitive information; The sensitive information of each target image and the sensitive information of each target text are treated as a single target sensitive information.
[0133] Optionally, module 803 is specifically used for: Based on the preset text encoding module, the image keywords of the sensitive information of the selected image are processed by text encoding to obtain the image keyword encoding vector; Based on the image bounding box of the sensitive information of the candidate image, the region image of the sensitive information of the candidate image is obtained by cropping from the engineering drawing to be desensitized; Based on the preset image encoding module, the region image of the selected image with sensitive information is processed by image encoding to obtain the image encoding vector; Based on the image keyword encoding vector and / or image encoding vector, a search is conducted from the sensitive information database to obtain search results; Based on the search results, determine whether to use the sensitive information of the candidate image as the sensitive information of the target image. If so, use the sensitive information of the candidate image as the sensitive information of the target image.
[0134] Optionally, module 803 is specifically used for: Based on the preset text encoding module, the text keywords of the selected sensitive text information are processed by text encoding to obtain the text keyword encoding vector; Based on the preset text encoding module, the text values of the selected sensitive information are processed by text encoding to obtain the text value encoding vector; Based on the text keyword encoding vector and / or text value encoding vector, a search is performed from the sensitive information database to obtain search results; Based on the search results, determine whether to treat the candidate sensitive text as a target sensitive text. If so, treat the candidate sensitive text as a target sensitive text.
[0135] Optionally, the desensitization processing module 804 is specifically used for: If the desensitization method is pixel coverage, then based on the target bounding box in the target sensitive information and the preset edge operator, the foreground outline and background area corresponding to the target bounding box are determined from the engineering drawing to be desensitized. Cluster the background areas to determine the dominant background color; By filling the foreground outline corresponding to the target bounding box with the dominant background color, the desensitized engineering drawing of the target sensitive information is obtained, and the desensitized engineering drawing is obtained after all the target sensitive information has been processed.
[0136] Optionally, the desensitization processing module 804 is specifically used for: Based on the target bounding box in the target sensitive information, obtain the local area from the engineering drawing to be desensitized; Within a local area, edge detection is performed using edge operators to generate a binary edge map; A morphological closing operation is performed on the binary edge map to obtain the foreground contour, and the background region is determined based on the foreground contour.
[0137] Optionally, the desensitization processing module 804 is specifically used for: Cluster the background region to determine at least one cluster and the number of pixels in each cluster; The color of the cluster center of the cluster with the most pixels is used as the dominant background color.
[0138] Optionally, the desensitization processing module 804 is specifically used for: If the desensitization method is watermark desensitization, then the target area corresponding to the target bounding box is determined from the engineering drawing to be desensitized based on the target bounding box in the target sensitive information; Obtain the target watermark image, which includes the transparency channel and the watermark color channel; Spatial registration is performed on the target watermark image according to the target bounding box to obtain the registered target watermark image. Based on the registered target watermark image and the target bounding box, the desensitized engineering drawings are obtained.
[0139] Optionally, module 803 is specifically used for: In response to a user's selection operation for at least one candidate image sensitive information and at least one candidate text sensitive information among a plurality of candidate image sensitive information, each selected candidate image sensitive information and each selected candidate text sensitive information is treated as a target sensitive information.
[0140] The processing flow of each module in the device and the interaction flow between each module can be referred to the relevant descriptions in the above method embodiments, and will not be detailed here.
[0141] This application also provides an electronic device, such as... Figure 9 As shown, Figure 9The schematic diagram of the electronic device structure provided in this application embodiment includes: a processor 901 and a memory 902, and optionally, a bus 903. The memory 902 stores machine-readable instructions executable by the processor 901. When the electronic device is running, the processor 901 and the memory 902 communicate through the bus 903, and the processor 901 executes the machine-readable instructions to perform the steps of the above-described method for desensitizing engineering drawing information based on a multimodal large model.
[0142] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the above-described method for desensitizing engineering drawing information based on a multimodal large model.
[0143] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the method embodiments, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some communication interfaces; the indirect coupling or communication connection of devices or modules can be electrical, mechanical, or other forms.
[0144] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.
[0145] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A method for desensitizing engineering drawing information based on a multimodal large model, characterized in that, include: Obtain the engineering drawings to be desensitized and the desensitization method, wherein the desensitization method includes at least one of the following: pixel overlay, watermark desensitization, and Gaussian blurring. The engineering drawings to be desensitized and the preset semantic screening prompts are input into a pre-trained multimodal large model. The multimodal large model performs semantic screening to obtain multiple candidate image sensitive information and multiple candidate text sensitive information corresponding to the engineering drawings to be desensitized. The candidate image sensitive information includes: image keywords and image bounding boxes. The candidate text sensitive information includes: text keywords, text values and text bounding boxes. Based on the multiple candidate image sensitive information and the multiple candidate text sensitive information, at least one target sensitive information is determined; Based on the desensitization method and the at least one target sensitive information, the engineering drawing to be desensitized is subjected to desensitization processing to obtain the desensitized engineering drawing; The process of desensitizing the engineering drawing to be desensitized according to the desensitization method and the at least one target sensitive information to obtain the desensitized engineering drawing includes: If the desensitization method is pixel coverage, then based on the target bounding box in the target sensitive information and the preset edge operator, the foreground contour and background area corresponding to the target bounding box are determined from the engineering drawing to be desensitized; the background area is clustered to determine the dominant background color; the foreground contour corresponding to the target bounding box is filled with the dominant background color to obtain the desensitized engineering drawing, and after all target sensitive information has been processed, the desensitized engineering drawing is obtained.
2. The method according to claim 1, characterized in that, The step of determining at least one target sensitive information based on the multiple candidate image sensitive information and the multiple candidate text sensitive information includes: Based on a preset sensitive information database, at least one target image sensitive information is determined from the plurality of candidate image sensitive information, and / or, based on a preset sensitive information database, at least one target text sensitive information is determined from the plurality of candidate text sensitive information; Each of the target image sensitive information and each of the target text sensitive information is treated as a single target sensitive information.
3. The method according to claim 2, characterized in that, The step of determining at least one target image sensitive information from the plurality of candidate image sensitive information according to a preset sensitive information database includes: According to the preset text encoding module, the image keywords of the sensitive information of the candidate image are processed by text encoding to obtain the image keyword encoding vector; Based on the image bounding box of the sensitive information of the candidate image, the region image of the sensitive information of the candidate image is obtained by cropping from the engineering drawing to be desensitized; According to the preset image encoding module, the region image of the sensitive information of the candidate image is subjected to image encoding processing to obtain the image encoding vector; Based on the image keyword encoding vector and / or the image encoding vector, a search is performed from the sensitive information database to obtain search results; Based on the search results, determine whether to use the candidate image sensitive information as target image sensitive information. If so, use the candidate image sensitive information as target image sensitive information.
4. The method according to claim 2, characterized in that, The step of determining at least one target text sensitive information from the plurality of candidate text sensitive information according to a preset sensitive information database includes: According to the preset text encoding module, the text keywords of the sensitive information of the candidate text are processed by text encoding to obtain the text keyword encoding vector; According to the preset text encoding module, the text value of the sensitive information of the candidate text is processed by text encoding to obtain the text value encoding vector; Based on the text keyword encoding vector and / or the text value encoding vector, a search is performed from the sensitive information database to obtain search results; Based on the search results, determine whether to use the candidate text sensitive information as target text sensitive information. If so, use the candidate text sensitive information as target text sensitive information.
5. The method according to claim 1, characterized in that, The step of determining the foreground contour and background region corresponding to the target bounding box from the engineering drawing to be desensitized, based on the target bounding box in the target sensitive information and a preset edge operator, includes: Based on the target bounding box in the target sensitive information, obtain the local area from the engineering drawing to be desensitized; Within the local area, edge detection is performed using the edge operator to generate a binary edge map; A morphological closing operation is performed on the binary edge map to obtain the foreground contour, and the background region is determined based on the foreground contour.
6. The method according to claim 1, characterized in that, The step of clustering the background region to determine the dominant background color includes: The background region is clustered to determine at least one cluster and the number of pixels in each cluster; The color of the cluster center of the cluster with the most pixels is used as the dominant background color.
7. The method according to claim 1, characterized in that, The process of desensitizing the engineering drawing to be desensitized according to the desensitization method and the at least one target sensitive information to obtain the desensitized engineering drawing includes: If the desensitization method is watermark desensitization, then based on the target bounding box in the target sensitive information, the target area corresponding to the target bounding box is determined from the engineering drawing to be desensitized; Acquire a target watermark image, the target watermark image including a transparency channel and a watermark color channel; Spatial registration is performed on the target watermark image according to the target bounding box to obtain the registered target watermark image; Based on the registered target watermark image and the target bounding box, the desensitized engineering drawing is obtained.
8. The method according to claim 1, characterized in that, The step of determining at least one target sensitive information based on the multiple candidate image sensitive information and the multiple candidate text sensitive information includes: In response to a user's selection operation for at least one of the plurality of candidate image sensitive information and at least one of the plurality of candidate text sensitive information, each selected candidate image sensitive information and each selected candidate text sensitive information is treated as a target sensitive information.
9. An electronic device, characterized in that, include: The device includes a processor and a memory, the memory storing machine-readable instructions executable by the processor. When the electronic device is running, the processor executes the machine-readable instructions to perform the steps of the method for desensitizing engineering drawing information based on a multimodal large model as described in any one of claims 1 to 8.