Image-text sign detection method and system based on large model and storage medium
Through the large-model-based graphic and text logo detection method, the graphic and text logo detection is automatically performed, which solves the problem of inefficient detection in the existing technology and realizes efficient and automatic detection result generation.
Patent Information
- Application Number
- CN202510126950.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-05-30
AI Technical Summary
The inefficient detection of graphic and text marks in the prior art is mainly due to the reliance on manual analysis and detection.
The graphic and text logo detection method based on the big model is adopted, and the image of the document to be detected is obtained, the feature vector transformation is performed, the vector similarity is calculated, the associated document is determined, and the semantic recognition is performed through the pre-trained big model to automatically generate the detection results.
Without manual intervention, the efficiency of graphic and text logo detection is significantly improved, and the risk pictures and associated documents can be automatically determined, and the detection results can be accurately generated.
Smart Images

Figure CN120071374A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a graphic logo detection method, system, and storage medium based on a large model. Background Art
[0002] With the rapid development of Internet technology and social economy, various graphic logos have emerged. Along with the increasingly frequent use of graphic logos in documents, the detection of the infringement use of graphic logos has attracted more and more attention.
[0003] In the existing graphic logo detection process, generally, manual analysis and detection are used, resulting in low efficiency of graphic logo detection. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a graphic logo detection method, system, and storage medium based on a large model to solve the problem of low efficiency of graphic logo detection in the prior art.
[0005] The embodiments of the present invention are implemented as follows. A graphic logo detection method based on a large model, the method includes:
[0006] Obtain a document to be detected, and perform image extraction on the document to be detected to obtain a to-be-detected image;
[0007] Obtain a target graphic logo, and perform feature vector conversion on the target graphic logo and the to-be-detected image to obtain a target graphic vector and a to-be-detected vector;
[0008] Calculate the vector similarity between the target graphic vector and the to-be-detected vector, and determine a risk image and the associated document of the risk image in the to-be-detected document according to the vector similarity;
[0009] Input the associated document into a pre-trained large model for semantic recognition to obtain document semantics, and determine a detection score corresponding to the risk image according to the document semantics;
[0010] Generate a graphic logo detection result of the to-be-detected document according to the detection score.
[0011] Preferably, before inputting the associated document into the pre-trained large model for semantic recognition, it further includes:
[0012] Obtain a sample document, and input the sample document into the large model for word segmentation to obtain sample word segments;
[0013] Perform vector conversion on the sample word segments to obtain word segment vectors, and perform attention mechanism processing on the word segment vectors to obtain attention features;
[0014] Perform a non - linear transformation on the attention feature, and perform normalization processing on the attention feature after the non - linear transformation to obtain a normalized feature;
[0015] Perform residual processing on the normalized feature to obtain a residual feature, and perform feature decoding on the residual feature to obtain a decoded feature;
[0016] Determine the model loss according to the decoded feature, and update the parameters of the large model according to the model loss until the large model converges to obtain the pre - trained large model.
[0017] Preferably, determining the detection score corresponding to the risk picture according to the document semantics includes:
[0018] Perform entity recognition on the semantic vocabulary in the document semantics, and determine the object name and object description vocabulary in the document semantics according to the entity recognition result;
[0019] Determine the descriptive semantic database according to the entity type of the object description vocabulary, and calculate the similarity between the object description vocabulary and the semantic vocabulary in the descriptive semantic database to obtain the vocabulary similarity;
[0020] Determine the detection score according to the object name and the vocabulary similarity.
[0021] Preferably, determining the detection score according to the object name and the vocabulary similarity includes:
[0022] Perform name matching between the object name and the target name corresponding to the target graphic - text logo;
[0023] If the name matching fails, determine the preset score as the detection score;
[0024] If the name matching is successful, calculate the sum of the maximum vocabulary similarities corresponding to different object description vocabularies to obtain the total similarity;
[0025] Calculate the average similarity of the object description words according to the total similarity, and match the average similarity with the score query table to obtain the detection score.
[0026] Preferably, generating the graphic - text logo detection result of the document to be detected according to the detection score includes:
[0027] If the detection score is greater than or equal to the score threshold, it is determined that the graphic - text logo detection of the document to be detected is qualified, and there is no infringement risk for the document to be detected;
[0028] If the detection score is greater than or equal to the preset score and less than the score threshold, it is determined that the graphic and text logo detection of the document to be detected is unqualified, and the document to be detected has a first-level infringement risk;
[0029] If the detection score is less than or equal to the score threshold, it is determined that the graphic and text logo detection of the document to be detected is unqualified, and the document to be detected has a second-level infringement risk.
[0030] Preferably, determining the risk picture and the associated document of the risk picture in the document to be detected according to the vector similarity includes:
[0031] If any of the vector similarities is greater than the similarity threshold, the picture to be detected corresponding to the vector similarity is determined as the risk picture, and the picture identifier of the risk picture in the document to be detected is obtained;
[0032] Determine the picture description paragraph according to the picture identifier, and determine the picture description paragraph as the associated document corresponding to the risk picture.
[0033] Preferably, determining the picture description paragraph according to the picture identifier includes:
[0034] Match the picture identifier with the document to be detected for identifier matching, and determine the identifier statement according to the identifier matching result;
[0035] Match the identifier statement with the preset identifier vocabulary, and determine the paragraph corresponding to the identifier statement that matches the preset identifier vocabulary as the picture description paragraph.
[0036] Another object of the embodiments of the present invention is to provide a graphic and text logo detection system based on a large model, and the system includes:
[0037] A picture extraction module, configured to obtain a document to be detected, and perform picture extraction on the document to be detected to obtain pictures to be detected;
[0038] A vector conversion module, configured to obtain a target graphic and text logo, and perform feature vector conversion on the target graphic and text logo and the pictures to be detected to obtain a target graphic and text vector and a vector to be detected;
[0039] An association determination module, configured to calculate the vector similarity between the target graphic and text vector and the vector to be detected, and determine the risk picture and the associated document of the risk picture in the document to be detected according to the vector similarity;
[0040] A semantic recognition module, configured to input the associated document into a pre-trained large model for semantic recognition to obtain document semantics, and determine the detection score corresponding to the risk picture according to the document semantics;
[0041] A result generation module, configured to generate a graphic and text logo detection result of the document to be detected according to the detection score.
[0042] Preferably, the semantic recognition module is further configured to:
[0043] Obtain a sample document, input the sample document into the large model for word segmentation, and obtain sample word segments;
[0044] Perform vector conversion on the sample word segments to obtain word segment vectors, and perform attention mechanism processing on the word segment vectors to obtain attention features;
[0045] Perform a non-linear transformation on the attention features, and perform normalization processing on the attention features after the non-linear transformation to obtain normalized features;
[0046] Perform residual processing on the normalized features to obtain residual features, and perform feature decoding on the residual features to obtain decoded features;
[0047] Determine a model loss according to the decoded features, and update the parameters of the large model according to the model loss until the large model converges, so as to obtain the pre-trained large model.
[0048] In the embodiment of the present invention, by extracting pictures from the document to be detected, the pictures to be detected in the document to be detected can be automatically obtained. By performing feature vector conversion on the target graphic and text logo and the pictures to be detected, the pictures can be effectively converted into vectors to obtain the target graphic and text vector and the vector to be detected. By calculating the vector similarity between the target graphic and text vector and the vector to be detected, the associated documents in the document to be detected can be effectively determined based on the vector similarity. By inputting the associated documents into the pre-trained large model for semantic recognition, the document semantics of the associated documents can be effectively extracted. Based on the document semantics, the detection score of the risk pictures can be automatically determined, and based on the detection score, the graphic and text logo detection result of the document to be detected can be automatically generated, without using manual methods for graphic and text logo detection, improving the efficiency of graphic and text logo detection. Description of the Drawings
[0049] Figure 1 is a flowchart of a graphic and text logo detection method based on a large model provided by the first embodiment of the present invention;
[0050] Figure 2 is a schematic structural diagram of a graphic and text logo detection system based on a large model provided by the second embodiment of the present invention;
[0051] Figure 3 is a specific implementation schematic diagram of a graphic and text logo detection system based on a large model provided by the second embodiment of the present invention;
[0052] Figure 4 It is a schematic structural diagram of a terminal device provided by the third embodiment of the present invention. Specific embodiments
[0053] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0054] In order to illustrate the technical solutions described in the present invention, specific embodiments will be used for illustration below.
[0055] Embodiment 1
[0056] Please refer to Figure 1 , which is a flowchart of a graphic and text logo detection method based on a large model provided by the first embodiment of the present invention. The graphic and text logo detection method based on a large model can be applied to any device or system. The graphic and text logo detection method based on a large model includes the following steps:
[0057] Step S10: Obtain a document to be detected, and extract pictures from the document to be detected to obtain pictures to be detected;
[0058] Among them, the document to be detected is obtained by web crawling public network articles. By performing document analysis on the document to be detected, pictures in the document to be detected are extracted based on the document analysis results to obtain pictures to be detected.
[0059] Step S20: Obtain a target graphic and text logo, and perform feature vector conversion on the target graphic and text logo and the pictures to be detected to obtain a target graphic and text vector and a vector to be detected;
[0060] Among them, the target graphic and text logo can be set according to requirements. The purpose of the graphic and text logo detection method based on a large model is to automatically determine whether there is an infringing use of the target graphic and text logo in the document to be detected. The target graphic and text logo can be any information such as a logo of a graphic or text.
[0061] In this step, based on a preset vector mapping relationship, feature vector conversion is performed on the target graphic and text logo and the pictures to be detected to obtain a target graphic and text vector and a vector to be detected. By converting the target graphic and text logo and the pictures to be detected into vectors, it effectively facilitates the determination of subsequent risk pictures and associated documents.
[0062] Step S30: Calculate the vector similarity between the target graphic and text vector and the vector to be detected, and determine the risk pictures and the associated documents of the risk pictures in the document to be detected according to the vector similarity;
[0063] Among them, the distance between the target graphic-text vector and the vector to be detected is calculated through the Euclidean distance formula to obtain the vector similarity. Based on the vector similarity, the risk pictures and the associated documents of the risk pictures in the document to be detected can be automatically determined, improving the detection efficiency of graphic-text marks.
[0064] Optionally, determining the risk pictures and the associated documents of the risk pictures in the document to be detected according to the vector similarity includes:
[0065] If any of the vector similarities is greater than the similarity threshold, the picture to be detected corresponding to the vector similarity is determined as a risk picture, and the picture identifier of the risk picture in the document to be detected is obtained; among them, the similarity threshold can be set according to requirements. When the vector similarity is greater than the similarity threshold, it is determined that the picture to be detected has a high similarity with the target graphic-text mark, and there is a risk of infringing use of the picture to be detected. Therefore, the picture to be detected corresponding to the vector similarity is determined as a risk picture. The picture identifier is used to represent the index number of the risk picture in the document to be detected. For example, the picture identifier can be identifiers such as "Figure a", "Figure b", or "Figure c", etc.;
[0066] Determine the picture description paragraph according to the picture identifier, and determine the picture description paragraph as the associated document corresponding to the risk picture; among them, based on the picture identifier, the description information of the risk picture in the document to be detected is determined to obtain the picture description paragraph.
[0067] Further, determining the picture description paragraph according to the picture identifier includes:
[0068] Match the picture identifier with the document to be detected for identifier matching, and determine the identifier statement according to the identifier matching result; among them, by matching the picture identifier with the document to be detected for identifier matching, the position of the picture identifier in the document to be detected can be effectively located. Based on the position positioning result of the picture identifier, the identifier statement can be effectively determined, that is, the statement where the located picture identifier is located is determined as the identifier statement;
[0069] Match the identifier statement with the preset identifier vocabulary, and determine the paragraph corresponding to the identifier statement that matches the preset identifier vocabulary as the picture description paragraph; among them, the preset identifier vocabulary can be set according to requirements. For example, the preset identifier vocabulary can be set as words such as "Please refer to", "Please consult", "Please view", etc. By matching the identifier statement with the preset identifier vocabulary, it is determined whether the identifier statement is a prompt word for reminding the user to view. When the identifier statement matches the preset identifier vocabulary, the paragraph corresponding to the identifier statement is the paragraph for describing the risk picture. Therefore, the paragraph corresponding to the identifier statement that matches the preset identifier vocabulary is determined as the picture description paragraph.
[0070] Step S40: Input the associated document into a pre-trained large model for semantic recognition to obtain the document semantics, and determine the detection score corresponding to the risk picture according to the document semantics;
[0071] Among them, by inputting the associated document into a pre-trained large model for semantic recognition, the document semantics of the associated document can be effectively extracted. Based on the document semantics, the detection score of the risk picture can be automatically determined. The detection score is used to represent the risk degree of infringement citation for the risk picture. When the detection score is larger, the risk degree of infringement citation for the risk picture is smaller; when the detection score is smaller, the risk degree of infringement citation for the risk picture is larger.
[0072] Optionally, before inputting the associated document into a pre-trained large model for semantic recognition, it further includes:
[0073] Obtain a sample document, and input the sample document into the large model for word segmentation to obtain sample word segments; among them, by performing word segmentation on the sample document, it effectively facilitates the subsequent vector conversion operation of the sample word segments;
[0074] Perform vector conversion on the sample word segments to obtain word segment vectors, and perform attention mechanism processing on the word segment vectors to obtain attention features; among them, by performing attention mechanism processing on the word segment vectors, the context features of the word segment vectors can be effectively fused to obtain attention features;
[0075] Perform a non-linear transformation on the attention features, and perform normalization processing on the attention features after the non-linear transformation to obtain normalized features; among them, by performing normalization processing on the attention features after the non-linear transformation, it effectively reduces the dimension of the attention features;
[0076] Perform residual processing on the normalized features to obtain residual features, perform feature decoding on the residual features to obtain decoded features, determine the model loss according to the decoded features, and update the parameters of the large model according to the model loss until the large model converges to obtain the pre-trained large model.
[0077] Furthermore, determining the detection score corresponding to the risk picture according to the document semantics includes:
[0078] Perform entity recognition on the semantic vocabulary in the document semantics, and determine the object name and object description vocabulary in the document semantics according to the entity recognition results; among them, the object name is used to represent the name of the object corresponding to the risk picture, and this name can be a person's name, a company name, or other names, etc. The object description vocabulary is used to represent the introduction information of the corresponding object;
[0079] Determine the descriptive semantic database according to the entity type of the object description vocabulary, calculate the similarity between the object description vocabulary and the semantic vocabulary in the descriptive semantic database to obtain the vocabulary similarity, and determine the detection score according to the object name and the vocabulary similarity; wherein, match the entity type of the object description vocabulary with the database query table to obtain the descriptive semantic database, and the database query table stores the corresponding relationship between the entity types of different object description vocabularies and the corresponding descriptive semantic databases, and calculate the similarity between the object description vocabulary and the semantic vocabulary in the descriptive semantic database to obtain the vocabulary similarity.
[0080] Further, determining the detection score according to the object name and the vocabulary similarity includes:
[0081] Perform name matching between the object name and the target name corresponding to the target graphic logo;
[0082] If the name matching fails, determine the preset score as the detection score; wherein, the preset score can be set according to requirements;
[0083] If the name matching is successful, calculate the sum of the corresponding maximum vocabulary similarities among different object description vocabularies to obtain the total similarity;
[0084] Calculate the average similarity of the object description words according to the total similarity, and match the average similarity with the score query table to obtain the detection score; wherein, the score query table stores the corresponding relationship between different average similarities and the corresponding detection scores.
[0085] Step S50, generate the graphic logo detection result of the document to be detected according to the detection score.
[0086] Optionally, generating the graphic logo detection result of the document to be detected according to the detection score includes:
[0087] If the detection score is greater than or equal to the score threshold, determine that the graphic logo detection of the document to be detected is qualified, and there is no infringement risk for the document to be detected;
[0088] If the detection score is greater than or equal to the preset score and less than the score threshold, determine that the graphic logo detection of the document to be detected is unqualified, and the document to be detected has a first-level infringement risk;
[0089] If the detection score is less than or equal to the score threshold, determine that the graphic logo detection of the document to be detected is unqualified, and the document to be detected has a second-level infringement risk;
[0090] Among them, when the object name corresponding to the risk picture matches the target name corresponding to the target graphic logo, and the object description words have a high similarity with the descriptions in the description semantic database, it is determined that the risk picture is a normal reference. Therefore, it is determined that the graphic logo detection of the document to be detected is qualified, and there is no infringement risk in the document to be detected. When the detection score is greater than or equal to the preset score and less than the score threshold, it is determined that the similarity between the object description words and the descriptions in the description semantic database is low. Therefore, it is determined that the graphic logo detection of the document to be detected is unqualified, and there is a first-level infringement risk in the document to be detected. The degree of the first-level infringement risk is lower than that of the second-level infringement risk.
[0091] In this embodiment, by extracting pictures from the document to be detected, the pictures to be detected in the document to be detected can be automatically obtained. By performing feature vector conversion on the target graphic logo and the pictures to be detected, the pictures can be effectively converted into vectors to obtain the target graphic vector and the vectors to be detected. By calculating the vector similarity between the target graphic vector and the vectors to be detected, the associated documents in the document to be detected can be effectively determined based on the vector similarity. By inputting the associated documents into a pre-trained large model for semantic recognition, the document semantics of the associated documents can be effectively extracted. Based on the document semantics, the detection score of the risk pictures can be automatically determined, and based on the detection score, the graphic logo detection result of the document to be detected can be automatically generated, without the need for manual graphic logo detection, improving the graphic logo detection efficiency.
[0092] Embodiment 2
[0093] Please refer to Figure 2 , which is a schematic structural diagram of the graphic logo detection system 100 based on a large model provided by the second embodiment of the present invention, including:
[0094] The picture extraction module 10 is used to obtain the document to be detected and extract pictures from the document to be detected to obtain the pictures to be detected.
[0095] The vector conversion module 11 is used to obtain the target graphic logo and perform feature vector conversion on the target graphic logo and the pictures to be detected to obtain the target graphic vector and the vectors to be detected.
[0096] The association determination module 12 is used to calculate the vector similarity between the target graphic vector and the vectors to be detected and determine the risk pictures and the associated documents of the risk pictures in the document to be detected according to the vector similarity.
[0097] Optionally, the association determination module 12 is further used to: if any of the vector similarities is greater than the similarity threshold, determine the picture to be detected corresponding to the vector similarity as a risk picture, and obtain the picture identifier of the risk picture in the document to be detected;
[0098] Determine the picture description paragraph according to the picture identifier, and determine the picture description paragraph as the associated document corresponding to the risk picture.
[0099] Further, the association determination module 12 is further configured to: match the picture identifier with the document to be detected, and determine an identification statement according to the identification matching result;
[0100] Match the identification statement with a preset identification vocabulary, and determine the paragraph corresponding to the identification statement that successfully matches the preset identification vocabulary as the picture description paragraph.
[0101] The semantic recognition module 13 is configured to input the associated document into a pre-trained large model for semantic recognition to obtain document semantics, and determine a detection score corresponding to the risk picture according to the document semantics.
[0102] Optionally, the semantic recognition module 13 is further configured to: obtain a sample document, and input the sample document into the large model for word segmentation to obtain sample word segmentation;
[0103] Perform vector conversion on the sample word segmentation to obtain word segmentation vectors, and perform attention mechanism processing on the word segmentation vectors to obtain attention features;
[0104] Perform a non-linear transformation on the attention features, and perform normalization processing on the attention features after the non-linear transformation to obtain normalized features;
[0105] Perform residual processing on the normalized features to obtain residual features, and perform feature decoding on the residual features to obtain decoded features;
[0106] Determine a model loss according to the decoded features, and update the parameters of the large model according to the model loss until the large model converges to obtain the pre-trained large model.
[0107] Further, the semantic recognition module 13 is further configured to: perform entity recognition on semantic vocabulary in the document semantics, and determine an object name and object description vocabulary in the document semantics according to the entity recognition result;
[0108] Determine a description semantic database according to the entity type of the object description vocabulary, and calculate the similarity between the object description vocabulary and semantic vocabulary in the description semantic database to obtain a vocabulary similarity;
[0109] Determine the detection score according to the object name and the vocabulary similarity.
[0110] Even further, the semantic recognition module 13 is further configured to: perform name matching between the object name and the target name corresponding to the target graphic logo;
[0111] If the name matching fails, determine the preset score as the detection score;
[0112] If the name matching is successful, calculate the sum of the corresponding maximum lexical similarities between different object description words to obtain the total similarity;
[0113] Calculate the average similarity of the object description words according to the total similarity, and match the average similarity with the score query table to obtain the detection score.
[0114] The result generation module 14 is used to generate the graphic logo detection result of the document to be detected according to the detection score.
[0115] Optionally, the result generation module 14 is further used to: if the detection score is greater than or equal to the score threshold, determine that the graphic logo detection of the document to be detected is qualified, and there is no infringement risk for the document to be detected;
[0116] If the detection score is greater than or equal to the preset score and less than the score threshold, determine that the graphic logo detection of the document to be detected is unqualified, and the document to be detected has a first-level infringement risk;
[0117] If the detection score is less than or equal to the score threshold, determine that the graphic logo detection of the document to be detected is unqualified, and the document to be detected has a second-level infringement risk.
[0118] Please refer to Figure 3 , this embodiment proposes an official logo retrieval scheme that combines semantic and image features. Through the public opinion monitoring system and web crawlers, relevant data of the official website logos cited in public network articles are crawled. The mainstream ViT model is used to extract the picture feature vectors, and the pictures in the document to be reviewed are converted into feature vectors and stored in the Faiss vector database. Use the vector database to perform similar picture retrieval. After obtaining the relevant document objects, enter the semantic discrimination link. In this inspection link, the semantic understanding ability of the large language model and the preset inspection rules are used to perform context retrieval on the article content. After the large model scores the retrieval results, the large model judgment result is given according to the preset threshold, assisting enterprise customers to identify whether the document is a normal citation of the logo or an infringement citation of the logo.
[0119] In this embodiment, the accuracy of image recognition is improved. By utilizing the image encoder function in the field of computer vision and together with the vector database, the retrieval function of similar images is realized. The problem of misjudgment is solved. By using the semantic understanding ability of the large language model and the manually preset checkpoints, it is determined whether the current document containing the enterprise logo picture belongs to legitimate citation. The retrieval efficiency is improved. By taking into account the features of both semantics and images, it can help enterprise users efficiently identify the documents that infringe on the citation of the enterprise logo and assist the enterprise in better protecting its own intellectual property rights and corporate image.
[0120] In this embodiment, by extracting pictures from the document to be detected, the pictures to be detected in the document to be detected can be automatically obtained. By performing feature vector conversion on the target graphic logo and the pictures to be detected, the pictures can be effectively converted into vectors, obtaining the target graphic vector and the vector to be detected. By calculating the vector similarity between the target graphic vector and the vector to be detected, based on the vector similarity, the associated documents in the document to be detected can be effectively determined. By inputting the associated documents into a pre-trained large model for semantic recognition, the document semantics of the associated documents can be effectively extracted. Based on the document semantics, the detection score of the risk pictures can be automatically determined. Based on the detection score, the graphic logo detection result of the document to be detected can be automatically generated, without the need to use manual methods for graphic logo detection, thus improving the efficiency of graphic logo detection.
[0121] Embodiment III
[0122] Figure 4 is a structural block diagram of a terminal device 2 provided in the third embodiment of the present application. As Figure 4 shown, the terminal device 2 of this embodiment includes: a processor 20, a memory 21, and a computer program 22 stored in the memory 21 and executable on the processor 20, such as a program for the graphic logo detection method based on a large model. When the processor 20 executes the computer program 22, the steps in each of the above embodiments of the graphic logo detection method based on a large model are implemented.
[0123] Exemplarily, the computer program 22 can be divided into one or more modules. The one or more modules are stored in the memory 21 and executed by the processor 20 to complete the present application. The one or more modules can be a series of computer program instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program 22 in the terminal device 2. The terminal device may include, but is not limited to, a processor 20 and a memory 21.
[0124] The so-called processor 20 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0125] The memory 21 may be an internal storage unit of the terminal device 2, such as the hard disk or memory of the terminal device 2. The memory 21 may also be an external storage device of the terminal device 2, such as a plug-in hard disk equipped on the terminal device 2, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 21 may also include both the internal storage unit of the terminal device 2 and the external storage device. The memory 21 is used to store the computer program and other programs and data required by the terminal device. The memory 21 may also be used to temporarily store the data that has been output or will be output.
[0126] In addition, in each embodiment of the present application, each functional module may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0127] When an integrated module is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium can be non-volatile or volatile. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of this application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium can include: any entity or device that can carry computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable storage medium does not include electrical carrier signals and telecommunication signals.
[0128] The above-mentioned embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of this application, and should all be included in the protection scope of this application.
Claims
1. A method for detecting graphic signs based on a large model, characterized in that: The method comprises: Obtaining a document to be detected, and performing image extraction on the document to be detected to obtain an image to be detected; Obtaining a target image and text mark, and performing feature vector conversion on the target image and text mark and the image to be detected, to obtain a target image and text vector and a vector to be detected; Calculating the vector similarity between the target image-text vector and the vector to be detected, and determining the risk image and the associated document of the risk image in the document to be detected according to the vector similarity; Inputting the associated document into a pre-trained large model for semantic recognition to obtain document semantics, and determining a detection score corresponding to the risk image according to the document semantics; A detection result of the graphic mark of the document to be detected is generated according to the detection score.
2. The large model-based image and text sign detection method according to claim 1, characterized in that: Before inputting the associated documents into the pre-trained large model for semantic recognition, the method further includes: Obtain a sample document, and input the sample document into the large model for word segmentation to obtain sample word segmentations; Performing vector conversion on the sample word segmentation to obtain a word segmentation vector, and performing attention mechanism processing on the word segmentation vector to obtain an attention feature; Performing a nonlinear transformation on the attention feature, and normalizing the attention feature after the nonlinear transformation to obtain a normalized feature; Performing residual processing on the normalized features to obtain residual features, and performing feature decoding on the residual features to obtain decoded features; The model loss is determined according to the decoding features, and the parameters of the large model are updated according to the model loss until the large model converges to obtain the pre-trained large model.
3. The large model-based image and text sign detection method according to claim 1, characterized in that: Determining a detection score corresponding to the risk image according to the document semantics includes: Performing entity recognition on the semantic words in the document semantics, and determining the object name and object description words in the document semantics according to the entity recognition result; Determine a description semantic database according to the entity type of the object description vocabulary, and calculate the similarity between the object description vocabulary and the semantic vocabulary in the description semantic database to obtain vocabulary similarity; The detection score is determined based on the object name and the vocabulary similarity.
4. The large model-based image and text sign detection method according to claim 3, characterized in that: Determining the detection score according to the object name and the vocabulary similarity includes: Perform name matching between the object name and the target name corresponding to the target graphic mark; If the name matching fails, the preset score is determined as the detection score; If the name matches successfully, then the sum of the corresponding maximum vocabulary similarities between different object description words is calculated to obtain the total similarity; The average similarity of the object description words is calculated according to the total similarity, and the average similarity is matched with the score query table to obtain the detection score.
5. The large model-based image and text sign detection method according to claim 4, characterized in that: Generating a detection result of the graphic mark of the document to be detected according to the detection score includes: If the detection score is greater than or equal to the score threshold, it is determined that the image and text mark detection of the document to be detected is qualified, and the document to be detected does not have an infringement risk; If the detection score is greater than or equal to the preset score and less than the score threshold, it is determined that the image and text mark detection of the document to be detected is unqualified, and the document to be detected has a first-level infringement risk; If the detection score is less than or equal to the score threshold, it is determined that the graphic mark detection of the document to be detected is unqualified, and the document to be detected has a second-level infringement risk.
6. The method for detecting graphic signs based on a large model as claimed in claim 1, characterized in that: Determining the risk image and the associated document of the risk image in the document to be detected according to the vector similarity includes: If any of the vector similarities is greater than a similarity threshold, the image to be detected corresponding to the vector similarity is determined as a risk image, and an image identifier of the risk image in the document to be detected is obtained; A picture description paragraph is determined according to the picture identifier, and the picture description paragraph is determined as the associated document corresponding to the risk picture.
7. The large model-based image and text sign detection method according to claim 6, characterized in that: Determining a picture description paragraph according to the picture identifier includes: Performing identification matching between the image identification and the document to be detected, and determining an identification statement according to the identification matching result; The identification sentence is matched with a preset identification vocabulary, and the paragraph corresponding to the identification sentence that successfully matches the preset identification vocabulary is determined as the picture description paragraph.
8. A large model-based graphic sign detection system, characterized in that: The system comprises: The image extraction module is used to obtain the document to be detected and extract the image of the document to be detected to obtain the image to be detected; A vector conversion module is used to obtain a target graphic mark and perform feature vector conversion on the target graphic mark and the image to be detected to obtain a target graphic vector and a vector to be detected; An association determination module, used to calculate the vector similarity between the target image-text vector and the vector to be detected, and determine the risk image and the associated document of the risk image in the document to be detected according to the vector similarity; A semantic recognition module, used for inputting the associated document into a pre-trained large model for semantic recognition, obtaining document semantics, and determining a detection score corresponding to the risk image according to the document semantics; The result generation module is used to generate the image and text mark detection result of the document to be detected according to the detection score.
9. The large model-based graphic sign detection system according to claim 8, characterized in that: The semantic recognition module is also used for: Obtain a sample document, and input the sample document into the large model for word segmentation to obtain sample word segmentations; Performing vector conversion on the sample word segmentation to obtain a word segmentation vector, and performing attention mechanism processing on the word segmentation vector to obtain an attention feature; Performing a nonlinear transformation on the attention feature, and normalizing the attention feature after the nonlinear transformation to obtain a normalized feature; Performing residual processing on the normalized features to obtain residual features, and performing feature decoding on the residual features to obtain decoded features; The model loss is determined according to the decoding features, and the parameters of the large model are updated according to the model loss until the large model converges to obtain the pre-trained large model.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Gift box design image-text detection and recognition system based on big data analysis
CN120747584A