Intelligent detection method and device for drug-related codewords based on multimodal semantic recognition

Through multimodal semantic recognition methods, text and image information is obtained, and the drug-related risks are judged using a codeword knowledge base and a large language model, which solves the problem of high missed detection rate in existing technologies and achieves more accurate drug-related detection.

CN120296681BActive Publication Date: 2025-09-19国家毒品实验室陕西分中心(陕西省公安厅毒品技术中心)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510774480.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-19
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

Existing technologies in drug-related detection rely on a single modality and have limitations in semantic understanding. They are unable to cope with the dynamic variations of code words and contextual semantic disguises, and lack cross-modal correlation analysis capabilities, resulting in a high missed detection rate.

Method used

By obtaining the target text segment and image set, determining the entity keywords and image information, using the drug-related codeword knowledge base to match the codeword root information, calculating the similarity, and inputting the large language model for semantic association, the drug-related risk confidence level is judged to achieve multimodal semantic recognition.

Benefits of technology

It improves the accuracy of drug-related identification, reduces missed detections, and can identify drug-related code words in text and images, improving detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296681B_ABST
    Figure CN120296681B_ABST
Patent Text Reader

Abstract

The present application provides an intelligent detection method and device for drug-related codewords based on multimodal semantic recognition, which belongs to the field of semantic recognition technology. In order to solve the problem of high missed detection rate of current drug-related codeword detection methods, the method can identify the first entity keyword in the target text segment and the first entity information in the target image set, and can embed the screened second codeword root information in a preset text frame, and perform semantic association through the target large language model to obtain the drug-related risk confidence. Finally, the drug-related risk confidence is used to judge whether the target text segment contains drug-related codewords. Drug-related codewords in the target text segment and the target image set can be identified, which can reduce the missed detection of drug-related identification to a certain extent, thereby improving the accuracy of drug-related identification to a certain extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method and device for intelligent detection of drug-related codewords based on multimodal semantic recognition. Background Art

[0002] Drugs refer to opium, heroin, methamphetamine (ice), morphine, marijuana, cocaine, and other addictive narcotics and psychotropic substances regulated by the state. Drugs can seriously damage physical health, leading to various diseases and organ failure. Long-term drug use can also cause mental disorders such as hallucinations and delusions, posing serious risks to personal health and social stability. In recent years, with the rapid development of mobile internet and encrypted communication technologies, drug-related criminal activities have become increasingly covert and intelligent. Criminals communicate through instant messaging applications and social media platforms, creating significant challenges for drug enforcement.

[0003] In related technologies, string matching can be performed based on a fixed keyword library (such as "ice" and "heroin"). When the corresponding sensitive string is matched, it is determined that the communication content involves drugs, so that corresponding drug enforcement work can be carried out.

[0004] However, the above methods have the following problems: (1) Single-modal dependence and semantic understanding limitations. Traditional detection systems are mostly based on fixed keyword libraries (such as "ice" and "heroin") for string matching, which cannot cope with the dynamic variation of code words (such as new aliases such as "rock sugar" and "jelly") and contextual semantic disguise (such as "tea quality is very good" refers to the purity of drugs in a specific context); (2) Multimodal fragmentation. Existing solutions usually process text and images independently and lack cross-modal association analysis capabilities. For example, drug information hidden by steganography in an image (such as the text watermark "50 grams of meat") and "goods have been shipped" in the chat text cannot form a chain of evidence, resulting in missed detection. Existing methods' isolated analysis of text is prone to missed detection, resulting in a high missed detection rate for drug-related code words. Summary of the Invention

[0005] In view of the above problems, the embodiments of the present application provide a method, device, electronic device and readable storage medium for intelligent detection of drug-related code words based on multimodal semantic recognition, so as to overcome the above problems or at least partially solve the above problems.

[0006] In a first aspect, embodiments of the present application provide a method for intelligently detecting drug-related codewords based on multimodal semantic recognition, the method comprising:

[0007] Obtaining a target text segment and a target image set suspected of being drug-related; wherein the target image set and the target text segment are contained in the same evidence data packet, and the target image set is associated with a generation timestamp of the target text segment in the evidence data packet;

[0008] Determining a first entity keyword in the target text segment and determining first entity information contained in the target image set;

[0009] Determining first codeword root information that matches the first entity keyword from a drug-related codeword knowledge base;

[0010] Determining a first similarity between each piece of the first codeword root information and each piece of the first entity information;

[0011] Determining, from the first coded word root information, second coded word root information corresponding to a first similarity greater than or equal to a first threshold;

[0012] The second codeword root information is embedded in a preset text framework to obtain a target text to be processed, and the target text to be processed is input into a target large language model to obtain a drug-related risk confidence score output by the target large language model; wherein the preset text framework is used to enable the target large language model to perform semantic association based on the first codeword root information; the target large language model is a large language model trained based on the preset text framework, and the target large language model is connected to the drug-related codeword knowledge base;

[0013] Based on the drug-related risk confidence level, it is determined whether the target text segment and the target image set contain drug-related code words.

[0014] Optionally, determining the first entity information contained in the target image set includes:

[0015] Determine a first resolution corresponding to each first image in the target image set;

[0016] Inputting a first image having a first resolution smaller than a preset resolution into a trained super-resolution generative adversarial network model to obtain a second image output by the super-resolution generative adversarial network model;

[0017] Inputting the second image and the first image having the first resolution greater than or equal to the preset resolution into a trained lightweight YOLOv7-tiny model to obtain a first recognition image output by the YOLOv7-tiny model; wherein the recognition scope of the YOLOv7-tiny model includes drugs and drug-related paraphernalia;

[0018] Determine first entity information included in the first recognition image.

[0019] Optionally, determining the first similarity between each piece of first codeword root information and each piece of first entity information includes:

[0020] Generating first semantic feature vectors corresponding to each of the first codeword root information and second semantic feature vectors corresponding to each of the first entity information;

[0021] A first cosine similarity between the first semantic feature vector and the second semantic feature vector is calculated to obtain a first similarity between the first codeword root information and the first entity information.

[0022] Optionally, determining the first entity keyword in the target text segment includes:

[0023] extracting a first Chinese text segment from the target text segment;

[0024] performing word segmentation processing on the first Chinese text segment to obtain a plurality of first Chinese word roots;

[0025] Determining first pronunciation features corresponding to each of the first Chinese word roots; the first pronunciation features include Mandarin pronunciation features and dialect pronunciation features;

[0026] Determining a second Chinese word root corresponding to the first pronunciation feature;

[0027] Semantic recognition is performed on the second Chinese word root to obtain a first entity keyword in the target text segment.

[0028] Optionally, determining first codeword root information matching the first entity keyword from a drug-related codeword knowledge base includes:

[0029] Determine the first inverted index corresponding to the drug-related codeword knowledge base;

[0030] Based on the first inverted index, first codeword root information matching the first entity keyword is determined from the drug-related codeword knowledge base.

[0031] Optionally, the determining, based on the drug-related risk confidence level, whether the target text segment and the target image set contain drug-related codewords includes:

[0032] Determining a first semantic similarity between the first entity keyword and the second codeword root information;

[0033] Determining first image similarities between each first image included in the target image set and the sample drug-related image;

[0034] Determining a comprehensive drug-related probability of the target text segment based on the drug-related risk confidence level, the first semantic similarity, and the first image similarity;

[0035] Based on the comprehensive drug-related probability, it is determined whether the target text segment and the target image set contain drug-related code words.

[0036] Optionally, determining the comprehensive drug-related probability of the target text segment based on the drug-related risk confidence, the first semantic similarity, and the first image similarity includes:

[0037] respectively determining a first weighting coefficient of the drug-related risk confidence, a second weighting coefficient of the first semantic similarity, and a third weighting coefficient of the first image similarity;

[0038] Based on the first weighting coefficient, the second weighting coefficient and the third weighting coefficient, the drug-related risk confidence, the first semantic similarity and the first image similarity are weightedly summed to obtain a comprehensive drug-related probability of the target text segment.

[0039] Optionally, the determining whether the target text segment and the target image set contain drug-related codewords based on the comprehensive drug-related probability includes:

[0040] If the comprehensive drug-related probability is less than a second threshold, determining that the target text segment does not contain drug-related codewords;

[0041] If the comprehensive drug-related probability is greater than or equal to the second threshold and less than a third threshold, determining that the target text segment is suspected of containing drug-related codewords;

[0042] When the comprehensive drug-related probability is greater than or equal to the third threshold, it is determined that the target text segment contains drug-related codewords.

[0043] In a second aspect, an embodiment of the present application provides an intelligent detection device for drug-related codewords based on multimodal semantic recognition, the device comprising:

[0044] an acquisition module, configured to acquire a target text segment and a target image set suspected of being drug-related; wherein the target image set and the target text segment are contained in the same evidence data packet, and the target image set is associated with a generation timestamp of the target text segment in the evidence data packet;

[0045] A first determining module is configured to determine a first entity keyword in the target text segment and determine first entity information contained in the target image set;

[0046] A second determining module is configured to determine first codeword root information matching the first entity keyword from a drug-related codeword knowledge base;

[0047] a third determining module, configured to determine a first similarity between each of the first codeword root information and each of the first entity information;

[0048] a fourth determining module, configured to determine, from the first codeword root information, second codeword root information corresponding to a first similarity greater than or equal to a first threshold;

[0049] An input / output module, configured to embed the second codeword root information into a preset text framework to obtain a target text to be processed, and input the target text to be processed into a target large language model to obtain a drug-related risk confidence score output by the target large language model; wherein the preset text framework is configured to enable the target large language model to perform semantic association based on the first codeword root information; the target large language model is a large language model trained based on the preset text framework, and the target large language model is connected to the drug-related codeword knowledge base;

[0050] The fifth determination module is used to determine whether the target text segment and the target image set contain drug-related code words based on the drug-related risk confidence level.

[0051] Optionally, the first determining module includes:

[0052] A first determination submodule, configured to determine first resolutions corresponding to respective first images in the target image set;

[0053] A first input-output submodule is configured to input a first image having a first resolution smaller than a preset resolution into a trained super-resolution generative adversarial network model, and obtain a second image output by the super-resolution generative adversarial network model;

[0054] A second input-output submodule is configured to input the second image and the first image having the first resolution greater than or equal to the preset resolution into a trained lightweight YOLOv7-tiny model to obtain a first recognition image output by the YOLOv7-tiny model; wherein the recognition scope of the YOLOv7-tiny model includes drugs and drug-related paraphernalia;

[0055] The second determining submodule is configured to determine first entity information contained in the first recognition image.

[0056] Optionally, the third determining module includes:

[0057] a generating submodule, configured to generate first semantic feature vectors corresponding to respective pieces of the first codeword root information and second semantic feature vectors corresponding to respective pieces of the first entity information;

[0058] The calculation submodule is configured to calculate a first cosine similarity between the first semantic feature vector and the second semantic feature vector to obtain a first similarity between the first codeword root information and the first entity information.

[0059] Optionally, the first determining module includes:

[0060] an extraction submodule, configured to extract a first Chinese text segment from the target text segment;

[0061] A word segmentation submodule, configured to perform word segmentation processing on the first Chinese text segment to obtain a plurality of first Chinese word roots;

[0062] A third determining submodule is configured to determine first pronunciation features corresponding to each of the first Chinese word roots; the first pronunciation features include Mandarin pronunciation features and dialect pronunciation features;

[0063] a fourth determining submodule, configured to determine a second Chinese word root corresponding to the first pronunciation feature;

[0064] The semantic recognition submodule is used to perform semantic recognition on the second Chinese word root to obtain the first entity keyword in the target text segment.

[0065] Optionally, the second determining module includes:

[0066] A fifth determination submodule is used to determine a first inverted index corresponding to the drug-related codeword knowledge base;

[0067] The sixth determining submodule is configured to determine, based on the first inverted index, first codeword root information that matches the first entity keyword from the drug-related codeword knowledge base.

[0068] Optionally, the fifth determining module includes:

[0069] a seventh determination submodule, configured to determine a first semantic similarity between the first entity keyword and the second codeword root information;

[0070] an eighth determining submodule, configured to determine first image similarities between each first image included in the target image set and the sample drug-related image;

[0071] a ninth determination submodule, configured to determine a comprehensive drug-related probability of the target text segment based on the drug-related risk confidence level, the first semantic similarity, and the first image similarity;

[0072] The tenth determination submodule is configured to determine whether the target text segment and the target image set contain drug-related code words based on the comprehensive drug-related probability.

[0073] Optionally, the ninth determining submodule includes:

[0074] a first determining unit, configured to respectively determine a first weighting coefficient of the drug-related risk confidence, a second weighting coefficient of the first semantic similarity, and a third weighting coefficient of the first image similarity;

[0075] A summing unit is used to perform weighted summation on the drug risk confidence, the first semantic similarity and the first image similarity based on the first weighting coefficient, the second weighting coefficient and the third weighting coefficient to obtain a comprehensive drug-related probability of the target text segment.

[0076] Optionally, the tenth determining submodule includes:

[0077] a second determining unit, configured to determine that the target text segment does not contain drug-related codewords when the comprehensive drug-related probability is less than a second threshold;

[0078] a third determining unit, configured to determine that the target text segment is suspected of containing drug-related codewords when the comprehensive drug-related probability is greater than or equal to the second threshold and less than a third threshold;

[0079] The fourth determining unit is configured to determine that the target text segment contains drug-related codewords when the comprehensive drug-related probability is greater than or equal to the third threshold.

[0080] In a third aspect, an embodiment of the present application provides an electronic device comprising a memory, a processor, and a computer program stored on the memory, wherein the processor executes the computer program to implement the intelligent detection method for drug-related code words based on multimodal semantic recognition as described in any one of the above items.

[0081] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the intelligent detection method for drug-related code words based on multimodal semantic recognition as described in any one of the above items is implemented.

[0082] The specific beneficial effects are:

[0083] The embodiment of the present application obtains a target text segment and a target image set suspected of being drug-related; wherein the target image set and the target text segment are contained in the same evidence data packet, and the target image set is associated with the generation timestamp of the target text segment in the evidence data packet, determines the first entity keyword in the target text segment, determines the first entity information contained in the target image set, determines the first cipher root information matching the first entity keyword from the drug-related cipher knowledge base, determines the first similarity between each first cipher root information and each first entity information, determines the second cipher root information corresponding to the first similarity greater than or equal to the first threshold from the first cipher root information, embeds the second cipher root information into a preset text framework, obtains the target text to be processed, and inputs the target text to be processed into the target large language model to obtain the drug-related risk confidence level output by the target large language model; wherein the preset text This framework is used to enable the target large language model to perform semantic association based on the first codeword root information; the target large language model is a large language model trained based on a preset text framework, and the target large language model is connected to a drug-related codeword knowledge base. Based on the drug-related risk confidence, it is determined whether the target text segment and the target image set contain drug-related codewords. The first entity keyword in the target text segment and the first entity information in the target image set can be identified, and the screened second codeword root information can be embedded in the preset text framework, and semantic association is performed through the target large language model to obtain the drug-related risk confidence. Finally, the drug-related risk confidence is used to determine whether the target text segment contains drug-related codewords. Drug-related identification can be performed on the drug-related codewords in the target text segment and the target image set, which can reduce the number of missed drug-related identifications to a certain extent, thereby improving the accuracy of drug-related identification to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0085] Figure 1 This is a flow chart of an intelligent drug-related codeword detection method based on multimodal semantic recognition provided by an embodiment of the present application;

[0086] Figure 2 This is a flow chart of another method for intelligent detection of drug-related codewords based on multimodal semantic recognition provided by an embodiment of the present application;

[0087] Figure 3 This is a logic block diagram of an intelligent drug-related codeword detection device based on multimodal semantic recognition provided by an embodiment of the present application;

[0088] Figure 4 This is a schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0089] The exemplary embodiments of the present application will be described in more detail below in conjunction with the accompanying drawings in the embodiments of the present application. Although the accompanying drawings show exemplary embodiments of the present application, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. Instead, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.

[0090] Reference Figure 1 , Figure 1 A flowchart of a method for intelligently detecting drug-related codewords based on multimodal semantic recognition is provided in an embodiment of the present application. The method may include:

[0091] Step 101: Obtain a target text segment and a target image set suspected of being drug-related; wherein the target image set and the target text segment are contained in the same evidence data packet, and the target image set is associated with a generation timestamp of the target text segment in the evidence data packet.

[0092] In an embodiment of the present application, the operator of the communication software can package the communication content of the user, thereby obtaining an electronic evidence data packet. By decompressing and analyzing the electronic evidence data packet, the target text segment suspected of being drug-related can be obtained. The target text segment is the communication text content between users. The evidence data packet is a chat data packet obtained after the chat software operator packages the chat data. The evidence data packet may also include some chat images, which may constitute a target image set. Thus, a target image set consisting of chat images may be obtained from the evidence data packet. Among them, the target image set and the target text segment may be derived from the same evidence data packet, and the generation timestamps of the target image set and the target text segment in the evidence data packet may be correlated, that is, the time difference between the generation timestamps of each image in the target image set and the generation timestamp of the target text segment may be less than a preset duration.

[0093] Step 102: Determine the first entity keyword in the target text segment and determine the first entity information contained in the target image set.

[0094] In an embodiment of the present application, the target text segment may contain entity keywords, action keywords, and modifiers. Entities may represent things that exist objectively and can be distinguished from each other, such as fans, water dispensers, cups, tea leaves, etc. Action keywords may represent the activity state of an entity, such as rotation, heating, etc. A language model (such as a BERT model) may be used to perform semantic analysis on the target text segment, thereby obtaining the first entity keyword in the target text segment. The individual images contained in the target image set may also be analyzed, thereby extracting the text information and item information contained in each image. Among them, the item information may be directly used as part of the first entity information. For the text information, a trained BERT model may be used for processing, thereby separating the entity information from the text information. By integrating the entity information and item information corresponding to the text information of each image, the first entity information contained in the target image set may be obtained.

[0095] Step 103: Determine first codeword root information that matches the first entity keyword from a drug-related codeword knowledge base.

[0096] In an embodiment of the present application, the process of constructing a drug-related jargon knowledge base may include crawling relevant data from public news sources such as court judgments, drug enforcement news, and jargon-related discussion threads on social media platforms. Drug-related jargon from platforms such as the dark web and Telegram is also captured in real time. Drug control experts then organize the jargon and its meaning to form structured data, categorizing the jargon types into categories such as common drug and ingredient jargon, institutional raw material and ingredient jargon, drug funding jargon, and drug processing and manufacturing processes. The collected images are annotated, each labeled with its category, and semantic features (including text and entity information within the image) are extracted and stored in the drug-related jargon knowledge base. Within the drug-related jargon knowledge base, the graph database Neo4j can be used to construct entity-jargon-image relationships and knowledge association rules, enabling the drug-related jargon knowledge base to support semantic queries. Neo4j is a high-performance graph database specifically designed for storing, managing, and querying graph-structured data (composed of nodes, relationships, and attributes). Unlike traditional relational databases, Neo4j uses a graph model to store data, making it more suitable for processing complex relational networks. The method for building a drug-related codeword knowledge base using Neo4j is to organize and manage drug-related information through the node and relationship model of a graph database. The specific steps include: first, defining three core node types: entities (drugs such as methamphetamine and heroin), codewords (disguised names such as rock sugar and jelly), and images (such as photos of drugs and images of drug-making tools). Attributes are added to each node, such as the chemical composition of the entity and the confidence score of the codeword. Next, three key relationships are established: "HAS_ALIAS" (the association between entities and codewords), "DETECTED_IN" (the association between images and codewords / entities), and "SIMILAR_TO" (the similarity between images). The Cypher query language is used for data import and relationship establishment. Finally, graph algorithms are used to implement intelligent reasoning. Based on this, the drug-related codeword knowledge base can be searched for first codeword root information that matches the first entity keyword. "Match" can refer to identical or similar characters. When building a drug-related codeword knowledge base, models can be used to extract codewords, build multimodal relationships, store entity-codeword-image relationships, store images of related codewords, and construct knowledge association rules. By integrating multimodal features and contextual semantics, these knowledge association rules significantly enhance the ability to identify drug-related codewords when using the drug-related codeword knowledge base.

[0097] For example, if the first entity keyword is "rock sugar" or "water sugar", the first codeword root information that can be matched is "rock sugar".

[0098] Step 104: Determine a first similarity between each piece of the first codeword root information and each piece of the first entity information.

[0099] In an embodiment of the present application, a first similarity between each first codeword root information and each first entity information can be calculated. The first codeword root information and the first entity information can be respectively input into an embedding module to obtain a first codeword root feature vector and a first entity feature vector. Then, a cosine similarity between the first codeword root feature vector and the first entity feature vector is calculated, and the cosine similarity is used as the first similarity between the first codeword root information and the first entity information.

[0100] Step 105: Determine, from the first codeword root information, second codeword root information corresponding to a first similarity greater than or equal to a first threshold.

[0101] In an embodiment of the present application, the first codeword root information having a first similarity greater than or equal to a first threshold may be determined as the second codeword root information.

[0102] Step 106: embed the second codeword root information into a preset text framework to obtain a target text to be processed, and input the target text to be processed into a target large language model to obtain the drug-related risk confidence level output by the target large language model; wherein the preset text framework is used to enable the target large language model to perform semantic association based on the first codeword root information; the target large language model is a large language model trained based on the preset text framework, and the target large language model is connected to the drug-related codeword knowledge base.

[0103] In an embodiment of the present application, the preset text frame is a data structure and can be used to train a large language model, so that the large language model can make semantic associations based on the specific text content contained in the preset text frame. The second cipher root information can be embedded in the preset text frame, so that the target text to be processed can be obtained, and then the target text to be processed can be input into the target large language model, so that the target large language model can make semantic associations based on the first cipher root information contained in the target text to be processed, thereby outputting the drug-related risk confidence. Among them, the target large language model can be connected to the drug-related cipher knowledge base, so that the target large language model can use the cipher data mapping contained in the drug-related cipher knowledge base to obtain the drug-related risk confidence.

[0104] For example, a target text to be processed is as follows:

[0105] "You are a narcotics analyst. Please generate a report based on the following evidence:

[0106] Text clue: The code word "rock sugar" (a substitute for methamphetamine, confidence level 92%) was detected, and the context is "tastier and purer"; Image clue: Methamphetamine crystals (confidence level 93%) and sealed bag packaging (confidence level 88%) were found: Association rule: It conforms to the pattern of "drug + quantifier + transaction". Please output: Risk level confidence (0 - 1) and an interpretable logical chain. The context part can be omitted.

[0107] The following is an example of the output result of a target large language model:

[0108] You are a drug analysis expert. Please generate a report based on the following evidence:

[0109] Text clue: The code word "rock sugar" (a substitute for methamphetamine, confidence level 92%) was detected, and the context is "tastier and purer"; Image clue: Methamphetamine crystals (confidence level 93%) and sealed bag packaging (confidence level 88%) were found; Association rule: It conforms to the pattern of "drug + quantifier + transaction". Please output: Risk level confidence (0 - 1) and an interpretable logical chain.

[0110] Model analysis process:

[0111] Okay, now I will analyze the user's needs and generate a risk level report on methamphetamine transactions. First, the user provided some text and image clues, as well as association rules.

[0112] Analysis of text clues: The code word "rock sugar" was detected with a confidence level of 92%. The context mentions "tastier and purer". This indicates that it may be discussing the characteristics of a certain drug, but it needs to be further confirmed whether it is methamphetamine.

[0113] Analysis of image clues: Methamphetamine crystals and sealed bag packaging were found, with confidence levels of 93% and 88% respectively. These evidences support the existence of a drug indicated by the code word "rock sugar", and there is a packaging method, which conforms to the characteristics of commercial transactions.

[0114] Application of association rules: According to the pattern of "drug + quantifier + transaction", "rock sugar" combined with a quantifier (such as "piece" or similar) usually appears in drug transactions. Detecting the crystals and sealed packaging further confirms this.

[0115] Considering these factors, the image evidences of methamphetamine crystals and packaging are very strong, and the text clues also point to the characteristics of drugs. Therefore, the risk level confidence is relatively high.

[0116] Model output:

[0117] Risk level confidence: 0.95

[0118] Logical chain:

[0119] 01. The text clue detected "rock sugar" with a confidence level of 92%, and mentioned that it has a purer taste, which indicates a specific description of a drug-related item.

[0120] 02. Image clues indicate the discovery of methamphetamine crystals, with a confidence level of 93%; sealed bag packaging was also found, with a confidence level of 88%. These evidences further confirm that the items referred to by the code word "rock sugar" are drug-related.

[0121] 03. Association rule matching: It matches the pattern of "drugs + quantifier (e.g., "granules") + "transaction." Combining these clues, we can infer that the item might be methamphetamine used in drug trafficking.

[0122] 04. Comprehensive judgment: Based on high-confidence text and image evidence, as well as association rules that conform to drug trafficking patterns, the risk of drug trafficking is determined.

[0123] Therefore, the document is likely to involve drug trafficking, with a risk level confidence of 0.95.

[0124] The output shown in the example above is the result obtained after inputting the target text to be processed into the target large language model. It indicates a drug risk confidence level of 0.95 and also provides a logical analysis chain.

[0125] Step 107: Based on the drug-related risk confidence level, determine whether the target text segment and the target image set contain drug-related code words.

[0126] In the embodiments of the present application, the drug risk confidence level indicates the probability that a target text segment contains drug-related codewords. Based on the drug risk confidence level, whether the target text segment and target image set contain drug-related codewords can be determined. For example, if the drug risk confidence level is greater than or equal to a first threshold (e.g., 0.9), the target text segment can be determined to contain drug-related codewords. If the drug risk confidence level is less than the first threshold, the target text segment can be determined to not contain drug-related codewords.

[0127] In an embodiment of the present application, a target text segment and a target image set suspected of being drug-related are obtained; wherein the target image set and the target text segment are contained in the same evidence data packet, and the target image set is associated with the generation timestamp of the target text segment in the evidence data packet, a first entity keyword in the target text segment is determined, the first entity information contained in the target image set is determined, the first cipher root information matching the first entity keyword is determined from a drug-related cipher knowledge base, the first similarity between each first cipher root information and each first entity information is determined, and the second cipher root information corresponding to the first similarity greater than or equal to the first threshold is determined from the first cipher root information, the second cipher root information is embedded in a preset text framework to obtain a target text to be processed, and the target text to be processed is input into a target large language model to obtain a drug-related risk confidence level output by the target large language model; wherein, the predetermined A text framework is provided for enabling a target large language model to perform semantic association based on the first codeword root information; the target large language model is a large language model trained based on a preset text framework, and the target large language model is connected to a drug-related codeword knowledge base. Based on the drug-related risk confidence, it is determined whether the target text segment and the target image set contain drug-related codewords, and the first entity keyword in the target text segment and the first entity information in the target image set can be identified. The screened second codeword root information can be embedded in the preset text framework, and semantic association is performed through the target large language model to obtain the drug-related risk confidence. Finally, the drug-related risk confidence is used to determine whether the target text segment contains drug-related codewords, and drug-related codewords in the target text segment and the target image set can be identified. To a certain extent, the situation of missed drug-related identification can be reduced, thereby improving the accuracy of drug-related identification to a certain extent.

[0128] Reference Figure 2 , Figure 2 A flowchart of another method for intelligently detecting drug-related codewords based on multimodal semantic recognition provided in an embodiment of the present application may include:

[0129] Step 201, obtaining a target text segment and a target image set suspected of being drug-related; wherein the target image set and the target text segment are contained in the same evidence data packet, and the target image set is associated with a generation timestamp of the target text segment in the evidence data packet.

[0130] In the embodiment of the present application, the implementation content of this step can refer to the embodiment content of step 101 and will not be repeated here.

[0131] Step 202: Determine the first entity keyword in the target text segment and determine the first entity information contained in the target image set.

[0132] In the embodiment of the present application, the implementation content of this step can refer to the embodiment content of step 102 and will not be repeated here.

[0133] Optionally, step 202 may include the following sub-steps:

[0134] Sub-step 2021: extracting a first Chinese text segment from the target text segment.

[0135] In an embodiment of the present application, the target text segment may contain multiple language types and some emoticon symbols. Therefore, a regular expression or string replacement method can be used to extract the first Chinese text segment from the target text segment.

[0136] For example, a regular expression to extract Chinese text in Python is as follows:

[0137] import re

[0138] text = "Hello, World! 123 ABC Chinese Test." #Mixed text

[0139] chinese_text = re.findall(r'[\u4e00-\u9fff]+', text) #Use regular expressions to match Chinese characters

[0140] result = ''.join(chinese_text) #Convert the list to a string

[0141] print(result) # Output: World Chinese Test

[0142] The "\u4e00-\u9fff" in the above code is the Unicode code range of Chinese characters.

[0143] Sub-step 2022: performing word segmentation processing on the first Chinese text segment to obtain a plurality of first Chinese word roots.

[0144] In an embodiment of the present application, a word segmentation tool can be used to perform word segmentation on the first Chinese text segment, so as to obtain multiple first Chinese word roots. Jieba is a Chinese word segmentation tool that supports three word segmentation methods: accurate mode, full mode, and search engine mode, and can meet the word segmentation requirements in different scenarios. Jieba supports custom dictionaries, and users can add exclusive vocabulary according to professional fields or specific needs to ensure the accuracy of word segmentation. Based on the prefix dictionary and dynamic programming algorithm, it realizes fast and accurate word segmentation processing and is widely used in fields such as information retrieval, text mining, and natural language processing. By processing the text with a word segmentation tool, it is possible to accurately segment drug-related words according to the custom dictionary, ensuring that professional terms and specific expressions are accurately extracted.

[0145] Sub-step 2023, determine the first pronunciation features respectively corresponding to each of the first Chinese word roots; the first pronunciation features include Mandarin pronunciation features and dialect pronunciation features.

[0146] In an embodiment of the present application, since each Chinese character has its corresponding pinyin pronunciation, therefore, the first pronunciation features respectively corresponding to each of the first Chinese word roots can be determined. Among them, the first pronunciation features can include Mandarin pronunciation features and dialect pronunciation features. Among them, the Mandarin pronunciation features can correspond to the pinyin phonetic symbols of the first Chinese word root in the Chinese dictionary; for the dialect pronunciation features, the pinyin phonetic symbols of the dialect can be obtained according to its pronunciation and the composition of the pinyin phonetic symbols of Mandarin. By using an embedding module to vectorize the pinyin phonetic symbols, the Mandarin pronunciation features and dialect pronunciation features can be obtained. [[ID=VIII]]

[0147] For example, for the Chinese character "鞋" (shoe), the Mandarin pinyin phonetic symbol is "xié", and in some regions, its dialect pinyin phonetic symbol is "hái". Thus, the first pronunciation features of this Chinese character can include Mandarin pronunciation features and dialect pronunciation features. Among them, the Mandarin pronunciation feature is the feature vector obtained by inputting the character "xié" into the embedding module, and the dialect pronunciation feature is the feature vector obtained by inputting the character "hái" into the embedding module.

[0148] Sub-step 2024, determine the second Chinese word roots corresponding to the first pronunciation features.

[0149] In an embodiment of the present application, after obtaining the first pronunciation features, the second Chinese word roots corresponding to the first pronunciation features can be further determined. There can be multiple second Chinese word roots corresponding to the first pronunciation features. Specifically, it is all Chinese characters or Chinese words pronounced in the Mandarin way under the pinyin phonetic symbols corresponding to this pronunciation feature.

[0150] Continuing with the above example, the second Chinese root can include all the Chinese characters corresponding to the Mandarin pinyin phonetic symbol "xié", such as "邪", "携", etc., and can also include all the Chinese characters corresponding to the Mandarin pinyin phonetic symbol "hái", such as "孩", "还", "骸", etc.

[0151] Sub-step 2025: Perform semantic recognition on the second Chinese root to obtain the first entity keyword in the target text segment.

[0152] In an embodiment of the present application, semantic recognition can be performed on the second Chinese root, so that the first entity keyword in the target text segment can be obtained. Here, a trained BERT model can be used to filter out non-entity roots, so that the first entity keyword can be obtained.

[0153] For example, a code example for root filtering is as follows:

[0154] from transformers import pipeline, AutoModelForTokenClassification,AutoTokenizer

[0155] import jieba.posseg as pseg # for word tagging

[0156] # Load the pre-trained Chinese NER model (take ckiplab as an example)

[0157] model_name = "ckiplab / bert-base-chinese-ner" # Replace with the actual available model name​​​​​​​​​​​​​​​​​text = "Li Ming bought a Huawei Mate60 phone and a book called "Introduction to Python Programming."

[0164] # Perform named entity recognition

[0165] results = ner_pipeline(text)

[0166] # Define item-related entity types (adjusted according to model tags)

[0167] # For example: items may belong to types such as PRODUCT, OBJ, etc., or be judged by part of speech (noun)

[0168] target_entities = ['PRODUCT', 'OBJ'] # Replace with the labels actually supported by the model

[0169] # Filter target entity type

[0170] filtered_results = [ent for ent in results if ent['entity'] intarget_entities]

[0171] # Further filter nouns by part-of-speech tagging

[0172] # Perform part-of-speech tagging on unmatched entities and extract nouns as items

[0173] words = pseg.cut(text)

[0174] for word, pos in words:

[0175] if pos.startswith('n'): # noun (n is a person's name, nr; a place name, ns; an organization name, nt; other nouns, nz)

[0176] # Exclude matched entities

[0177] if all(not (ent['start']<= word.start() <ent['end'] and ent['end']<= word.end()) for ent in filtered_results):

[0178] filtered_results.append({'entity': 'OBJ', 'word': word,'start': word.start(), 'end': word.end()})

[0179] # Remove duplicates and sort

[0180] filtered_results = sorted(set(filtered_results), key=lambda x: x['start'])

[0181] # Output the results

[0182] print("Identified item type roots:")

[0183] for entity in filtered_results:

[0184] print(f"Root: {entity['word']}, Type: {entity['entity']}")

[0185] The output result of the above code is as follows:

[0186] Identified item type roots:

[0187] Root: Huawei Mate 60 mobile phone, Type: PRODUCT

[0188] Root: "Python Programming Introduction", Type: PRODUCT

[0189] Root: book, Type: OBJ

[0190] For example, "frozen" belongs to a non-entity root, so this root will be filtered out by the BERT model; "rock sugar" belongs to an entity root, so the BERT model will output this root as the first entity keyword. In the above example, the action root "buy" and the quantifier roots "one" and "one book" as well as other modifiers will be filtered out.

[0191] In an embodiment of the present application, a first Chinese text segment is extracted from a target text segment, the first Chinese text segment is segmented to obtain multiple first Chinese word roots, and the first pronunciation features corresponding to each first Chinese word root are determined; the first pronunciation features include Mandarin pronunciation features and dialect pronunciation features, and the second Chinese word root corresponding to the first pronunciation feature is determined. The second Chinese word root is semantically recognized to obtain the first entity keyword in the target text segment. The first Chinese word root can be data expanded through Mandarin pronunciation and dialect pronunciation to obtain the second Chinese word root, and the first entity keyword in the target text segment can be further determined through the second Chinese word root, so that the first entity keyword can include some of the latest drug-related code words, which can avoid the missed detection of the target text segment to a certain extent and improve the accuracy and reliability of the first entity keyword.

[0192] Sub-step 2026: Determine the first resolution corresponding to each first image in the target image set.

[0193] In the embodiment of the present application, resolution is an inherent property of an image. The first resolution corresponding to each first image can be determined based on the image property of each first image in the target image set.

[0194] Sub-step 2027: Input the first image whose first resolution is smaller than the preset resolution into a trained super-resolution generative adversarial network model to obtain a second image output by the super-resolution generative adversarial network model.

[0195] In an embodiment of the present application, a first image with a first resolution less than a preset resolution can be input into a trained super-resolution generative adversarial network (SRGAN) model, thereby obtaining a second image output by the super-resolution GAN model. The super-resolution GAN model is an image super-resolution reconstruction network model based on a generative adversarial network (GAN). It uses deep learning techniques to restore low-resolution (LR) images to high-resolution (HR) images while enhancing texture detail and visual realism. After processing by the SRGAN model, a higher-resolution image can be obtained. In other words, the second image has a higher resolution (typically four times higher) than the first image input into the SRGAN model.

[0196] Sub-step 2028: Input the second image and the first image whose first resolution is greater than or equal to the preset resolution into a trained lightweight YOLOv7-tiny model to obtain a first recognition image output by the YOLOv7-tiny model; wherein the recognition range of the YOLOv7-tiny model includes drugs and drug-related smoking tools.

[0197] In an embodiment of the present application, the YOLOv7-tiny model is an ultra-lightweight target detection model in the YOLO series, designed for resource-constrained scenarios, and can achieve high-precision target detection while maintaining a high processing speed. In an embodiment of the present application, the YOLOv7-tiny model can be trained to identify drugs and drug-related smoking tools. The second image and the first image with a first resolution greater than or equal to the preset resolution can be input into the trained lightweight YOLOv7-tiny model, so that the first recognition image output by the YOLOv7-tiny model can be obtained.

[0198] Sub-step 2029: determining first entity information contained in the first recognition image.

[0199] In an embodiment of the present application, a suspected drug or a suspected drug-related smoking paraphernalia may be identified in the first recognition image by a colored frame. Therefore, the first entity information contained in the first recognition image may be determined by the first recognition image.

[0200] In an embodiment of the present application, by determining the first resolutions corresponding to each first image in the target image set, the first image whose first resolution is less than the preset resolution is input into a trained super-resolution generative adversarial network model to obtain a second image output by the super-resolution generative adversarial network model, and the second image and the first image whose first resolution is greater than or equal to the preset resolution are input into a trained lightweight YOLOv7-tiny model to obtain a first recognition image output by the YOLOv7-tiny model; wherein the recognition range of the YOLOv7-tiny model includes drugs and drug-related smoking tools, and the first entity information contained in the first recognition image is determined, the image can be resolution enhanced by the super-resolution generative adversarial network model, and then the enhanced second image and the qualified first image can be subjected to target recognition by the YOLOv7-tiny model to obtain the first entity information, which can improve the accuracy and reliability of the first entity information to a certain extent.

[0201] Step 203: Determine first codeword root information that matches the first entity keyword from a drug-related codeword knowledge base.

[0202] In the embodiment of the present application, the implementation content of this step can refer to the embodiment content of step 103 and will not be repeated here.

[0203] Optionally, step 203 may include the following sub-steps:

[0204] Sub-step 2031: determining the first inverted index corresponding to the drug-related codeword knowledge base.

[0205] In an embodiment of the present application, an inverted index is an index structure used for full-text search, which is commonly found in search engines, databases and other systems. It achieves fast retrieval by mapping words or phrases in a document to a list of documents containing these words. The inverted index uses words, phrases or specific symbols (such as punctuation) as the smallest unit to record the occurrence of each term in a document set. The inverted index greatly improves the efficiency of full-text search by reverse mapping documents through terms. The data structure established in the drug-related codeword knowledge base is "codeword-drug name-drug picture", and these data have fixed storage addresses. Therefore, a first inverted index is formed in the drug-related codeword knowledge base. Therefore, the first inverted index corresponding to the drug-related codeword knowledge base can be determined from the drug-related codeword knowledge base. The structural framework of the first inverted index can be "query keyword-data storage address".

[0206] Sub-step 2032: Based on the first inverted index, determine the first codeword root information that matches the first entity keyword from the drug-related codeword knowledge base.

[0207] In an embodiment of the present application, the first inverted index can be used to determine the first codeword root information that matches the first entity keyword from the drug-related codeword knowledge base, thereby determining all drug-related codewords that may be contained in the target text segment.

[0208] In an embodiment of the present application, by determining the first inverted index corresponding to the drug-related codeword knowledge base, based on the first inverted index, the first codeword root information matching the first entity keyword is determined from the drug-related codeword knowledge base, and the first codeword root information can be queried in the drug-related codeword knowledge base through the inverted index method, which can improve the query speed and query accuracy of the first codeword root information to a certain extent.

[0209] Step 204: Determine a first similarity between each piece of the first codeword root information and each piece of the first entity information.

[0210] In the embodiment of the present application, the implementation content of this step can refer to the embodiment content of step 104 and will not be repeated here.

[0211] Optionally, step 204 may include the following sub-steps:

[0212] Sub-step 2041 : generating first semantic feature vectors corresponding to each of the first codeword root information and second semantic feature vectors corresponding to each of the first entity information.

[0213] In an embodiment of the present application, the first coded word root information and the first entity information can be respectively input into a trained Word2Vec model, so that the first semantic feature vector corresponding to each first coded word root information and the second semantic feature vector corresponding to each first entity information can be generated by the Word2Vec model.

[0214] Sub-step 2042 : calculating a first cosine similarity between the first semantic feature vector and the second semantic feature vector to obtain a first similarity between the first codeword root information and the first entity information.

[0215] In an embodiment of the present application, a first cosine similarity can be calculated between a first semantic feature vector and a second semantic feature vector. This first cosine similarity can serve as a first similarity between the first codeword root information corresponding to the first semantic feature vector and the first entity information corresponding to the second semantic feature vector. By traversing the first semantic feature vector and the second semantic feature vector, multiple first similarities can be obtained.

[0216] In an embodiment of the present application, by respectively generating first semantic feature vectors corresponding to each first coded word root information and second semantic feature vectors corresponding to each first entity information, calculating the first cosine similarity between the first semantic feature vector and the second semantic feature vector, the first similarity between the first coded word root information and the first entity information is obtained. The first cosine similarity between the first semantic feature vector corresponding to each first coded word root information and the second semantic feature vector corresponding to each first entity information can be used as the first similarity between the first coded word root information and the first entity information, which can avoid semantic omissions and improve the accuracy of the first similarity to a certain extent.

[0217] Step 205: Determine, from the first codeword root information, second codeword root information corresponding to a first similarity greater than or equal to a first threshold.

[0218] In the embodiment of the present application, the implementation content of this step can refer to the embodiment content of step 105 and will not be repeated here.

[0219] Step 206: embed the second codeword root information into a preset text framework to obtain a target text to be processed, and input the target text to be processed into a target large language model to obtain the drug-related risk confidence level output by the target large language model; wherein the preset text framework is used to enable the target large language model to perform semantic association based on the first codeword root information; the target large language model is a large language model trained based on the preset text framework, and the target large language model is connected to the drug-related codeword knowledge base.

[0220] In the embodiment of the present application, the implementation content of this step can refer to the embodiment content of step 106 and will not be repeated here.

[0221] Step 207: Determine a first semantic similarity between the first entity keyword and the second codeword root information.

[0222] In an embodiment of the present application, a first semantic similarity between the first entity keyword and the second coded word root information can be determined by first extracting semantic feature vectors of the first entity keyword and the first coded word root information, respectively, and then calculating a cosine similarity between the semantic feature vector of the first entity keyword and the semantic feature vector of the first coded word root information, and using the cosine similarity as the first semantic similarity between the first entity keyword and the first coded word root information.

[0223] Step 208: Determine the first image similarity between each first image included in the target image set and the sample drug-related image.

[0224] In an embodiment of the present application, the first image similarity between each first image included in the target image set and the sample drug-related image can be calculated, and the calculation method is shown in the following formula 1:

[0225] (Formula 1);

[0226] In the above formula 1, and are the average brightness of the first image and the sample drug-related image respectively. and are the standard deviations of the first image and the sample drug-related images, respectively. is the covariance between the first image and the sample drug-related images. and is a constant used to avoid the denominator being zero. Usually, and ,in and , is the dynamic range (e.g., for an 8-bit image, L=255). The sample drug-related images may be sourced from a drug-related codeword knowledge base.

[0227] Step 209 : Determine the comprehensive drug-related probability of the target text segment based on the drug-related risk confidence, the first semantic similarity, and the first image similarity.

[0228] In an embodiment of the present application, a comprehensive drug-related probability of the target text segment can be calculated based on the drug-related risk confidence level, the first semantic similarity, and the first image similarity. For example, an arithmetic average method can be used to calculate the average of the drug-related risk confidence level, the first semantic similarity, and the first image similarity, and this average value can be used as the comprehensive drug-related probability of the target text segment.

[0229] Optionally, step 209 may include the following sub-steps:

[0230] Sub-step 2091 , respectively determining a first weighting coefficient of the drug-related risk confidence, a second weighting coefficient of the first semantic similarity, and a third weighting coefficient of the first image similarity.

[0231] In an embodiment of the present application, a first weighting coefficient for the drug risk confidence, a second weighting coefficient for the first semantic similarity, and a third weighting coefficient for the first image similarity may be determined separately, wherein the sum of the first weighting coefficient, the second weighting coefficient, and the third weighting coefficient may be 1.

[0232] Sub-step 2092, based on the first weighting coefficient, the second weighting coefficient and the third weighting coefficient, weighted summation is performed on the drug risk confidence, the first semantic similarity and the first image similarity to obtain the comprehensive drug-related probability of the target text segment.

[0233] In an embodiment of the present application, the drug risk confidence, the first semantic similarity and the first image similarity can be weighted and summed according to the first weighting coefficient, the second weighting coefficient and the third weighting coefficient, so that the comprehensive drug-related probability of the target text segment can be calculated.

[0234] In an embodiment of the present application, by respectively determining the first weighted coefficient of the drug risk confidence, the second weighted coefficient of the first semantic similarity and the third weighted coefficient of the first image similarity, the drug risk confidence, the first semantic similarity and the first image similarity are weightedly summed based on the first weighted coefficient, the second weighted coefficient and the third weighted coefficient to obtain the comprehensive drug-related probability of the target text segment. The drug risk confidence, the first semantic similarity and the first image similarity can be weightedly summed to obtain the comprehensive drug-related probability of the target text segment, which can improve the degree of fit between the comprehensive drug-related probability and the actual scene to a certain extent, thereby improving the accuracy and reliability of the comprehensive drug-related probability.

[0235] Step 210: Based on the comprehensive drug-related probability, determine whether the target text segment and the target image set contain drug-related code words.

[0236] In an embodiment of the present application, whether a target text segment or target image set contains drug-related codewords can be determined based on a comprehensive drug-related probability. If the comprehensive drug-related probability is greater than or equal to a preset threshold, the target text segment or target image set can be considered to contain drug-related codewords. If the comprehensive drug-related probability is less than the preset threshold, the target text segment or target image set can be considered not to contain drug-related codewords.

[0237] Optionally, step 210 may include the following sub-steps:

[0238] Sub-step 2101, when the comprehensive drug-related probability is less than a second threshold, determining that the target text segment does not contain drug-related code words.

[0239] In an embodiment of the present application, if the comprehensive drug-related probability is less than or equal to the second threshold, it can be determined that the target text segment does not contain drug-related code words.

[0240] Sub-step 2102, when the comprehensive drug-related probability is greater than or equal to the second threshold and less than the third threshold, determines that the target text segment is suspected of containing drug-related code words.

[0241] In the embodiment of the present application, if the comprehensive drug-related probability is greater than or equal to the second threshold and less than the third threshold, it can be considered that the target text segment is suspected of containing drug-related code words. In this case, further drug-related judgment can be made manually on the target text segment.

[0242] Sub-step 2103, when the comprehensive drug-related probability is greater than or equal to the third threshold, determining that the target text segment contains drug-related codewords.

[0243] In an embodiment of the present application, if the comprehensive drug-related probability is greater than or equal to a third threshold, it can be determined that the target text segment contains drug-related codewords.

[0244] In an embodiment of the present application, by determining that the target text segment does not contain drug-related codewords when the comprehensive drug-related probability is less than the second threshold, determining that the target text segment is suspected of containing drug-related codewords when the comprehensive drug-related probability is greater than or equal to the second threshold and less than the third threshold, and determining that the target text segment contains drug-related codewords when the comprehensive drug-related probability is greater than or equal to the third threshold, the target text segment can be graded and judged based on the comprehensive drug-related probability, thereby improving the accuracy of the drug-related judgment of the target text segment to a certain extent.

[0245] In an embodiment of the present application, by determining the first semantic similarity between the first entity keyword and the first codeword root information, the first image similarity between each first image contained in the target image set and the sample drug-related image is determined, and based on the drug-related risk confidence, the first semantic similarity and the first image similarity, the comprehensive drug-related probability of the target text segment is determined. Based on the comprehensive drug-related probability, it is determined whether the target text segment and the target image set contain drug-related codewords. The drug-related risk confidence, the first semantic similarity and the first image similarity can be jointly calculated to obtain the comprehensive drug-related probability, and the comprehensive drug-related probability can be used to determine whether the target text segment and the target image set contain drug-related codewords. Since the comprehensive drug-related probability includes evaluation indicators of multiple dimensions, when the comprehensive drug-related probability is used to make drug-related judgments, the accuracy and reliability of the judgment can be improved.

[0246] Reference Figure 3 , Figure 3 This is a logical block diagram of an intelligent drug-related codeword detection device based on multimodal semantic recognition provided in an embodiment of the present application. The device 300 may include:

[0247] An acquisition module 301 is configured to acquire a target text segment and a target image set suspected of being drug-related; wherein the target image set and the target text segment are contained in the same evidence data packet, and the target image set is associated with a generation timestamp of the target text segment in the evidence data packet;

[0248] A first determining module 302 is configured to determine a first entity keyword in the target text segment and determine first entity information contained in the target image set;

[0249] The second determining module 303 is configured to determine first codeword root information matching the first entity keyword from a drug-related codeword knowledge base;

[0250] A third determining module 304 is configured to determine a first similarity between each of the first codeword root information and each of the first entity information;

[0251] A fourth determining module 305 is configured to determine, from the first codeword root information, second codeword root information corresponding to a first similarity greater than or equal to a first threshold;

[0252] Input / output module 306 is configured to embed the second codeword root information into a preset text framework to obtain a target text to be processed, and input the target text to be processed into a target large language model to obtain a drug-related risk confidence score output by the target large language model; wherein the preset text framework is used to enable the target large language model to perform semantic association based on the first codeword root information; the target large language model is a large language model trained based on the preset text framework, and the target large language model is connected to the drug-related codeword knowledge base;

[0253] The fifth determination module 307 is configured to determine whether the target text segment and the target image set contain drug-related code words based on the drug-related risk confidence level.

[0254] Optionally, the first determining module 302 includes:

[0255] A first determination submodule, configured to determine first resolutions corresponding to respective first images in the target image set;

[0256] A first input-output submodule is configured to input a first image having a first resolution smaller than a preset resolution into a trained super-resolution generative adversarial network model, and obtain a second image output by the super-resolution generative adversarial network model;

[0257] A second input-output submodule is configured to input the second image and the first image having the first resolution greater than or equal to the preset resolution into a trained lightweight YOLOv7-tiny model to obtain a first recognition image output by the YOLOv7-tiny model; wherein the recognition scope of the YOLOv7-tiny model includes drugs and drug-related paraphernalia;

[0258] The second determining submodule is configured to determine first entity information contained in the first recognition image.

[0259] Optionally, the third determining module 304 includes:

[0260] a generating submodule, configured to generate first semantic feature vectors corresponding to respective pieces of the first codeword root information and second semantic feature vectors corresponding to respective pieces of the first entity information;

[0261] The calculation submodule is configured to calculate a first cosine similarity between the first semantic feature vector and the second semantic feature vector to obtain a first similarity between the first codeword root information and the first entity information.

[0262] Optionally, the first determining module 302 includes:

[0263] an extraction submodule, configured to extract a first Chinese text segment from the target text segment;

[0264] A word segmentation submodule, configured to perform word segmentation processing on the first Chinese text segment to obtain a plurality of first Chinese word roots;

[0265] A third determining submodule is configured to determine first pronunciation features corresponding to each of the first Chinese word roots; the first pronunciation features include Mandarin pronunciation features and dialect pronunciation features;

[0266] a fourth determining submodule, configured to determine a second Chinese word root corresponding to the first pronunciation feature;

[0267] The semantic recognition submodule is used to perform semantic recognition on the second Chinese word root to obtain the first entity keyword in the target text segment.

[0268] Optionally, the second determining module 303 includes:

[0269] A fifth determination submodule is used to determine a first inverted index corresponding to the drug-related codeword knowledge base;

[0270] The sixth determining submodule is configured to determine, based on the first inverted index, first codeword root information that matches the first entity keyword from the drug-related codeword knowledge base.

[0271] Optionally, the fifth determining module 307 includes:

[0272] a seventh determination submodule, configured to determine a first semantic similarity between the first entity keyword and the second codeword root information;

[0273] an eighth determining submodule, configured to determine first image similarities between each first image included in the target image set and the sample drug-related image;

[0274] a ninth determination submodule, configured to determine a comprehensive drug-related probability of the target text segment based on the drug-related risk confidence level, the first semantic similarity, and the first image similarity;

[0275] The tenth determination submodule is configured to determine whether the target text segment and the target image set contain drug-related code words based on the comprehensive drug-related probability.

[0276] Optionally, the ninth determining submodule includes:

[0277] a first determining unit, configured to respectively determine a first weighting coefficient of the drug-related risk confidence, a second weighting coefficient of the first semantic similarity, and a third weighting coefficient of the first image similarity;

[0278] A summing unit is used to perform weighted summation on the drug risk confidence, the first semantic similarity and the first image similarity based on the first weighting coefficient, the second weighting coefficient and the third weighting coefficient to obtain a comprehensive drug-related probability of the target text segment.

[0279] Optionally, the tenth determining submodule includes:

[0280] a second determining unit, configured to determine that the target text segment does not contain drug-related codewords when the comprehensive drug-related probability is less than a second threshold;

[0281] a third determining unit, configured to determine that the target text segment is suspected of containing drug-related codewords when the comprehensive drug-related probability is greater than or equal to the second threshold and less than a third threshold;

[0282] The fourth determining unit is configured to determine that the target text segment contains drug-related codewords when the comprehensive drug-related probability is greater than or equal to the third threshold.

[0283] The intelligent drug-related codeword detection device based on multimodal semantic recognition in the embodiments of the present application can be an electronic device or a component of an electronic device, such as an integrated circuit or chip. The electronic device can be a terminal or other device other than a terminal. For example, the electronic device can be a GPU box, a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a mobile internet device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook or a personal digital assistant (PDA), etc. It can also be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine or a self-service machine, etc., and the embodiments of the present application are not specifically limited.

[0284] The intelligent drug-related codeword detection device based on multimodal semantic recognition in the embodiments of the present application can be a device having an operating system. The operating system can be an Android operating system, a Linux operating system, a Windows operating system, or other possible operating systems, which are not specifically limited in the embodiments of the present application.

[0285] The intelligent drug-related codeword detection device based on multimodal semantic recognition provided by the embodiment of the present application can achieve Figures 1 to 2 To avoid repetition, the various processes implemented in the method embodiment are not described here.

[0286] The present application provides an electronic device. Figure 4 The electronic device 40 includes: a processor 401, a memory 402, and a computer program 4021 stored in the memory 402 and executable on the processor 401. When the processor 401 executes the program, the intelligent drug-related codeword detection method based on multimodal semantic recognition of the aforementioned embodiment is implemented.

[0287] An embodiment of the present application also provides a computer-readable storage medium having a computer program / instruction stored thereon. When the computer program / instruction is executed by a processor, the steps of the intelligent detection method for drug-related code words based on multimodal semantic recognition as disclosed in the embodiment of the present application are implemented.

[0288] An embodiment of the present application also provides a computer program product. When the computer program product is run on an electronic device, the processor is enabled to implement the steps of the intelligent detection method for drug-related codewords based on multimodal semantic recognition as disclosed in the embodiment of the present application.

[0289] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0290] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, apparatuses, electronic devices, and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0291] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1The function specified in one or more boxes.

[0292] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0293] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0294] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "includes," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements that are inherent to such process, method, article, or terminal device. In the absence of further restrictions, an element defined by the phrase "comprises a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.

[0295] The above is a detailed introduction to the intelligent detection method and device for drug-related codewords based on multimodal semantic recognition provided by this application. Specific examples are used in this article to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method of this application and its core idea; at the same time, for general technical personnel in this field, based on the ideas of this application, there will be changes in the specific implementation methods and application areas. In summary, the content of this specification should not be understood as a limitation on this application.

Claims

1. An intelligent drug-related codeword detection method based on multimodal semantic recognition, characterized in that: The method comprises: Obtaining a target text segment and a target image set suspected of being drug-related; wherein the target image set and the target text segment are contained in the same evidence data packet, and the target image set is associated with a generation timestamp of the target text segment in the evidence data packet; Determining a first entity keyword in the target text segment and determining first entity information contained in the target image set; Determining first codeword root information that matches the first entity keyword from a drug-related codeword knowledge base; Determining a first similarity between each piece of the first codeword root information and each piece of the first entity information; Determining, from the first coded word root information, second coded word root information corresponding to a first similarity greater than or equal to a first threshold; The second codeword root information is embedded in a preset text framework to obtain a target text to be processed, and the target text to be processed is input into a target large language model to obtain a drug-related risk confidence score output by the target large language model; wherein the preset text framework is used to enable the target large language model to perform semantic association based on the first codeword root information; the target large language model is a large language model trained based on the preset text framework, and the target large language model is connected to the drug-related codeword knowledge base; Based on the drug-related risk confidence level, determining whether the target text segment and the target image set contain drug-related codewords; The step of determining whether the target text segment and the target image set contain drug-related codewords based on the drug-related risk confidence level includes: Determining a first semantic similarity between the first entity keyword and the second codeword root information; Determining first image similarities between each first image included in the target image set and the sample drug-related image; Determining a comprehensive drug-related probability of the target text segment based on the drug-related risk confidence level, the first semantic similarity, and the first image similarity; Based on the comprehensive drug-related probability, determining whether the target text segment and the target image set contain drug-related code words; The determining of the comprehensive drug-related probability of the target text segment based on the drug-related risk confidence, the first semantic similarity, and the first image similarity includes: respectively determining a first weighting coefficient of the drug-related risk confidence, a second weighting coefficient of the first semantic similarity, and a third weighting coefficient of the first image similarity; Based on the first weighting coefficient, the second weighting coefficient and the third weighting coefficient, the drug-related risk confidence, the first semantic similarity and the first image similarity are weightedly summed to obtain a comprehensive drug-related probability of the target text segment.

2. The method according to claim 1, characterized in that The determining the first entity information included in the target image set includes: Determine a first resolution corresponding to each first image in the target image set; Inputting a first image having a first resolution smaller than a preset resolution into a trained super-resolution generative adversarial network model to obtain a second image output by the super-resolution generative adversarial network model; Inputting the second image and the first image having the first resolution greater than or equal to the preset resolution into a trained lightweight YOLOv7-tiny model to obtain a first recognition image output by the YOLOv7-tiny model; wherein the recognition scope of the YOLOv7-tiny model includes drugs and drug-related paraphernalia; Determine first entity information included in the first recognition image.

3. The method according to claim 1, characterized in that The determining of the first similarity between each of the first codeword root information and each of the first entity information includes: Generating first semantic feature vectors corresponding to each of the first codeword root information and second semantic feature vectors corresponding to each of the first entity information; A first cosine similarity between the first semantic feature vector and the second semantic feature vector is calculated to obtain a first similarity between the first codeword root information and the first entity information.

4. The method according to claim 1, wherein Determining the first entity keyword in the target text segment includes: extracting a first Chinese text segment from the target text segment; performing word segmentation processing on the first Chinese text segment to obtain a plurality of first Chinese word roots; Determining first pronunciation features corresponding to each of the first Chinese word roots; the first pronunciation features include Mandarin pronunciation features and dialect pronunciation features; Determining a second Chinese word root corresponding to the first pronunciation feature; Semantic recognition is performed on the second Chinese word root to obtain a first entity keyword in the target text segment.

5. The method according to claim 1, wherein The step of determining the first codeword root information that matches the first entity keyword from the drug-related codeword knowledge base includes: Determine the first inverted index corresponding to the drug-related codeword knowledge base; Based on the first inverted index, first codeword root information matching the first entity keyword is determined from the drug-related codeword knowledge base.

6. The method according to claim 1, characterized in that The determining, based on the comprehensive drug-related probability, whether the target text segment and the target image set contain drug-related codewords includes: If the comprehensive drug-related probability is less than a second threshold, determining that the target text segment does not contain drug-related codewords; If the comprehensive drug-related probability is greater than or equal to the second threshold and less than a third threshold, determining that the target text segment is suspected of containing drug-related codewords; When the comprehensive drug-related probability is greater than or equal to the third threshold, it is determined that the target text segment contains drug-related codewords.

7. An intelligent drug-related codeword detection device based on multimodal semantic recognition, characterized in that: The device comprises: an acquisition module, configured to acquire a target text segment and a target image set suspected of being drug-related; wherein the target image set and the target text segment are contained in the same evidence data packet, and the target image set is associated with a generation timestamp of the target text segment in the evidence data packet; A first determining module is configured to determine a first entity keyword in the target text segment and determine first entity information contained in the target image set; A second determining module is configured to determine first codeword root information matching the first entity keyword from a drug-related codeword knowledge base; a third determining module, configured to determine a first similarity between each of the first codeword root information and each of the first entity information; a fourth determining module, configured to determine, from the first codeword root information, second codeword root information corresponding to a first similarity greater than or equal to a first threshold; An input / output module, configured to embed the second codeword root information into a preset text framework to obtain a target text to be processed, and input the target text to be processed into a target large language model to obtain a drug-related risk confidence score output by the target large language model; wherein the preset text framework is configured to enable the target large language model to perform semantic association based on the first codeword root information; the target large language model is a large language model trained based on the preset text framework, and the target large language model is connected to the drug-related codeword knowledge base; a fifth determination module, configured to determine whether the target text segment and the target image set contain drug-related codewords based on the drug-related risk confidence level; Wherein, the fifth determining module includes: a seventh determination submodule, configured to determine a first semantic similarity between the first entity keyword and the second codeword root information; an eighth determining submodule, configured to determine first image similarities between each first image included in the target image set and the sample drug-related image; a ninth determination submodule, configured to determine a comprehensive drug-related probability of the target text segment based on the drug-related risk confidence level, the first semantic similarity, and the first image similarity; a tenth determination submodule, configured to determine whether the target text segment and the target image set contain drug-related codewords based on the comprehensive drug-related probability; The ninth determining submodule includes: a first determining unit, configured to respectively determine a first weighting coefficient of the drug-related risk confidence, a second weighting coefficient of the first semantic similarity, and a third weighting coefficient of the first image similarity; A summing unit is used to perform weighted summation on the drug risk confidence, the first semantic similarity and the first image similarity based on the first weighting coefficient, the second weighting coefficient and the third weighting coefficient to obtain a comprehensive drug-related probability of the target text segment.