Sales forbidding and limiting behavior identification method and device, storage medium and program product
By extracting multi-dimensional features from product images and text, and integrating visual and textual information, this technology identifies prohibited and restricted sales behaviors in online transactions, solving the problem of difficulty in identifying concealed prohibited and restricted sales behaviors in existing technologies, and achieving efficient and accurate identification results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA ACADEMY OF ELECTRONICS AND INFORMATION TECHNOLOGY OF CHINA ELECTRONICS TECHNOLOGY GROUP CORPORATION
- Filing Date
- 2026-01-23
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies are insufficient to effectively identify prohibited or restricted goods in online transactions, especially those presented through text distortion, image steganography, and semantic spoofing, which increases the difficulty of regulation.
By extracting semantic, frequency domain, and reconstruction features from product images, and combining them with semantic, speech, and glyph features of product text, the system integrates visual and textual features to identify the predicted probability of prohibited or restricted sales behaviors.
It significantly improves the accuracy and efficiency of identifying and processing obscure prohibited and restricted sales behaviors, and can cope with diverse image steganography and text deformation methods to achieve efficient and accurate identification of prohibited and restricted sales behaviors.
Smart Images

Figure CN122089435A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, device, storage medium, and program product for identifying prohibited or restricted sales behaviors. Background Technology
[0002] With the rapid development of online transactions, the variety of goods and services on platforms is becoming increasingly diverse, and the transaction volume continues to expand. To maintain market order, protect public safety and consumer rights, effectively managing goods and services prohibited or restricted from sale in online transactions has become a key task for the compliant operation of these platforms.
[0003] Prohibited or restricted sales activities refer to the actions of businesses in online transactions that directly or indirectly publish, promote, or facilitate transactions involving goods or services that are prohibited or restricted from sale by laws, regulations, or platform rules. Such activities not only include explicitly listing prohibited items but also often employ covert methods, such as text distortion, image steganography, semantic spoofing, and cross-regional traffic diversion, to circumvent the supervision of trading platforms.
[0004] Therefore, there is an urgent need for a scheme to identify prohibited and restricted sales behaviors, which can efficiently and accurately identify various prohibited and restricted sales behaviors in transaction scenarios. Summary of the Invention
[0005] This application provides a method, device, storage medium, and program product for identifying prohibited and restricted sales behaviors, which can efficiently and accurately identify various prohibited and restricted sales behaviors in transaction scenarios.
[0006] In a first aspect, embodiments of this application provide a method for identifying prohibited or restricted sales behaviors, the method comprising: Obtain the product images and text descriptions displayed on the trading platform for the target product; Extract the image semantic features and frequency domain features corresponding to the product image, as well as the restoration features obtained after the product image is processed for content completion; the image semantic features are used to describe the product category and explicit prohibited or restricted items in the product image, the frequency domain features are used to describe digital tampering traces in the product image, and the restoration features are used to describe information that is physically occluded in the product image; Based on the image semantic features, the frequency domain features, and the restored features, the visual features of the target product are determined; Extract the text semantic features, speech features, and glyph features corresponding to the product text description; Based on the text semantic features, the speech features, and the glyph features, the text features of the target product are determined; The visual features and the text features are fused to obtain the fused features of the target product; Based on the fusion features, the predicted probability that the act of providing the target product on the trading platform constitutes a prohibited or restricted sales act is calculated.
[0007] Secondly, embodiments of this application provide a device for identifying prohibited or restricted sales behaviors, the device comprising: The acquisition module is used to acquire the product images and text descriptions of the target product displayed on the trading platform.
[0008] The processing module is used to extract the image semantic features and frequency domain features corresponding to the product image, as well as the restored features obtained after content completion processing of the product image; the image semantic features are used to describe the product category and explicit prohibited or restricted sales content in the product image, the frequency domain features are used to describe digital tampering traces in the product image, and the restored features are used to describe information physically obscured in the product image; based on the image semantic features, the frequency domain features, and the restored features, the visual features of the target product are determined; the text semantic features, speech features, and glyph features corresponding to the product text description are extracted; based on the text semantic features, the speech features, and the glyph features, the text features of the target product are determined; the visual features and the text features are fused to obtain the fused features of the target product; based on the fused features, the predicted probability that the act of providing the target product on the trading platform belongs to prohibited or restricted sales behavior is calculated.
[0009] Thirdly, embodiments of this application provide an electronic device, including: a memory, a processor, and a communication interface; wherein, the memory stores a computer program, and when the computer program is executed by the processor, the processor can at least implement the prohibited and restricted sales behavior identification method as described in the first aspect.
[0010] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor of an electronic device, enables the processor to at least implement the prohibited or restricted sales behavior identification method as described in the first aspect.
[0011] Fifthly, embodiments of this application provide a computer program product, including: a computer program or instructions, which, when executed by a processor of an electronic device, enable the processor to at least implement the prohibited or restricted sales behavior identification method as described in the first aspect.
[0012] The solution provided in this application extracts multi-dimensional visual features from product images that can characterize global semantics, frequency domain tampering traces, and content restored from occluded areas. It also extracts multi-dimensional text features from product text descriptions that can reflect word meaning, pronunciation, and character structure. This allows for a comprehensive perception of the methods used by operators to hide prohibited or restricted product information, such as image steganography and text deformation. Based on this, by fusing visual and text features, it can effectively capture deep semantic relationships or conflicts between product images and text, thereby significantly improving the accuracy and efficiency of identifying obscure prohibited or restricted sales behaviors. This enables efficient and accurate identification of various prohibited or restricted sales behaviors in transaction scenarios. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 A flowchart illustrating a method for identifying prohibited or restricted sales behaviors provided in this application embodiment; Figure 2 A schematic diagram of a prohibited or restricted sales behavior recognition model provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a device for identifying prohibited or restricted sales behaviors provided in an embodiment of this application; Figure 4 To and Figure 3 The illustrated embodiment provides a schematic diagram of the electronic device corresponding to the prohibited and restricted sales behavior recognition device. Detailed Implementation
[0015] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0016] It should be noted that, in the cases involving user information in the embodiments of this application, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse. In addition, the various models involved in this application (including but not limited to language models or large models) comply with relevant laws and standards.
[0017] Furthermore, the timing of the steps in the following method embodiments is merely an example and not a strict limitation.
[0018] Prohibited or restricted sales activities refer to the actions of operators in online transactions that directly or indirectly publish, promote, or facilitate transactions involving goods or services that are prohibited or restricted from sale by laws, regulations, or platform rules.
[0019] Prohibited and restricted sales activities not only include openly listing prohibited items, but also often present them in a covert manner, such as through text distortion, image steganography, semantic spoofing, and cross-regional traffic diversion, in order to circumvent the supervision of trading platforms.
[0020] These methods include: text distortion, such as using homophones, similar-looking characters, pinyin abbreviations, misspellings, or word order changes to distort information about prohibited or restricted goods in product titles or descriptions; image steganography, such as embedding information about prohibited or restricted goods in product main images or detail images using watermarks, handwritten fonts, partial obscuring (such as mosaics or stickers), image flipping, or color interference; semantic camouflage, such as indirectly conveying information about prohibited or restricted goods through industry jargon, metaphorical patterns, or background elements; and cross-regional traffic diversion, which involves transferring information about prohibited or restricted goods to areas that are difficult for the platform to monitor, such as publishing information about prohibited or restricted goods in buyer review sections, private chat interfaces, or third-party traffic links, to achieve a disguised sales behavior that is ostensibly compliant but secretly conducted.
[0021] In practice, businesses continuously update information on prohibited and restricted goods, resulting in diverse and rapidly evolving prohibited and restricted goods behaviors that are difficult to identify during transaction supervision.
[0022] To address at least one of the aforementioned technical problems, this application provides a method for identifying prohibited and restricted sales behaviors. This method extracts multi-dimensional visual features from product images that characterize global semantics, frequency domain tampering traces, and content recovered from occluded areas. It also extracts multi-dimensional text features from product text descriptions that reflect word meaning, pronunciation, and character structure. This allows for a comprehensive perception of the methods used by businesses to hide prohibited and restricted sales information, such as image steganography and text deformation. Furthermore, by fusing visual and text features, the method effectively captures deep semantic relationships or conflicts between product images and text, thereby significantly improving the accuracy and efficiency of identifying subtle prohibited and restricted sales behaviors. This enables efficient and accurate identification of various prohibited and restricted sales behaviors in transaction scenarios.
[0023] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.
[0024] The method for identifying prohibited or restricted sales behaviors provided in this application can be executed by an electronic device, which can be a terminal device such as a PC, laptop, or smartphone, or a server. The server can be a physical server containing an independent host, a virtual server, a cloud server, or a server cluster.
[0025] Figure 1 A flowchart of a method for identifying prohibited or restricted sales behaviors provided in this application embodiment is shown below. Figure 1 As shown, it may include the following steps: 101. Obtain the product images and text descriptions of the target product displayed on the trading platform.
[0026] The target product refers to any product displayed by the operator on a trading platform (such as an e-commerce platform or social networking platform) that is available for users to browse or initiate transactions. The product image refers to the main image or details page image of the product published by the operator on the trading platform, or image frames from a video showcasing the target product. The product text description includes text information such as the product title and product details. The product text description may include text information from the product image.
[0027] In the specific implementation process, the relevant interface can be called to obtain the product information (i.e., product images and product text descriptions) of the target products published by the operator on the trading platform. Then, based on the obtained product information of the target products, it can be identified whether the operator's behavior of displaying the product information of the target products to the outside world, allowing users to browse or initiate transactions on the trading platform is a prohibited or restricted sales behavior.
[0028] 102. Extract the semantic features and frequency domain features corresponding to the product image, as well as the restoration features obtained after the product image is processed for content completion; the semantic features are used to describe the product category and explicit prohibited or restricted items in the product image, the frequency domain features are used to describe digital tampering traces in the product image, and the restoration features are used to describe information that is physically obscured in the product image.
[0029] In practice, in order to evade supervision of prohibited and restricted goods, operators will modify the product information of prohibited and restricted goods in various ways, so that they are presented on the trading platform in a more obscure and difficult-to-detect form.
[0030] In order to accurately perceive the information provided by the product image, this application embodiment uses three different methods to extract features from the product image in different dimensions, thereby comprehensively perceiving the explicit content, digital tampering traces and physically obscured implicit information in the image from different dimensions, in order to deal with diverse image steganography avoidance methods.
[0031] It is worth noting that the features mentioned in the embodiments of this application are expressed in vector form.
[0032] The first feature extraction method involves extracting semantic features of the product image using models such as Convolutional Neural Networks (CNN). Image semantic features Used to describe the product category and explicit prohibited or restricted items in product images.
[0033] Here, "product category" refers to the classification determined based on the appearance, use, or function of the product in the product image, such as police equipment, medical devices, tobacco products, and wildlife products. Optionally, a corresponding product category knowledge base can be pre-built so that, when extracting semantic features from the image, the product category corresponding to the product contained in the product image can be determined using this product category knowledge base.
[0034] Among them, explicit prohibited or restricted content refers to items, logos, or patterns that are directly displayed in product images and are explicitly prohibited or restricted from sale by laws, regulations, or platform rules, such as police badges, gun outlines, and endangered animal fur. These can be directly identified as illegal elements without additional context.
[0035] The second feature extraction method involves separating high-frequency residual components from the product image using Discrete Cosine Transform (DCT) or wavelet transform. Based on these high-frequency residual components, convolutional layers are used to extract the frequency domain features of the product image. , Used to describe digital tampering traces in product images, such as watermark embedding, image flipping, color interference, etc.
[0036] In the specific implementation process, optionally, the product image can first be divided into multiple fixed-size image blocks (e.g., 8×8 pixels), and then a DCT transform can be performed on each image block to obtain the coefficient matrix of each image block in the frequency domain. The coefficients of the high-frequency region in the DCT coefficient matrix are then extracted as high-frequency residual components. These high-frequency residual components reflect the detailed texture, edge variations, and noise information of the product image, and are the main carriers of digital tampering (such as watermark embedding, image flipping, and image stitching). Subsequently, the high-frequency residual components are reconstructed into a two-dimensional spatial domain feature map (i.e., the DCT residual map), and this map is fed as input into a convolutional layer to extract frequency domain features representing abnormal distributions of the product image in the spectrum from the DCT residual map. .
[0037] The third feature extraction method involves using image completion models such as Convolutional Long Short-Term Memory (ConvLSTM) networks to recover features from areas physically obscured by elements like mosaics or stickers in product images. This process generates feature vectors that represent the original image content of the obscured areas, thus restoring the features. .
[0038] In an optional embodiment, in order to improve recovery features To improve accuracy, optionally, a product knowledge base can be pre-built, which stores product images of various products (including prohibited and restricted products and non-prohibited and restricted products). Thus, when performing feature completion (i.e., feature restoration) on product images using an image completion model, the product knowledge base can serve as prior knowledge, guiding the image completion model in feature completion.
[0039] This solution, through the joint extraction of three types of features—image semantic features, frequency domain features, and restoration features—can establish a complete perception chain covering everything from global semantics to local tampering and content loss. This significantly improves the ability to perceive the visual features of product images and can counter various micro-level countermeasures by businesses against product images, such as steganography, occlusion, tampering, and camouflage. It has strong robustness and high anti-interference capability.
[0040] 103. Determine the visual features of the target product based on the image's semantic features, frequency domain features, and restoration features.
[0041] In an alternative embodiment, image semantic features can be addressed through a channel attention mechanism. Frequency domain characteristics and recovery features Adaptive weighted fusion processing is performed to integrate multi-source heterogeneous image features into a unified and compact visual feature set. : in, This represents the concatenation operation of feature vectors. The weights are dynamically learned through a channel attention mechanism, representing the semantic features of the image. Frequency domain characteristics and recovery features The degree of importance or contribution to visual features. Among them, .
[0042] In this scheme, visual features not only preserve the original semantic information, but also integrate tampering traces and restored content, providing a comprehensive and discriminative visual input for subsequent identification of prohibited and restricted sales behaviors.
[0043] 104. Extract the semantic features, speech features, and glyph features corresponding to the product text description.
[0044] In this embodiment, by decoupling text information from three independent dimensions—semantics, pronunciation, and character shape—it addresses the text distortion and circumvention methods employed by businesses in practical applications, such as homophonic substitution, character splitting and deformation, and symbol insertion, thereby comprehensively capturing the potential prohibition or restriction intent in text information such as product titles and details.
[0045] The first dimension involves encoding the product text description using a pre-trained language model (such as BERT, RoBERTa, etc.) and extracting its high-level semantic embedding representation to obtain the text semantic features. Among them, text semantic features Used to describe the meaning of words and contextual semantics in product text descriptions.
[0046] The second dimension involves converting the product text description into a pinyin sequence, and then extracting its acoustic features, i.e., speech features, using a text-to-speech (TTS) model. Speech features Used to describe the pronunciation information of text, based on speech features It can identify variant words at the pronunciation level, such as homophonic avoidance through homophones or near-homophones.
[0047] The third dimension involves rendering the text corresponding to the product description into an image and using techniques such as CNNs to extract the stroke structure features of the text, i.e., the glyph features. Character features Used to describe the stroke structure and shape of Chinese characters or letters, based on glyph features. It can identify glyph disguises using character decomposition or similar - shaped characters, such as similar - shaped characters (e.g., "烟" → "咽"), etc.
[0048] In an optional embodiment, to improve the accuracy of text feature extraction, a knowledge base of prohibited and restricted goods can be pre - constructed, which stores variant forms such as the names of common prohibited items, homophones, similar - shaped characters, and black - market expressions. When extracting text semantic features,词性标注 or rule filtering can be performed in combination with this knowledge base to further enhance the ability to identify prohibited and restricted information in the veiled expressions in the commodity text description.
[0049] In this solution, by jointly extracting these three types of features: text semantic features, speech features, and glyph features, the true referential intention of the commodity text description can be restored, significantly improving the ability to perceive text features of the commodity text description, and it can cope with diverse micro - countermeasure means such as lexical variants and black - market expressions used by operators in the commodity text description.
[0050] 105. Determine the text features of the target commodity according to the text semantic features, speech features, and glyph features.
[0051] In an optional embodiment, the text semantic features , speech features and glyph features can be interactively fused through the multi - head self - attention mechanism to deeply fuse multi - dimensional text features into a unified and compact text feature : where: represents the calculation function of the multi - head self - attention mechanism, are learnable weight matrices corresponding to the query Q, key K, and value V respectively.
[0052] In this solution, the multi - head self - attention mechanism allows different features to pay attention to and weight each other, thereby learning the internal association and complementary relationship between the three types of features. For example, when "枪" is written as "强", the speech feature may prompt the same pronunciation, the semantic feature may identify that there is a semantic association between "强" and "枪" in the context, and the glyph feature may find that "强" and "枪" are structurally similar; through the self - attention mechanism, these clues can be integrated into a comprehensive text feature.
[0053] In this solution, the text features retain the original semantic meaning of the commodity text description and also incorporate variant information at the pronunciation and glyph levels, providing a more discriminative input for judging the consistency between text and images in the subsequent prohibited and restricted behavior recognition process.
[0054] By fusing visual and textual features, the fused features of the target product can be obtained.
[0055] In real-world online transactions, businesses often employ cross-modal circumvention strategies to evade platform oversight of prohibited or restricted goods. For example, they might display actual prohibited items (such as replica guns or controlled knives) in product images, while using legal terminology in the text description (such as model toys or outdoor tools). Alternatively, they might imply prohibited or restricted content in the text (such as special purposes or internal channels), while the images present a harmless appearance. These practices make it difficult to accurately determine whether a business has illegal intentions based solely on images or text, easily leading to missed or incorrect judgments.
[0056] Based on this, this solution deeply integrates visual and textual features, not only preserving the discriminative information of each modality but also explicitly modeling the semantic alignment relationship between the image and the text. When the content presented by the image is highly consistent with the semantics described by the text, the fused features will strengthen this consistency signal; while when there is a significant semantic deviation or logical conflict between the two, the fused features will highlight the inconsistency, serving as an important basis for identifying the operator's deceptive behavior.
[0057] As an optional feature fusion method, fusing visual and textual features to obtain the fused features of the target product can be implemented as follows: Calculate the feature similarity between visual features and text features, wherein the feature similarity is used to reflect the degree of semantic conflict between the visual features and the text features; Based on feature similarity, determine the first weight corresponding to the visual feature and the second weight corresponding to the text feature; Based on the first weight and the second weight, the visual features and text features are weighted and concatenated to obtain the fused features of the target product.
[0058] To elaborate further, in the specific implementation process, optionally, visual features can be first... and text features Mapping to the same semantic space Z to obtain visual features Corresponding visual mapping features and text features Corresponding text mapping features .
[0059] Next, calculate the visual mapping features. Text mapping features Cosine similarity between them, as a visual feature Text features Feature similarity between : Among them, feature similarity Can be used as a visual feature Text features Semantic consistency score between features. Feature similarity. The value ranges from [-1, 1]. The closer it is to 1, the more consistent the text and images are; the closer it is to -1, the more serious the conflict is.
[0060] Alternatively, visual mapping features can also be calculated. Text mapping features Euclidean distance, Manhattan distance, etc., are used as visual features. Text features Feature similarity between them.
[0061] Optionally, a semantic conflict-aware gating network can be pre-built, so that when feature similarity... If the similarity is less than a preset similarity threshold τ (e.g., τ=0.5), it is considered that there is a semantic conflict or inconsistency between the image and text, which may be a disguised behavior to circumvent the supervision of prohibited and restricted goods. Furthermore, a gating network is used to amplify visual features. The corresponding first weight and reduce text features The corresponding second weight The goal is to generate a collision mask. Dynamically adjust visual features The corresponding first weight Text features The corresponding second weight .
[0062] Alternatively, the gating network can be implemented by a neural network, with feature similarity as input and weight adjustment coefficient as output.
[0063] In this solution, visual features are enhanced. The corresponding first weight This can enhance reliance on image information and reduce text features. The corresponding second weight This can reduce trust in text information that may have been tampered with.
[0064] Finally, based on the first weight Second weight visual features Text features Weighted splicing is performed to obtain the fusion characteristics of the target product. : in, .
[0065] This solution introduces a multimodal semantic conflict gating and joint inference mechanism, which can accurately calculate the inconsistency score between the visual content and text description of a target product. When a serious semantic conflict is detected between the visual content and the text description, the solution can dynamically enhance the weight of visual features and suppress the weight of text features, thereby effectively identifying and combating prohibited or restricted sales behaviors that are inconsistent with the actual content of the product.
[0066] 107. Based on the fusion characteristics, calculate the predicted probability that the act of providing the target product on the trading platform belongs to the prohibited or restricted sales behavior.
[0067] The act of providing target goods on a trading platform refers to the act of an operator displaying product information (including image information and text description information) to the outside world on the trading platform, making it available for users to browse or initiate transactions.
[0068] After obtaining the fusion features Subsequently, based on this fusion characteristic... This study aims to determine the degree of risk associated with a business operator's act of publishing product information on a trading platform, thereby providing an interpretable probabilistic basis for subsequent regulatory decisions.
[0069] Alternatively, it can be based on fusion features A normalized prediction probability P is generated through, for example, a nonlinear transformation. This prediction probability P characterizes the likelihood that the display of a target product on the trading platform constitutes a prohibited or restricted activity. For instance, if the prediction probability P exceeds a preset threshold θ (e.g., θ=0.8), then offering the target product on the trading platform is considered highly likely to be a prohibited or restricted activity, triggering manual review or automatic removal. If the prediction probability P does not exceed the preset threshold θ, it is considered low-risk, and offering the target product on the trading platform is permitted.
[0070] In this embodiment, by jointly extracting semantic features, frequency domain features, and restoration features of product images, and simultaneously decoupling the semantic, speech, and glyph features of text, a multimodal recognition mechanism covering the dual perception dimensions of "explicit content - digital tampering - physical occlusion" and "word meaning - pronunciation - glyph" is constructed. On this basis, by calculating the similarity between visual and text features in a unified semantic space, the fusion weight of the two is dynamically adjusted so that when there is a semantic conflict between the image and the text, the dependence on the high-confidence modality (such as the image) is automatically enhanced, and the interference of the disguised text is suppressed. This effectively identifies the diverse micro-adversarial methods used by operators, such as image steganography, homophonic substitution, character splitting and deformation, and inconsistency between images and text, significantly improving the detection accuracy and anti-interference ability of obscure prohibited and restricted sales behaviors.
[0071] In an optional embodiment, after calculating the prediction probability that the act of providing a target commodity on the trading platform belongs to a prohibited or restricted act, further, the attention degrees for different image regions in the process of calculating the above prediction probability can be reversely mapped to the commodity image to generate a visual heat map corresponding to the prediction probability, where the visual heat map is used to provide an explanatory basis for the accuracy evaluation of the prediction probability.
[0072] In the foregoing embodiment, in step 103, when adaptively weighted fusion processing is performed on the image semantic features , frequency domain features and restored features through the channel attention mechanism, the weight coefficients of each channel learned reflect the attention degrees in different feature dimensions; in step 105, when interactive fusion processing is performed on the text semantic features , speech features and glyph features through the multi-head self-attention mechanism, the query Q-key K-value V response matrix output by the attention mechanism contains the attention distribution information of the internal spatial positions of different features; in step 106, when visual features and text features are fused, the first weight and the second weight also reflect the attention degrees for the visual features and text features .
[0073] To achieve the visual interpretation of the prediction result (i.e., the prediction probability), the attention degrees for different image regions in the process of calculating the above prediction probability (i.e., the above-mentioned attention distribution) can be back-projected through the Class Activation Mapping (CAM) technology or its variants (such as Grad-CAM) technology, so as to associate the feature map that contributes more to the final prediction in the high-level neural network with its corresponding spatial position, and generate a two-dimensional heat map with the same size as the original commodity image, which is called the visual heat map in the embodiments of this application.
[0074] In the visual heat map, the brighter the color (such as red) of a region, the higher the attention degree given to this region when judging the prohibited or restricted act, usually corresponding to the position of the key violation elements, such as the outline of a firearm, the pattern of a police badge, an occluded area, etc.; while the darker the color (such as blue) of a region, the lower the possibility that there are violation elements in this region.
[0075] Visual heatmaps not only visually demonstrate the decision-making basis for calculating prediction probabilities, but also serve as an auxiliary tool for manual review to assess the reliability of prediction results. For example, if brighter areas (i.e., high-attention areas) in a visual heatmap are concentrated in non-sensitive areas, it may indicate a risk of misjudgment, requiring further manual verification. Conversely, if brighter areas in a visual heatmap focus on clearly prohibited or restricted goods, it can enhance the credibility of automated identification and processing of prohibited or restricted goods.
[0076] Furthermore, based on the high-interest areas identified by the heatmap, an active learning loop can be constructed by combining an uncertainty sampling strategy: product samples with low prediction confidence and suspicious traces shown in the heatmap are labeled and sent to the training set to continuously optimize the model's ability to identify new evasion methods, thereby maintaining the system's long-term recognition accuracy and adaptability.
[0077] In an optional embodiment, the method for identifying prohibited or restricted sales behaviors described in the foregoing embodiments can be executed through a prohibited or restricted sales behavior identification model.
[0078] Figure 2 A schematic diagram of a prohibited or restricted sales behavior recognition model provided in this application embodiment is shown below. Figure 2 As shown, the prohibited and restricted sales behavior recognition model may include: an image feature extraction module, a text feature extraction module, a multimodal fusion module, and a classifier.
[0079] The image feature extraction module is used to extract semantic features, frequency domain features, and restored features from product images.
[0080] Optionally, the image feature extraction module can construct a three-stream parallel network structure, specifically a spatial semantic stream, a frequency domain adversarial stream, and an occlusion completion stream. The spatial semantic stream is used to extract semantic features from the product image, the frequency domain adversarial stream is used to extract frequency domain features from the product image, and the occlusion completion stream is used to extract restored features from the product image. The specific extraction processes for image semantic features, frequency domain features, and restored features can be referred to the aforementioned embodiments and will not be repeated here.
[0081] The text feature extraction module is used to extract semantic features, speech features, and glyph features from product text descriptions.
[0082] Optionally, the text feature extraction module may include a text encoder with three independent embedding layers: a semantic embedding layer, a pinyin embedding layer, and a glyph embedding layer. The semantic embedding layer extracts semantic features from the product text description, the pinyin embedding layer extracts speech features, and the glyph embedding layer extracts glyph features. The specific extraction processes for the semantic, speech, and glyph features can be found in the aforementioned embodiments and will not be repeated here.
[0083] The multimodal fusion module is used to determine the visual features of the target product based on image semantic features, frequency domain features, and restoration features; to determine the text features of the target product based on text semantic features, speech features, and glyph features; and to fuse the visual features and text features to obtain the fused features of the target product.
[0084] It is worth noting that the multimodal fusion module can be implemented as a single module or as three independent sub-modules. When implemented as three independent sub-modules, the first fusion sub-module is used to fuse image semantic features, frequency domain features, and reconstruction features to determine visual features; the second fusion sub-module is used to fuse text semantic features, speech features, and glyph features to determine text features; and the third fusion sub-module is used to fuse visual features and text features to obtain fused features.
[0085] A classifier is used to calculate the predicted probability that the act of offering a target product on a trading platform belongs to a prohibited or restricted sale behavior based on the fused features, and outputs the predicted probability.
[0086] Understandably, the prohibited and restricted sales behavior recognition model needs to be trained using training samples before use. Specifically, firstly, training samples labeled with "prohibited and restricted sales" or "not prohibited and restricted sales" are obtained. These training samples include product image samples and product text description samples. Then, the product image samples and product text description samples are input into the prohibited and restricted sales behavior recognition model to be trained, so that the model will sequentially perform the following processing steps: Extract the semantic features, frequency domain features, and reconstruction features corresponding to the product image samples, and fuse them to obtain visual features; Extract the semantic features, speech features, and glyph features of the product text description samples, and fuse them to obtain the text features; Based on visual and textual features, the similarity between the images and text is calculated, and the fusion weights of the visual and text features are dynamically adjusted based on this similarity to generate fused features. Based on the fusion features, the predicted probability that the training sample belongs to the prohibited or restricted sales behavior is calculated.
[0087] Subsequently, based on the difference between the predicted probability and the true label of the training sample, the loss value is calculated, and a gradient descent algorithm (such as the Adam optimizer) is used to iteratively update the learnable parameters in the restricted sales behavior recognition model (including but not limited to the convolutional kernel weights of each feature extraction module, the query / key / value mapping matrix in the attention mechanism, the parameters of the fusion gating network, and the fully connected weights of the prediction layer, etc.) to minimize the loss function.
[0088] Repeat the above process until the model converges, thereby obtaining the trained prohibited and restricted sales behavior recognition model. This model can accurately determine whether newly input product information constitutes prohibited or restricted sales behavior.
[0089] However, in the actual model training process, there may be problems such as insufficient number of training samples or limited types, resulting in weak model generalization ability and difficulty in dealing with the ever-evolving avoidance strategies of operators.
[0090] To this end, this application provides an adversarial multimodal sample augmentation method to expand the high-quality training dataset and improve the model's ability to identify micro-adversarial behaviors, specifically including: Obtain initial product samples, which include initial product image samples and initial product text description samples.
[0091] The initial product image sample is subjected to product image information occlusion processing and / or frequency domain digital noise embedding processing to obtain the target product image sample.
[0092] The initial product text description sample is processed by at least one of the following methods: homophonic replacement, character splitting and deformation, symbol insertion, or semantic synonym replacement, in order to obtain the target product text description sample.
[0093] The target product image sample and the target product text description sample, along with the initial product sample, are used as training samples to train the prohibited and restricted sales behavior recognition model.
[0094] In the specific implementation process, an initial product sample can be obtained first, which includes an initial product image sample. and initial product text description sample Among them, the initial product image sample The initial product text description sample consists of genuine product images that have not been altered in any way (such as images of replica guns, controlled knives, and other prohibited items). Provide its corresponding legal or neutral description (such as model toys, outdoor tools, etc.).
[0095] For the initial product image sample Optionally, at least one of the following image enhancement processes can be performed using a Generative Adversarial Network (GAN) to generate target product image samples. This allows us to simulate the micro-level countermeasures employed by business operators in actual transactions.
[0096] 1) Physical occlusion handling: This is achieved through methods such as GANs on the initial product image samples. Random noise blocks, mosaics, or images of irrelevant objects can be added to key areas (such as the barrel of a firearm, the blade of a knife, or the markings on controlled items) to obscure them. For example, semi-transparent stickers can be added to key areas of a police uniform to disguise it as an ordinary uniform, such as a security guard uniform, simulating the behavior of business operators using physical objects to cover up prohibited key information; blurring can also be inserted into key information areas using image editing tools to create a visual information hiding effect.
[0097] 2) Digital adversarial examples: using methods such as GANs to sample initial product images. Frequency domain embedding of digital noise invisible to the human eye but perceptible to the model can simulate traces of image tampering. For example, by injecting weak perturbations into specific frequency components through DCT transform, the model can simulate image resampling, compression artifacts, or copy-and-move forgeries. Alternatively, adversarial image samples can be generated through GANs to introduce local distortions without changing the overall semantics, thereby improving the model's robustness against digital spoofing.
[0098] For the initial product text description sample Optionally, a process including at least one of the following can be constructed to obtain a sample text description of the target product: homophones, character splitting, insertion of special symbols, or semantic synonym replacement. For example, the initial product text description sample. The word "gun" is replaced with "forced," "robbed," or "clawed," and "controlled" is replaced with "special supply" or "internal channels," etc. Through diversified semantic masquerading, a sample text description of the target product is obtained. .
[0099] Finally, the enhanced target product image samples Sample text description of the target product Perform preprocessing operations such as normalization and serialization, and compare with the initial product image samples. and initial product text description sample Together, they form an enhanced multimodal dataset D={ + + + }, and used for model training of the restricted sales behavior recognition model.
[0100] In this scheme, sample augmentation is used to effectively simulate common real-world evasion techniques such as image steganography, physical occlusion, digital tampering, homophonic substitution, and character decomposition, which can significantly improve the model's generalization ability and anti-interference performance in complex scenarios.
[0101] In practical applications, the model for recognizing prohibited and restricted sales behaviors needs to be used on trading platforms for a long time, facing the ever-evolving circumvention methods of operators (such as new homophones, image occlusion methods, and digital tampering techniques). Traditional statically trained models are difficult to adapt to such dynamic changes and are prone to missed or false judgments.
[0102] To address this, this application designs a closed-loop active learning mechanism: During use, the prohibited and restricted sales behavior recognition model not only outputs predicted probabilities but also evaluates the consistency between the product image and text description of the target product input to the model, as well as the uncertainty of the predicted probability. For product images and text descriptions with high predicted probability uncertainty and significant semantic conflicts between the images and text, they are marked as "high-risk difficult examples" and pushed to a manual review queue for annotation. Subsequently, the manually annotated data is re-injected into the training set for periodic fine-tuning of the model. Through this process, a shift from "passive response" to "active evolution" is achieved, enabling the prohibited and restricted sales behavior recognition model to possess continuous learning capabilities, thereby effectively responding to the rapid iteration of circumvention methods for prohibited and restricted products in practical applications.
[0103] In the specific implementation process, the closed-loop active learning mechanism includes: after calculating the predicted probability, determining the degree of semantic conflict between visual features and text features; if the degree of semantic conflict exceeds the preset condition, calculating the uncertainty of the predicted probability calculated by the classifier; if the uncertainty is greater than the set threshold, using the product image and product text description as supplementary training samples to fine-tune the restricted sales behavior recognition model.
[0104] The degree of semantic conflict between visual features and text features can be represented by the feature similarity (e.g., cosine similarity) between visual features and text features in the aforementioned embodiments. For example, when the feature similarity is less than a certain preset threshold... θconflict When the value is 0.3, a serious semantic conflict is determined, and the process proceeds to the uncertainty assessment stage of the subsequent prediction probability.
[0105] Alternatively, the entropy value of the predicted probability output by the classifier can be calculated. To measure the uncertainty vector: Where C is the total number of categories classified by the classifier. For the first The predicted probability of a class.
[0106] Assuming the classifier performs a binary classification task, that is, determining whether offering a target product on a trading platform constitutes a prohibited or restricted activity, then C=2. It can predict the probability that offering a target product on a trading platform constitutes a prohibited or restricted activity. It can predict the probability that offering a target product on a trading platform does not constitute a prohibited or restricted sale behavior.
[0107] Where H(P) is larger, it indicates that the model's prediction for that sample is more uncertain. For example, when =0.45, When H(P) is close to its maximum value at 0.55, it indicates that the model cannot clearly determine whether the behavior constitutes a sales restriction or prohibition.
[0108] When uncertainty Greater than the set threshold θuncertainty If the value is 0.7, the product image and product text description of the target product will be judged as a "high conflict-high uncertainty" difficult case and pushed to the manual review queue for label confirmation.
[0109] Finally, manually labeled difficult examples are added to the original training set to construct an incremental training dataset. The restricted sales behavior recognition model is fine-tuned using mini-batch gradient descent (such as the Adam optimizer) to update the model parameters and enable it to better distinguish new types of spoofing behaviors.
[0110] Through the aforementioned closed-loop active learning mechanism, accurate identification of novel, subtle, and cross-modal prohibited and restricted sales behaviors is achieved. On the one hand, the semantic conflict detection mechanism can effectively capture "image-text mismatch" disguised behaviors; on the other hand, uncertainty sampling ensures that the model prioritizes learning the most challenging samples, avoiding overfitting on simple samples. In addition, this mechanism enables the model to continuously absorb newly emerging prohibited and restricted sales behavior disguised strategies without full retraining, maintaining long-term recognition accuracy. For novel and subtle prohibited and restricted sales behaviors that are not yet covered by existing rules but have already resulted in real transactions, recognition can be quickly established through human feedback, significantly improving the accuracy and robustness of identifying "difficult cases."
[0111] The following will describe in detail one or more embodiments of the prohibited or restricted sales behavior identification device of this application. Those skilled in the art will understand that these devices can all be configured using commercially available hardware components through the steps taught in this solution.
[0112] Figure 3 This is a schematic diagram of the structure of a device for identifying prohibited or restricted sales behaviors provided in an embodiment of this application, as shown below. Figure 3 As shown, the device includes: an acquisition module 11 and a processing module 12.
[0113] The acquisition module 11 is used to acquire the product image and product text description displayed on the trading platform for the target product.
[0114] Processing module 12 is used to extract the image semantic features and frequency domain features corresponding to the product image, as well as the restored features obtained after content completion processing of the product image; the image semantic features are used to describe the product category and explicit prohibited or restricted sales content in the product image, the frequency domain features are used to describe digital tampering traces in the product image, and the restored features are used to describe information physically obscured in the product image; based on the image semantic features, the frequency domain features, and the restored features, the visual features of the target product are determined; the text semantic features, speech features, and glyph features corresponding to the product text description are extracted; based on the text semantic features, the speech features, and the glyph features, the text features of the target product are determined; the visual features and the text features are fused to obtain the fused features of the target product; based on the fused features, the predicted probability that the act of providing the target product on the trading platform belongs to prohibited or restricted sales behavior is calculated.
[0115] Figure 3 The device shown can perform the steps described in the foregoing embodiments. For detailed execution process and technical effects, please refer to the description in the foregoing embodiments, which will not be repeated here.
[0116] In one possible design, the above Figure 3 The structure of the prohibited and restricted sales behavior recognition device shown can be implemented as an electronic device, such as... Figure 4 As shown, the electronic device may include: a memory 21, a processor 22, and a communication interface 23. The memory 21 stores a computer program, which, when executed by the processor 22, enables the processor 22 to at least implement the restricted sales behavior identification method provided in the foregoing embodiments.
[0117] The aforementioned memory 21 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0118] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above-described method embodiments. The computer-readable storage medium includes volatile or non-volatile components, or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, CD-ROM, Digital Video Disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium. Accordingly, this application also provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are executed by a processor, the processor is able to implement the steps in the above method embodiments. It should be understood that each step or combination of steps in the above method flow can be implemented by the computer program or instructions. In addition, these computer programs or instructions can be applied to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable prohibited and restricted sales behavior recognition device, so that the processor of the general-purpose computer, special-purpose computer, embedded processor, or other programmable prohibited and restricted sales behavior recognition device can be implemented as a means to implement the corresponding functions in the above method embodiments.
[0119] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separate. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0120] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of a necessary general-purpose hardware platform, or by a combination of hardware and software. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a computer product. This application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0121] Finally, it should be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0122] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for identifying prohibited or restricted sales behaviors, characterized in that, The method includes: Obtain the product images and text descriptions displayed on the trading platform for the target product; Extract the image semantic features and frequency domain features corresponding to the product image, as well as the restoration features obtained after the product image is processed for content completion; the image semantic features are used to describe the product category and explicit prohibited or restricted items in the product image, the frequency domain features are used to describe digital tampering traces in the product image, and the restoration features are used to describe information that is physically occluded in the product image; Based on the image semantic features, the frequency domain features, and the restored features, the visual features of the target product are determined; Extract the text semantic features, speech features, and glyph features corresponding to the product text description; Based on the text semantic features, the speech features, and the glyph features, the text features of the target product are determined; The visual features and the text features are fused to obtain the fused features of the target product; Based on the fusion features, the predicted probability that the act of providing the target product on the trading platform constitutes a prohibited or restricted sales act is calculated.
2. The method according to claim 1, characterized in that, The process of fusing the visual features and the text features to obtain the fused features of the target product includes: Calculate the feature similarity between the visual features and the text features, whereby the feature similarity reflects the degree of semantic conflict between the visual features and the text features; Based on the feature similarity, determine the first weight corresponding to the visual feature and the second weight corresponding to the text feature; Based on the first weight and the second weight, the visual features and the text features are weighted and concatenated to obtain the fused features of the target product.
3. The method according to claim 2, characterized in that, The calculation of the feature similarity between the visual features and the text features includes: The visual features and the text features are mapped to the same semantic space to obtain the visual mapping features corresponding to the visual features and the text mapping features corresponding to the text features; Calculate the cosine similarity between the visual mapping features and the text mapping features, and use it as the feature similarity between the visual features and the text features; The step of determining the first weight corresponding to the visual feature and the second weight corresponding to the text feature based on the feature similarity includes: If the feature similarity is less than a preset similarity threshold, the first weight corresponding to the visual feature is increased and the second weight corresponding to the text feature is decreased through a gating network.
4. The method according to claim 1, characterized in that, After calculating the predicted probability, the method further includes: The attention given to different image regions during the calculation of the predicted probability is back-mapped to the product image to generate a visual heatmap corresponding to the predicted probability. The visual heatmap is used to provide an interpretive basis for evaluating the accuracy of the predicted probability.
5. The method according to any one of claims 1 to 4, characterized in that, The method is executed through a restricted sales behavior identification model, which includes: The image feature extraction module is used to extract the image semantic features, the frequency domain features, and the restored features from the product image; The text feature extraction module is used to extract the text semantic features, the speech features, and the glyph features from the product text description; A multimodal fusion module is used to determine the visual features of the target product based on the image semantic features, the frequency domain features, and the recovery features; determine the text features of the target product based on the text semantic features, the speech features, and the glyph features; and fuse the visual features and text features to obtain the fused features of the target product. A classifier is used to calculate, based on the fused features, the predicted probability that the act of providing the target product on the trading platform belongs to the prohibited or restricted sales behavior, and output the predicted probability.
6. The method according to claim 5, characterized in that, After calculating the predicted probability, the method further includes: Determine the degree of semantic conflict between the visual features and the text features; If the degree of semantic conflict exceeds a preset condition, the uncertainty of the classifier in calculating the predicted probability is calculated. If the uncertainty exceeds a set threshold, the product image and the product text description will be used as supplementary training samples to fine-tune the restricted sales behavior recognition model.
7. The method according to claim 5, characterized in that, The method further includes: Obtain an initial product sample, which includes an initial product image sample and an initial product text description sample; The initial product image sample is subjected to product image information occlusion processing and / or frequency domain digital noise embedding processing to obtain the target product image sample; The initial product text description sample is processed by at least one of the following: homophonic replacement, character splitting and deformation, symbol insertion, or semantic synonym replacement, to obtain the target product text description sample. The target product image sample and the target product text description sample, along with the initial product sample, are used as training samples to train the prohibited and restricted sales behavior recognition model.
8. An electronic device, characterized in that, include: The device includes a memory, a processor, and a communication interface; wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the method for identifying prohibited or restricted sales behaviors as described in any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor of an electronic device, causes the processor to perform the method for identifying prohibited or restricted sales activities as described in any one of claims 1 to 7.
10. A computer program product, characterized in that, include: A computer program or instruction that, when executed by a processor of an electronic device, causes the processor to perform the method for identifying prohibited or restricted sales as described in any one of claims 1 to 7.