E-commerce platform multi-mode commodity data automatic quality inspection method and system
By constructing a multimodal entity scene graph and an e-commerce knowledge graph using multimodal learning technology, and combining semantic extraction and quality inspection rules, the problem of identifying violations of inconsistent images and text in e-commerce platforms was solved, achieving efficient and accurate quality inspection results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGZHOU ANJIE BIOTECHNOLOGY CO LTD
- Filing Date
- 2026-02-06
- Publication Date
- 2026-05-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing e-commerce platform quality inspection methods are inefficient, costly, and unable to accurately identify false advertising patterns with deep semantic inconsistencies across modalities, especially contradictions between images and text and hidden violations.
By employing multimodal learning technology, we preprocess and extract features from product data acquired from e-commerce platforms to construct a multimodal entity scene graph. We then combine this graph with an e-commerce knowledge graph for entity recognition and attribute extraction, utilize a semantic extraction model for semantic feature extraction and alignment, establish a quality inspection rule semantic graph for violation inspection, and finally use a violation prediction model for judgment.
It achieves efficient, accurate, and interpretable multimodal semantic fusion of product data from e-commerce platforms, enabling it to identify complex violation patterns and provide accurate quality inspection results.
Smart Images

Figure CN122048152A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of e-commerce and artificial intelligence technology, specifically relating to an automatic quality inspection method and system for multimodal product data on e-commerce platforms. Background Technology
[0002] With the rapid development of e-commerce, the real-time release and updating of massive amounts of product data (including text descriptions, display images, videos, etc.) has become the norm for platform operations. However, the authenticity, accuracy, and compliance of product information face severe challenges. False advertising, discrepancies between images and text, counterfeit materials, and illegal wording are rampant, seriously infringing on consumer rights, damaging platform reputation, and placing enormous pressure on market supervision. Currently, e-commerce platforms mainly rely on two quality inspection methods: first, manual review, which suffers from inherent defects such as low efficiency, high cost, inconsistent standards, and susceptibility to fatigue and errors when dealing with hundreds of millions of products, making it difficult to achieve large-scale real-time monitoring; second, rule-based automated filtering, which typically uses keyword matching, sensitive word databases, and simple image tag comparison techniques. While it can handle some explicit violations, its intelligence level is low, it cannot understand semantics, and it cannot identify deep semantic inconsistencies across modalities or implicit false advertising patterns. For example, the system cannot determine whether there is a contradiction between the "leather texture" displayed in the image and the "environmentally friendly PU" in the text description, nor can it identify soft violations such as implying "medical effects" on ordinary food.
[0003] In recent years, artificial intelligence technologies, especially multimodal learning and natural language processing, have provided new approaches for automated content moderation. Existing technologies are mainly applied to single-modal detection of misinformation or coarse-grained calculation of image-text similarity. For example, they extract image and text features separately and then perform simple concatenation or calculate global similarity, but fail to model the fine-grained correspondence between objects within an image and entities in the text. This makes it impossible to accurately judge violations such as misattribution (e.g., an image depicting a flood, text describing an earthquake) or carefully designed misrepresentation (e.g., superficially related text and image but with deep semantic contradictions).
[0004] As mentioned above, how to provide an automatic quality inspection method and system for multimodal product data of e-commerce platforms that can deeply integrate multimodal semantics and make accurate, efficient and interpretable quality inspection judgments on complex violation patterns has become an urgent problem to be solved in this field. Summary of the Invention
[0005] The purpose of this invention is to provide an automatic quality inspection method and system for multimodal product data on e-commerce platforms, in order to solve the above-mentioned problems existing in the prior art.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides an automatic quality inspection method for multimodal product data on an e-commerce platform, comprising: The original product data of the e-commerce platform is obtained, and the original product data of the e-commerce platform is preprocessed and multimodal feature extracted to obtain multimodal feature data of e-commerce platform products. The multimodal feature data of e-commerce platform products includes text feature data of e-commerce platform products, image feature data of e-commerce platform products, and image tampering feature data of e-commerce platform products. A preset e-commerce knowledge graph is obtained. Based on the multimodal feature data of products on the e-commerce platform, a multimodal entity scene graph of product data is constructed. Product entity recognition and entity attribute information extraction are performed on the multimodal entity scene graph of product data and the e-commerce knowledge graph. The identified product entities and the entity attribute information are fused in the multimodal entity scene graph of product data to obtain a multimodal enhanced scene graph of product data. Obtain a pre-trained semantic extraction model, use the semantic extraction model to extract semantic features from the multimodal enhanced scene graph of the product data to obtain multimodal semantic features of the product data, and perform semantic alignment and semantic consistency processing on the multimodal semantic features of the product data to obtain semantic matching features of the product data image and text. Establish a semantic graph of quality inspection rules for commodity data, and use the semantic graph of quality inspection rules for commodity data to perform violation inspection on the multimodal enhanced scenario graph of commodity data, so as to generate corresponding semantic features of commodity data violations; A pre-trained violation prediction model is obtained, and the multimodal feature data of the e-commerce platform products, the image and text semantic matching features of the product data, and the violation semantic features of the product data are fused into the violation features of the e-commerce platform product data. The violation prediction model is then used to perform quality inspection and judgment on the violation features of the e-commerce platform product data to obtain the quality inspection results of the e-commerce platform product data.
[0007] In one possible design, raw product data from an e-commerce platform is acquired, and the raw product data is preprocessed and multimodal feature extracted to obtain multimodal feature data of the e-commerce platform products, including: Obtain raw product data from e-commerce platforms, wherein the raw product data from e-commerce platforms includes product text description data and product image data; The product text description data is subjected to invalid character removal and unified encoding filtering. Natural language processing technology is used to extract keywords from the product text description data after invalid character removal and unified encoding filtering to obtain the corresponding product text description keyword sequence, so as to use the product text description keyword sequence as the product text description feature. Using the long short-term attention mechanism, the context of the product text description features is extracted to obtain the product text description context information. The product text description context information and the product text description features are then concatenated to form product text feature data of the e-commerce platform. A preset standard size is obtained, and the product image data is scaled and normalized using the standard size. Multi-scale visual feature extraction is performed on the product image data after image scaling and normalization to obtain a multi-scale visual feature map of the product image. Obtain a preset compression ratio, compress the product image data according to the compression ratio to obtain compressed product image data, perform differential calculation on the product image data and the compressed product image data to obtain differential product image data, and partition the product image differential data into image segments to obtain differential data for each sub-product image. Error calculations are performed on the differential data of each sub-product image to identify sub-product image differential data with different error levels, and differential processing is performed on the sub-product image differential data with different error levels to generate a product image tampering feature map. Feature vectors are extracted from the multi-scale visual feature map of the product image and the tampering feature map of the product image, respectively, to obtain the corresponding e-commerce platform product image feature data and e-commerce platform product image tampering feature data; The e-commerce platform integrates the product text feature data, the e-commerce platform product image feature data, and the e-commerce platform product image tampering feature data to form e-commerce platform product multimodal feature data.
[0008] In one possible design, based on the multimodal feature data of products from the e-commerce platform, a multimodal entity scene graph of product data is constructed, including: Based on named entity recognition technology, entity recognition is performed on the product text feature data of the e-commerce platform to obtain multiple product text entities, and the corresponding semantic connection relationship is identified between each product text entity. Each product text entity includes product category tag information. Each of the product text entities is used as a product text feature node, and the semantic connection relationship between each of the product text entities is used as a product text feature edge. Each of the product text feature edges is used to connect the product text feature nodes accordingly to form a product data text modal entity scene graph. The product image feature data and the product image tampering feature data of the e-commerce platform are subjected to convolution processing to identify multiple product image entities, wherein each product image entity includes product category label information, product image feature information and product spatial location feature information; A pre-trained product association prediction model is obtained. The product image feature information and product spatial location feature information in each product image entity are fused to obtain product association features. The product association features are used as input to the product association prediction model to output a product association prediction probability distribution map. The product association prediction probability distribution map is then selected as the final product association. Each of the product image entities is used as a product image feature node, and the final product association relationship between each of the product image entities is used as a product image feature edge. Each of the product image feature edges is used to connect the product image feature nodes accordingly to form a product data image modal entity scene graph. Obtain a preset standard graph structure. Based on the standard graph structure, perform isomorphic processing on the product data text modal entity scene graph and the product data image modal entity scene graph, and integrate the isomorphic product data text modal entity scene graph and the product data image modal entity scene graph to form a product data multimodal entity scene graph.
[0009] In one possible design, the product data multimodal entity scene graph and the e-commerce knowledge graph are used to perform product entity recognition and entity attribute information extraction. The identified product entities and their attribute information are then fused in the product data multimodal entity scene graph to obtain a product data multimodal enhanced scene graph, including: All nodes are extracted from the multimodal entity scene graph of the product data as product entities, wherein the product entities include product text entities and product image entities; Each of the product entities is used as a query condition. The e-commerce knowledge graph is traversed according to the graph structure to extract the entity attribute information corresponding to each product entity from the e-commerce knowledge graph. The entity attribute information includes the product entity's superordinate concept, the product entity's inherent attribute tags, and the product entity's typical attribute values. Based on the entity attribute information corresponding to each of the product entities, a corresponding product entity attribute node is generated for each product entity, wherein the product entity attribute node corresponding to each product entity includes a product entity concept node and / or a product entity attribute node; An attribute feature edge is generated between each product entity node and its corresponding product entity attribute nodes, wherein the attribute feature edge is used to represent the subordinate relationship between the product entity node and its corresponding product entity attribute nodes; By utilizing the attribute feature edges between each product entity node and its corresponding product entity attribute nodes, each product entity node is connected to the corresponding product entity attribute nodes to form a multimodal enhanced scene graph of product data.
[0010] In one possible design, a pre-trained semantic extraction model is obtained, and the semantic extraction model is used to extract semantic features from the multimodal enhanced scene graph of the product data to obtain multimodal semantic features of the product data, including: Initial semantic feature extraction is performed on the multimodal augmented scene graph of the product data to extract multiple product entity nodes from the multimodal augmented scene graph of the product data; Adjacency node search is performed on each of the product entity nodes to obtain the product entity attribute nodes connected to each of the product entity nodes, and the connection relationship between each product entity node and the corresponding product entity nodes is extracted to form the corresponding initial semantic features of each product entity node. Using multiple weight matrices in the semantic extraction model, the initial semantic features of each product entity node are linearly transformed to generate product node feature vectors and product node adjacency feature vectors corresponding to each weight matrix. For each product entity node, in each set of weight matrices, the attention scores of each neighboring node corresponding to the product entity node are calculated, and all attention scores obtained for each product entity node are normalized. The adjacent feature vectors of each product entity node are then weighted and summed to obtain the adjacent aggregated feature vector corresponding to each product entity node. In each set of weight matrices, the adjacency aggregation feature vector corresponding to each product entity node is residually connected with the product node feature vector corresponding to each product entity node to obtain the product data semantic feature vector corresponding to each product entity node. The semantic feature vectors of the product data generated for each product entity node are integrated from each set of weight matrices, and the vector average is calculated accordingly to obtain the average semantic feature vector of the product data for each product entity node. The average semantic feature vectors of the product data corresponding to each of the product entity nodes are integrated to form an average semantic feature vector matrix of product data. The average semantic feature vector matrix of product data is used as the multimodal semantic features of product data, wherein the multimodal semantic features of product data include textual semantic features and image semantic features of product data.
[0011] In one possible design, semantic alignment and semantic consistency processing are performed on the multimodal semantic features of the product data to obtain the text-image semantic matching features of the product data, including: Cross-attention calculation is performed on the product data text modal semantic features and the product data image modal semantic features to obtain the product data image modal related semantic features corresponding to the product data text modal semantic features and the product data text modal related semantic features corresponding to the product data image modal semantic features; The semantic features of the product data text modality are concatenated with the corresponding semantic features of the product data image modality to obtain the product data text modality enhanced semantic features. The semantic features of the product data image modality are concatenated with the corresponding semantic features of the product data text modality to obtain the product data image modality enhanced semantic features. Obtain preset text attention weights, perform global average pooling on the product data text modality enhanced semantic features and the product data text modality enhanced semantic features, and perform weighted fusion on the product data text modality enhanced semantic features after global average pooling based on the text attention weights to complete semantic alignment processing and obtain product data multimodal collaborative semantic features; A preset semantic similarity perception model is obtained, and the multimodal collaborative semantic features of the product data are used as input to the semantic similarity perception model to generate image-text semantic matching features of the product data through the semantic similarity perception model.
[0012] In one possible design, a semantic graph of product data quality inspection rules is established, and the semantic graph of product data quality inspection rules is used to perform violation inspection on the multimodal enhanced scenario graph of product data to generate corresponding semantic features of product data violations, including: By obtaining basic quality inspection rules and commodity data supervision regulations from e-commerce platforms, semantic recognition is performed on the basic quality inspection rules and commodity data supervision regulations to define multiple commodity data quality inspection violation patterns, and the corresponding pattern feature relationships are identified for each commodity data quality inspection violation pattern. By using each of the aforementioned product data quality inspection violation patterns as quality inspection rule nodes and identifying the corresponding pattern feature relationships of each product data quality inspection violation pattern as quality inspection rule edges, a semantic graph of product data quality inspection rules is constructed. Each product entity node and its corresponding product entity attribute node are extracted from the multimodal enhanced scene graph of the product data. The product data quality inspection rule semantic graph is used to perform violation inspection on each product entity attribute node corresponding to each product entity node. When the product entity attribute node corresponding to the product entity node conforms to the product data quality inspection violation pattern in the product data quality inspection rule semantic graph, the product entity node is considered to have violated the rules. Traverse all the product entity nodes, perform violation inspection on each product entity node to filter out each product entity node that has a violation, and add corresponding product data quality inspection violation pattern tags to each product entity node that has a violation according to the semantic graph of the product data quality inspection rules to form a violation product entity node. The various non-compliant product entity nodes are integrated to form a non-compliant product entity node matrix, which is then used as the semantic feature of non-compliant product data.
[0013] In one possible design, a pre-trained violation prediction model is obtained, and the multimodal feature data of the e-commerce platform products, the image-text semantic matching features of the product data, and the violation semantic features of the product data are fused into violation features of the e-commerce platform product data. The violation prediction model is then used to perform quality inspection and judgment on the violation features of the e-commerce platform product data to obtain the quality inspection results of the e-commerce platform product data, including: The e-commerce platform's multimodal feature data, the product data's image-text semantic matching features, and the product data's violation semantic features are subjected to dimensional standardization processing to obtain standard e-commerce platform product multimodal feature data, standard product data's image-text semantic matching features, and standard product data's violation semantic features; Obtain preset feature importance weights, and then perform a weighted summation on the standard e-commerce platform product multimodal feature data, standard product data image-text semantic matching features, and standard product data violation semantic features according to the feature importance weights to obtain the e-commerce platform product data violation features; A pre-trained violation prediction model is obtained, and the violation features of the e-commerce platform's product data are used as input to the violation prediction model so that the violation prediction model can output the product data violation confidence level. A preset violation confidence threshold is obtained, and the violation confidence threshold is used to perform quality inspection on the violation confidence of the product data to obtain a quality inspection result. The quality inspection result is used as the quality inspection result of the product data on the e-commerce platform. If the violation confidence of the product data is higher than the violation confidence threshold, the quality inspection result is output as "confirmed violation". If the violation confidence of the product data is not higher than the violation confidence threshold, the quality inspection result is output as "confirmed compliance".
[0014] Secondly, this invention provides an automatic quality inspection system for multimodal product data on e-commerce platforms, comprising: The product data acquisition unit is used to acquire the original product data of the e-commerce platform, and to perform data preprocessing and multimodal feature extraction on the original product data of the e-commerce platform to obtain multimodal feature data of the e-commerce platform products. The multimodal feature data of the e-commerce platform products includes text feature data of the e-commerce platform products, image feature data of the e-commerce platform products, and image tampering feature data of the e-commerce platform products. The scene graph generation unit is used to acquire a preset e-commerce knowledge graph, construct a multimodal entity scene graph of product data based on the multimodal feature data of the e-commerce platform, and perform product entity recognition and entity attribute information extraction on the multimodal entity scene graph of product data and the e-commerce knowledge graph. The identified product entities and the entity attribute information are fused in the multimodal entity scene graph of product data to obtain a multimodal enhanced scene graph of product data. The image-text semantic matching unit is used to obtain a pre-trained semantic extraction model, use the semantic extraction model to extract semantic features from the multimodal enhanced scene graph of the product data to obtain multimodal semantic features of the product data, and perform semantic alignment and semantic consistency processing on the multimodal semantic features of the product data to obtain image-text semantic matching features of the product data. The data violation inspection unit is used to establish a semantic graph of product data quality inspection rules, and to use the semantic graph of product data quality inspection rules to perform violation inspection on the multimodal enhanced scene graph of product data, so as to generate corresponding semantic features of product data violations. The feature quality inspection and judgment unit is used to acquire a pre-trained violation prediction model, integrate the multimodal feature data of the e-commerce platform products, the image-text semantic matching features of the product data, and the violation semantic features of the product data into the violation features of the e-commerce platform products, and use the violation prediction model to perform quality inspection and judgment on the violation features of the e-commerce platform products to obtain the quality inspection results of the e-commerce platform products.
[0015] Thirdly, the present invention provides an electronic device comprising a memory, a processor, and a transceiver connected in sequence and communication, wherein the memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the automatic quality inspection method for multimodal product data of e-commerce platforms as described in the first aspect or any possible design of the first aspect.
[0016] Fourthly, the present invention provides a computer-readable storage medium storing instructions that, when executed on a computer, perform the automatic quality inspection method for multimodal product data of an e-commerce platform as described in the first aspect or any possible design of the first aspect.
[0017] Fifthly, the present invention provides a computer program product containing instructions that, when the instructions are executed on a computer, cause the computer to perform the automatic quality inspection method for multimodal product data of an e-commerce platform as described in the first aspect or any possible design of the first aspect.
[0018] Beneficial Effects: This invention provides an automatic quality inspection method and system for multimodal product data on e-commerce platforms, comprising: First, acquiring original product data from the e-commerce platform, and performing data preprocessing and multimodal feature extraction on the original product data to obtain multimodal feature data of e-commerce platform products, wherein the multimodal feature data of e-commerce platform products includes e-commerce platform product text feature data, e-commerce platform product image feature data, and e-commerce platform product image tampering feature data; Second, acquiring a preset e-commerce knowledge graph, constructing a multimodal entity scene graph of product data based on the multimodal feature data of e-commerce platform products, and performing product entity recognition and entity attribute information extraction on the multimodal entity scene graph of product data and the e-commerce knowledge graph, and fusing the identified product entities and their attribute information in the multimodal entity scene graph of product data to obtain multimodal product data. The process involves several steps: First, an enhanced scene graph is constructed. Then, a pre-trained semantic extraction model is obtained, and semantic features are extracted from the enhanced scene graph of the multimodal product data using this model to obtain multimodal semantic features of the product data. Semantic alignment and consistency processing are then performed on these multimodal semantic features to obtain text-image semantic matching features of the product data. Next, a semantic graph of product data quality inspection rules is established, and this graph is used to perform violation inspections on the enhanced scene graph of the multimodal product data to generate corresponding product data violation semantic features. Finally, a pre-trained violation prediction model is obtained, and the multimodal product feature data from the e-commerce platform, the text-image semantic matching features of the product data, and the violation semantic features of the product data are fused into e-commerce platform product data violation features. The violation prediction model is then used to perform quality inspection judgments on these e-commerce platform product data violation features to obtain the e-commerce platform product data quality inspection results. By constructing and enhancing image and text scene graphs, a multimodal enhanced scene graph of commodity data is generated. This allows for semantic feature extraction and semantic consistency processing, resulting in accurate image-text semantic matching features for commodity data. This achieves deep fusion of multimodal semantics, facilitating a deeper understanding of complex data. Furthermore, by constructing a semantic graph of commodity data quality inspection rules, efficient violation pattern recognition is performed on the multimodal enhanced scene graph of commodity data. This generates commodity data violation semantic features containing accurate violation pattern information. Finally, a violation prediction model is used to perform accurate, efficient, and interpretable quality inspection judgments on commodity data. Attached Figure Description
[0019] Figure 1 A flowchart illustrating the automatic quality inspection method for multimodal product data on an e-commerce platform provided in an embodiment of the present invention; Figure 2 This is a functional structure diagram of the automatic quality inspection system for multimodal product data on an e-commerce platform provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the present invention will be briefly introduced below in conjunction with the accompanying drawings and descriptions of the embodiments or the prior art. Obviously, the following description of the structure of the accompanying drawings is only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. It should be noted that the description of these embodiments is for the purpose of helping to understand the present invention, but does not constitute a limitation of the present invention.
[0021] It should be understood that although the terms first, second, etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit, without departing from the scope of the exemplary embodiments of the invention.
[0022] It should be understood that the term "and / or" that may appear in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, B exists alone, and A and B exist simultaneously. The term " / and" that may appear in this document describes another relationship between related objects, indicating that two relationships can exist. For example, A / and B can mean: A exists alone, and A and B exist alone. In addition, the character " / " that may appear in this document generally indicates that the related objects before and after it are in an "or" relationship.
[0023] Example: like Figure 1 As shown, the first aspect of this embodiment provides an automatic quality inspection method for multimodal product data on an e-commerce platform, which may include, but is not limited to, the following steps: S1. Obtain the original product data from the e-commerce platform, and perform data preprocessing and multimodal feature extraction on the original product data from the e-commerce platform to obtain multimodal feature data of the e-commerce platform products. The multimodal feature data of the e-commerce platform products includes text feature data, image feature data, and image tampering feature data of the e-commerce platform products. In one possible implementation, step S1 involves acquiring raw product data from an e-commerce platform and performing data preprocessing and multimodal feature extraction on the raw product data to obtain multimodal feature data of the e-commerce platform products. This step can be broken down into, but is not limited to, the following steps S11-S18, specifically including: S11. Obtain original product data from the e-commerce platform, wherein the original product data from the e-commerce platform includes product text description data and product image data; S12. Perform invalid character removal and unified encoding filtering on the product text description data, and use natural language processing technology to extract keywords from the product text description data after invalid character removal and unified encoding filtering to obtain the corresponding product text description keyword sequence, so as to use the product text description keyword sequence as the product text description feature; S13. Using the long short-term attention mechanism, the context of the product text description features is extracted to obtain the product text description context information. The product text description context information and the product text description features are then concatenated to form e-commerce platform product text feature data. S14. Obtain a preset standard size, and use the standard size to perform image scaling and normalization processing on the product image data, and perform multi-scale visual feature extraction on the product image data after image scaling and normalization processing to obtain a multi-scale visual feature map of the product image. S15. Obtain a preset compression ratio, compress the product image data according to the compression ratio to obtain compressed product image data, perform differential calculation on the product image data and the compressed product image data to obtain differential product image data, and partition the product image differential data into image segments to obtain differential data for each sub-product image. S16. Perform error calculation on each of the sub-product image difference data to identify sub-product image difference data with different error levels, and perform differential processing on the sub-product image difference data with different error levels to generate a product image tampering feature map. S17. Extract feature vectors from the multi-scale visual feature map of the product image and the tampering feature map of the product image respectively to obtain the corresponding e-commerce platform product image feature data and e-commerce platform product image tampering feature data; S18. Integrate the e-commerce platform product text feature data, the e-commerce platform product image feature data, and the e-commerce platform product image tampering feature data to form e-commerce platform product multimodal feature data.
[0024] It should be noted that, in specific applications, the automatic quality inspection method for multimodal product data of e-commerce platforms provided in this embodiment requires that the product text description features be first input into a pre-trained BERT model for the generation process of e-commerce platform product text feature data. The BERT model is built on the Transformer architecture and generates a hidden state vector that integrates bidirectional context for each keyword in the e-commerce platform product text feature data through a self-attention mechanism. The length of this vector is the sequence length of the product text description keyword sequence. In order to further capture the long-distance dependencies in the product text description data, the hidden state vector can be input into a bidirectional long short-term memory network (including forward LSTM units and backward LSTM units) as input. The bidirectional long short-term memory network performs bidirectional processing from the beginning and end of the hidden state vector respectively. Finally, the forward hidden state and backward hidden state generated at each time step are used as product text description context information. The product text description context information and the product text description features are concatenated to obtain the e-commerce platform product text feature data.
[0025] In practical applications, the generation of product image feature data for e-commerce platforms requires first pre-training a Transformer-based multi-scale attention visual feature extraction model based on historical e-commerce platform product data. Product image data that has undergone image scaling and normalization is used as input to the multi-scale attention visual feature extraction model (which includes at least three sequentially connected scales: local detail scale, overall structure scale, and global semantic scale). The multi-scale attention visual model divides the product image data into multiple non-overlapping image blocks, extracts features from these blocks according to different scales, and merges adjacent image blocks after feature extraction at each scale, thereby gradually expanding the receptive field of each feature point until the last scale. The model then outputs and integrates the image features generated at each scale to form a multi-scale visual feature map of the product image. Finally, global average pooling is performed on the multi-scale visual feature map of the product image to aggregate the image features from all spatial locations into a global feature vector, which serves as the product image feature data for the e-commerce platform.
[0026] In addition, in one possible implementation, the multi-scale attention visual feature extraction model can also introduce a sliding window to perform self-attention calculation within the local window for each image block divided in the product image data, and slide the local window (to ensure coverage of the entire image). Each time it slides, only the self-attention of the formed local window is calculated to connect the information images in different windows, avoiding the complex calculation of global attention and completing the efficient extraction of multi-scale visual features.
[0027] Correspondingly, the tampering feature data of e-commerce platform product images is achieved through compression differential calculation. Since the loss of untampered areas of the image is consistent under the same compression ratio, while tampered areas will exhibit different error levels due to different sources, the differences will be highlighted, forming sub-product image differential data with different error levels. Then, the sub-product image differential data is subjected to differential processing such as contrast enhancement or color mapping to generate product image tampering feature maps. The product image tampering feature maps are input into the multi-scale attention visual feature extraction model for feature vector extraction to obtain the e-commerce platform product image tampering feature data.
[0028] S2. Obtain a preset e-commerce knowledge graph, construct a multimodal entity scene graph of product data based on the multimodal feature data of the e-commerce platform, and perform product entity recognition and entity attribute information extraction on the multimodal entity scene graph of product data and the e-commerce knowledge graph. Then, fuse the identified product entities and the entity attribute information in the multimodal entity scene graph of product data to obtain a multimodal enhanced scene graph of product data. In one possible implementation, step S2, based on the multimodal feature data of the e-commerce platform's products, constructs a multimodal entity scene graph of the product data. This can be decomposed, but is not limited to, the following steps S21-S26, specifically including: S21. Based on named entity recognition technology, entity recognition is performed on the product text feature data of the e-commerce platform to obtain multiple product text entities, and corresponding semantic connection relationships are identified between each product text entity, wherein each product text entity includes product category label information; S22. Each of the product text entities is taken as a product text feature node, and the semantic connection relationship between each of the product text entities is taken as a product text feature edge. Each of the product text feature edges is used to connect the product text feature nodes accordingly to form a product data text modal entity scene graph. S23. Perform convolution processing on the product image feature data and the product image tampering feature data of the e-commerce platform to identify multiple product image entities, wherein each product image entity includes product category label information, product image feature information and product spatial location feature information; S24. Obtain a pre-trained product association prediction model, perform feature fusion on the product image feature information and product spatial location feature information in each product image entity to obtain product association features, use the product association features as input to the product association prediction model, output a product association prediction probability distribution map through the product association prediction model, and select the product association with the highest probability from the product association prediction probability distribution map as the final product association; S25. Each of the product image entities is used as a product image feature node, and the final product association relationship between each of the product image entities is used as a product image feature edge. Each of the product image feature edges is used to connect the product image feature nodes accordingly to form a product data image modal entity scene graph. S26. Obtain a preset standard graph structure, and according to the standard graph structure, perform isomorphic processing on the product data text modal entity scene graph and the product data image modal entity scene graph, and integrate the isomorphic product data text modal entity scene graph and the product data image modal entity scene graph as a product data multimodal entity scene graph.
[0029] It should be noted that in the automatic quality inspection method for multimodal product data of e-commerce platforms provided in this embodiment, the product category label information is used to represent the classification of products, such as "dress", "snow boots" or "handbag", etc. The product image feature information can be directly extracted from the product image feature data and the product image tampering feature data of the e-commerce platform, while the product spatial location feature information is used to represent the position coordinate range of each product (object) in the product image, so as to provide a boundary for product identification.
[0030] In practical applications, the product association prediction model described in this embodiment is trained using historical image data from e-commerce platforms. Specifically, its base model structure is constructed using a convolutional neural network, with the probability distribution of all possible product association categories as the model output. By iteratively training this convolutional neural network with a large number of historical product images, an accurate product association prediction model is obtained. This product association prediction model can process product association features to obtain the final product association that best reflects reality.
[0031] Furthermore, in one possible implementation, to avoid the influence of habitual biases introduced by historical product image data on the final product association relationship, thereby leading to prediction errors in the actual product association relationship, a correction mechanism for the final product association relationship can be introduced: For each product image entity, its product category label information and product spatial location feature information are retained, but its product image feature information is replaced with a neutral benchmark feature. The neutral benchmark feature and the product spatial location feature information in each product image entity are then fused to obtain a corrected product association relationship feature. This corrected product association relationship feature is input into the product association relationship prediction model to obtain a corrected product association relationship prediction probability distribution map. The difference between the product association relationship prediction probability distribution map and the corrected product association relationship prediction probability distribution map is calculated. This difference reflects the contribution of product image features to the relationship judgment. Based on this difference, the product association relationship prediction probability distribution map is weighted and adjusted to obtain the adjusted final product association relationship.
[0032] The standard graph structure generally adopts the form of "node-edge-node" triples to complete the isomorphism processing of the product data text modal entity scene graph and the product data image modal entity scene graph.
[0033] In one possible implementation, step S2 involves performing product entity recognition and entity attribute information extraction on the product data multimodal entity scene graph and the e-commerce knowledge graph. The identified product entities and their attribute information are then fused in the product data multimodal entity scene graph to obtain a product data multimodal enhanced scene graph. This step can be, but is not limited to, decomposed into the following steps S27-S211, specifically including: S27. Extract all nodes from the multimodal entity scene graph of the product data as product entities, wherein the product entities include product text entities and product image entities; S28. Using each of the product entities as query conditions, the e-commerce knowledge graph is traversed according to the graph structure to extract the entity attribute information corresponding to each of the product entities from the e-commerce knowledge graph. The entity attribute information includes the product entity's superordinate concept, the product entity's inherent attribute tags, and the product entity's typical attribute values. S29. Based on the entity attribute information corresponding to each of the product entities, generate a corresponding product entity attribute node for each product entity, wherein the product entity attribute node corresponding to each product entity includes a product entity concept node and / or a product entity attribute node; S210. Generate attribute feature edges between each product entity node and its corresponding product entity attribute nodes, wherein the attribute feature edges are used to represent the subordinate relationship between the product entity node and its corresponding product entity attribute nodes; S211. Using the attribute feature edges between each product entity node and the corresponding product entity attribute nodes, connect each product entity node to each product entity attribute node to form a multimodal enhanced scene graph of product data.
[0034] In practical applications, before querying entity attribute information for each product entity, candidate entity attribute information can be identified in the e-commerce knowledge graph. Information similarity is calculated between product description context information (including product text description context information and product image context information) and each candidate entity attribute information. The information similarity is then used to disambiguate each candidate entity attribute information to select the candidate entity attribute information most relevant to the product entity as the entity attribute information. This achieves accurate matching of attribute information to product entities and provides an accurate node foundation for subsequent scene graph enhancement.
[0035] S3. Obtain a pre-trained semantic extraction model, use the semantic extraction model to extract semantic features from the multimodal enhanced scene graph of the product data to obtain multimodal semantic features of the product data, perform semantic alignment and semantic consistency processing on the multimodal semantic features of the product data to obtain semantic matching features of the product data image and text; In one possible implementation, step S3 involves obtaining a pre-trained semantic extraction model and using the semantic extraction model to extract semantic features from the multimodal enhanced scene graph of the product data to obtain multimodal semantic features of the product data. This step can be decomposed into, but is not limited to, the following steps S31-S37, specifically including: S31. Perform initial semantic feature extraction on the multimodal augmented scene graph of the product data to extract multiple product entity nodes from the multimodal augmented scene graph of the product data; S32. Perform an adjacency node search on each of the product entity nodes to obtain the product entity attribute nodes connected to each of the product entity nodes, and extract the connection relationship between each product entity node and the corresponding product entity nodes to form a corresponding initial semantic feature for each product entity node. S33. Using the multiple weight matrices in the semantic extraction model, perform linear transformations on the initial semantic features of each product entity node to generate product node feature vectors and product node adjacency feature vectors corresponding to each weight matrix. S34. For each product entity node, in each set of weight matrices, calculate the attention score of each neighboring node corresponding to the product entity node to the product entity node, and normalize all the attention scores obtained by each product entity node to perform weighted summation on the adjacent feature vectors of each product node corresponding to each product entity node to obtain the adjacent aggregated feature vector corresponding to each product entity node. S35. In each set of weight matrices, the adjacency aggregation feature vector corresponding to each product entity node is residually connected with the product node feature vector corresponding to each product entity node to obtain the product data semantic feature vector corresponding to each product entity node. S36. Integrate the semantic feature vectors of the product data generated for each product entity node by each set of weight matrices, and calculate the average vector of the product data for each product entity node. S37. Integrate the average semantic feature vectors of the product data corresponding to each of the product entity nodes to form an average semantic feature vector matrix of product data, so as to use the average semantic feature vector matrix of product data as the multimodal semantic features of product data, wherein the multimodal semantic features of product data include text modal semantic features of product data and image modal semantic features of product data.
[0036] The initial semantic features of each product entity node include the features inherent to the product entity node (product category label information, product text description context information, product image feature information, and product spatial location feature information), the features of the product entity attribute node (product entity superordinate concept, product entity inherent attribute label, and product entity typical attribute value), and the connection relationship between the product entity node and the corresponding product entity node. Among these features, the features inherent to the product entity node are used as the initial node features, the features of the product entity attribute node are used as the node adjacency features, and the connection relationship between the product entity node and the corresponding product entity node is used as the initial edge features. It should be noted that in the automatic quality inspection method for multimodal product data of e-commerce platforms provided in this embodiment, in order to ensure the accuracy of the output multimodal semantic features of product data and to better evaluate global features, the semantic extraction model can be set as a multi-layer stacked semantic extraction model. Each layer independently executes the semantic extraction process of step SS, and the average semantic feature vector of product data corresponding to each product entity node extracted by each layer is used as the input vector of the next layer. Furthermore, as the number of layers increases, each product entity node can be aggregated into multi-hop adjacent nodes as product entity attribute nodes, thereby capturing a wider range of graph structure semantics. Finally, the average semantic feature vector of product data corresponding to each product entity node output by the last layer is integrated into the average semantic feature vector matrix of product data, forming a more comprehensive and holistic multimodal semantic feature of product data. The semantic extraction model is actually a multi-head attention scene graph attention network model. When extracting semantic features from the multimodal enhanced scene graph of product data, it needs to extract features from the text modality enhanced scene graph and the image modality enhanced scene graph of product data separately to obtain the text modality semantic features and the image modality semantic features of product data, respectively.
[0037] In one possible implementation, step S3 involves performing semantic alignment and semantic consistency processing on the multimodal semantic features of the product data to obtain the text-image semantic matching features of the product data. This can be, but is not limited to, decomposed into the following steps S38-S311, specifically including: S38. Perform cross-attention calculation on the product data text modal semantic features and the product data image modal semantic features to obtain the product data image modal related semantic features corresponding to the product data text modal semantic features and the product data text modal related semantic features corresponding to the product data image modal semantic features; S39. The semantic features of the commodity data text modality are concatenated with the corresponding semantic features of the commodity data image modality to obtain the enhanced semantic features of the commodity data text modality; the semantic features of the commodity data image modality are concatenated with the corresponding semantic features of the commodity data text modality to obtain the enhanced semantic features of the commodity data image modality. S310. Obtain a preset text attention weight, perform global average pooling on the product data text modality enhanced semantic features and the product data text modality enhanced semantic features, and perform weighted fusion on the product data text modality enhanced semantic features after global average pooling based on the text attention weight to complete semantic alignment processing and obtain product data multimodal collaborative semantic features; S311. Obtain a preset semantic similarity perception model, and input the multimodal collaborative semantic features of the product data as input to the semantic similarity perception model, so as to generate image-text semantic matching features of the product data through the semantic similarity perception model.
[0038] It should be noted that in the automatic quality inspection method for multimodal product data of e-commerce platforms provided in this embodiment, the semantic similarity perception model is built through a preset similarity-aware network (SAN), forming a multi-layer perceptron (preferably 3-layer) architecture.
[0039] S4. Establish a semantic graph of product data quality inspection rules, and use the semantic graph of product data quality inspection rules to perform violation inspection on the multimodal enhanced scenario graph of product data, so as to generate corresponding semantic features of product data violations; In one possible implementation, step S4 involves establishing a semantic graph of product data quality inspection rules and using this semantic graph to perform violation inspection on the multimodal enhanced scenario graph of product data to generate corresponding semantic features of product data violations. This step can be broken down into, but is not limited to, the following steps S41-S45, specifically including: S41. Obtain basic quality inspection rules and commodity data supervision regulations through e-commerce platforms, perform semantic recognition on the basic quality inspection rules and commodity data supervision regulations to define multiple commodity data quality inspection violation patterns, and identify the corresponding pattern feature relationships for each commodity data quality inspection violation pattern; S42. Using each of the product data quality inspection violation patterns as quality inspection rule nodes and the corresponding pattern feature relationships identified for each product data quality inspection violation pattern as quality inspection rule edges, construct a product data quality inspection rule semantic graph; S43. Extract each product entity node and its corresponding product entity attribute node from the multimodal enhanced scene graph of the product data. Use the product data quality inspection rule semantic graph to perform violation inspection on each product entity attribute node corresponding to each product entity node. When the product entity attribute node corresponding to the product entity node conforms to the product data quality inspection violation pattern in the product data quality inspection rule semantic graph, it is considered that the product entity node has violated the rules. S44. Traverse all the product entity nodes, perform violation inspection on each product entity node to filter out each product entity node that has a violation, and add corresponding product data quality inspection violation pattern tags to each product entity node that has a violation according to the product data quality inspection rule semantic graph to form a violation product entity node. S45. Integrate the various non-compliant product entity nodes to form a non-compliant product entity node matrix, so as to use the non-compliant product entity node matrix as the semantic feature of product data non-compliance.
[0040] S5. Obtain a pre-trained violation prediction model, integrate the e-commerce platform product multimodal feature data, the product data image-text semantic matching features, and the product data violation semantic features into e-commerce platform product data violation features, and use the violation prediction model to perform quality inspection and judgment on the e-commerce platform product data violation features to obtain the e-commerce platform product data quality inspection results.
[0041] In one possible implementation, step S5 involves obtaining a pre-trained violation prediction model, fusing the e-commerce platform's multimodal product feature data, the product data's image-text semantic matching features, and the product data's violation semantic features into e-commerce platform product data violation features, and using the violation prediction model to perform quality inspection on the e-commerce platform's product data violation features to obtain the e-commerce platform's product data quality inspection results. This can be broken down into, but is not limited to, the following steps S51-S54, specifically including: S51. Perform dimensional standardization processing on the e-commerce platform product multimodal feature data, the product data image-text semantic matching features, and the product data violation semantic features to obtain standard e-commerce platform product multimodal feature data, standard product data image-text semantic matching features, and standard product data violation semantic features; S52. Obtain preset feature importance weights, and perform weighted summation on the standard e-commerce platform product multimodal feature data, standard product data image-text semantic matching features, and standard product data violation semantic features according to the feature importance weights, so as to obtain the e-commerce platform product data violation features; S53. Obtain a pre-trained violation prediction model, and input the violation features of the e-commerce platform's product data as input to the violation prediction model, so as to output the product data violation confidence level through the violation prediction model; S54. Obtain a preset violation confidence threshold, use the violation confidence threshold to perform quality inspection on the violation confidence of the product data, and obtain a quality inspection result. Use the quality inspection result as the quality inspection result of the product data on the e-commerce platform. If the violation confidence of the product data is higher than the violation confidence threshold, the quality inspection result is output as "confirmed violation". If the violation confidence of the product data is not higher than the violation confidence threshold, the quality inspection result is output as "confirmed compliance".
[0042] like Figure 2 As shown, the second aspect of this embodiment provides a hardware system for implementing the automatic quality inspection method for multimodal product data on e-commerce platforms described in the first aspect of the embodiment, including: The product data acquisition unit is used to acquire the original product data of the e-commerce platform, and to perform data preprocessing and multimodal feature extraction on the original product data of the e-commerce platform to obtain multimodal feature data of the e-commerce platform products. The multimodal feature data of the e-commerce platform products includes text feature data of the e-commerce platform products, image feature data of the e-commerce platform products, and image tampering feature data of the e-commerce platform products. The scene graph generation unit is used to acquire a preset e-commerce knowledge graph, construct a multimodal entity scene graph of product data based on the multimodal feature data of the e-commerce platform, and perform product entity recognition and entity attribute information extraction on the multimodal entity scene graph of product data and the e-commerce knowledge graph. The identified product entities and the entity attribute information are fused in the multimodal entity scene graph of product data to obtain a multimodal enhanced scene graph of product data. The image-text semantic matching unit is used to obtain a pre-trained semantic extraction model, use the semantic extraction model to extract semantic features from the multimodal enhanced scene graph of the product data to obtain multimodal semantic features of the product data, and perform semantic alignment and semantic consistency processing on the multimodal semantic features of the product data to obtain image-text semantic matching features of the product data. The data violation inspection unit is used to establish a semantic graph of product data quality inspection rules, and to use the semantic graph of product data quality inspection rules to perform violation inspection on the multimodal enhanced scene graph of product data, so as to generate corresponding semantic features of product data violations. The feature quality inspection and judgment unit is used to acquire a pre-trained violation prediction model, integrate the multimodal feature data of the e-commerce platform products, the image-text semantic matching features of the product data, and the violation semantic features of the product data into the violation features of the e-commerce platform products, and use the violation prediction model to perform quality inspection and judgment on the violation features of the e-commerce platform products to obtain the quality inspection results of the e-commerce platform products.
[0043] The working process, working details and technical effects of the system provided in this embodiment can be found in the first aspect of the embodiment, and will not be repeated here.
[0044] like Figure 3 As shown, the third aspect of this embodiment provides an electronic device, including: a memory, a processor, and a transceiver that are sequentially and communicatively connected, wherein the memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the automatic quality inspection method for multimodal product data of e-commerce platforms as described in the first aspect of the embodiment.
[0045] For specific examples, the memory may include, but is not limited to, random access memory (RAM), read-only memory (ROM), flash memory, first-in-first-out (FIFO) memory, and / or first-in-last-out (FILO) memory, etc.; specifically, the processor may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor may be implemented using at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), PLA (Programmable Logic Array). The processor may also include a main processor and a coprocessor. The main processor, also known as the CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state.
[0046] In some embodiments, the processor may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. For example, the processor may not be limited to microprocessors of the STM32F105 series, reduced instruction set computer (RISC) microprocessors, x86 architecture processors, or processors with integrated neural network processing units (NPUs). The transceiver may be, but is not limited to, a Wi-Fi transceiver, a Bluetooth transceiver, a General Packet Radio Service (GPRS) transceiver, a ZigBee (a low-power LAN protocol based on the IEEE 802.15.4 standard) transceiver, a 3G transceiver, a 4G transceiver, and / or a 5G transceiver. Furthermore, the device may also include, but is not limited to, a power module, a display screen, and other necessary components.
[0047] The working process, working details and technical effects of the electronic device provided in this embodiment can be found in the first aspect of the embodiment, and will not be repeated here.
[0048] The fourth aspect of this embodiment provides a storage medium that stores instructions containing the automatic quality inspection method for multimodal product data of an e-commerce platform as described in the first aspect of the embodiment. That is, the storage medium stores instructions, and when the instructions are run on a computer, the automatic quality inspection method for multimodal product data of an e-commerce platform as described in the first aspect of the embodiment is executed.
[0049] The storage medium refers to a carrier for storing data, which may include, but is not limited to, floppy disks, optical disks, hard disks, flash memory, USB flash drives, and / or memory sticks. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.
[0050] The working process, working details and technical effects of the storage medium provided in this embodiment can be found in the first aspect of the embodiment, and will not be repeated here.
[0051] The fifth aspect of this embodiment provides a computer program product containing instructions that, when executed on a computer, cause the computer to perform the automatic quality inspection method for multimodal product data on an e-commerce platform as described in the first aspect of this embodiment. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.
[0052] Finally, it should be noted that the above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. An automatic quality inspection method for multimodal product data on an e-commerce platform, characterized in that, include: The original product data of the e-commerce platform is obtained, and the original product data of the e-commerce platform is preprocessed and multimodal feature extracted to obtain multimodal feature data of e-commerce platform products. The multimodal feature data of e-commerce platform products includes text feature data of e-commerce platform products, image feature data of e-commerce platform products, and image tampering feature data of e-commerce platform products. A preset e-commerce knowledge graph is obtained. Based on the multimodal feature data of products on the e-commerce platform, a multimodal entity scene graph of product data is constructed. Product entity recognition and entity attribute information extraction are performed on the multimodal entity scene graph of product data and the e-commerce knowledge graph. The identified product entities and the entity attribute information are fused in the multimodal entity scene graph of product data to obtain a multimodal enhanced scene graph of product data. Obtain a pre-trained semantic extraction model, use the semantic extraction model to extract semantic features from the multimodal enhanced scene graph of the product data to obtain multimodal semantic features of the product data, and perform semantic alignment and semantic consistency processing on the multimodal semantic features of the product data to obtain semantic matching features of the product data image and text. Establish a semantic graph of quality inspection rules for commodity data, and use the semantic graph of quality inspection rules for commodity data to perform violation inspection on the multimodal enhanced scenario graph of commodity data, so as to generate corresponding semantic features of commodity data violations; A pre-trained violation prediction model is obtained, and the multimodal feature data of the e-commerce platform products, the image and text semantic matching features of the product data, and the violation semantic features of the product data are fused into the violation features of the e-commerce platform product data. The violation prediction model is then used to perform quality inspection and judgment on the violation features of the e-commerce platform product data to obtain the quality inspection results of the e-commerce platform product data.
2. The automatic quality inspection method for multimodal product data on e-commerce platforms according to claim 1, characterized in that, Obtain raw product data from e-commerce platforms, and perform data preprocessing and multimodal feature extraction on the raw product data to obtain multimodal feature data of e-commerce platform products, including: Obtain raw product data from e-commerce platforms, wherein the raw product data from e-commerce platforms includes product text description data and product image data; The product text description data is subjected to invalid character removal and unified encoding filtering. Natural language processing technology is used to extract keywords from the product text description data after invalid character removal and unified encoding filtering to obtain the corresponding product text description keyword sequence, so as to use the product text description keyword sequence as the product text description feature. Using the long short-term attention mechanism, the context of the product text description features is extracted to obtain the product text description context information. The product text description context information and the product text description features are then concatenated to form product text feature data of the e-commerce platform. A preset standard size is obtained, and the product image data is scaled and normalized using the standard size. Multi-scale visual feature extraction is performed on the product image data after image scaling and normalization to obtain a multi-scale visual feature map of the product image. Obtain a preset compression ratio, compress the product image data according to the compression ratio to obtain compressed product image data, perform differential calculation on the product image data and the compressed product image data to obtain differential product image data, and partition the product image differential data into image segments to obtain differential data for each sub-product image. Error calculations are performed on the differential data of each sub-product image to identify sub-product image differential data with different error levels, and differential processing is performed on the sub-product image differential data with different error levels to generate a product image tampering feature map. Feature vectors are extracted from the multi-scale visual feature map of the product image and the tampering feature map of the product image, respectively, to obtain the corresponding e-commerce platform product image feature data and e-commerce platform product image tampering feature data; The e-commerce platform integrates the product text feature data, the e-commerce platform product image feature data, and the e-commerce platform product image tampering feature data to form e-commerce platform product multimodal feature data.
3. The automatic quality inspection method for multimodal product data on e-commerce platforms according to claim 1, characterized in that, Based on the multimodal feature data of products from the e-commerce platform, a multimodal entity scene graph of product data is constructed, including: Based on named entity recognition technology, entity recognition is performed on the product text feature data of the e-commerce platform to obtain multiple product text entities, and the corresponding semantic connection relationship is identified between each product text entity. Each product text entity includes product category tag information. Each of the product text entities is used as a product text feature node, and the semantic connection relationship between each of the product text entities is used as a product text feature edge. Each of the product text feature edges is used to connect the product text feature nodes accordingly to form a product data text modal entity scene graph. The product image feature data and the product image tampering feature data of the e-commerce platform are subjected to convolution processing to identify multiple product image entities, wherein each product image entity includes product category label information, product image feature information and product spatial location feature information; A pre-trained product association prediction model is obtained. The product image feature information and product spatial location feature information in each product image entity are fused to obtain product association features. The product association features are used as input to the product association prediction model to output a product association prediction probability distribution map. The product association prediction probability distribution map is then selected as the final product association. Each of the product image entities is used as a product image feature node, and the final product association relationship between each of the product image entities is used as a product image feature edge. Each of the product image feature edges is used to connect the product image feature nodes accordingly to form a product data image modal entity scene graph. Obtain a preset standard graph structure. Based on the standard graph structure, perform isomorphic processing on the product data text modal entity scene graph and the product data image modal entity scene graph, and integrate the isomorphic product data text modal entity scene graph and the product data image modal entity scene graph to form a product data multimodal entity scene graph.
4. The automatic quality inspection method for multimodal product data on e-commerce platforms according to claim 3, characterized in that, The product data multimodal entity scene graph and the e-commerce knowledge graph are used to perform product entity recognition and entity attribute information extraction. The identified product entities and their attribute information are then fused in the product data multimodal entity scene graph to obtain a product data multimodal enhanced scene graph, including: All nodes are extracted from the multimodal entity scene graph of the product data as product entities, wherein the product entities include product text entities and product image entities; Each of the product entities is used as a query condition. The e-commerce knowledge graph is traversed according to the graph structure to extract the entity attribute information corresponding to each product entity from the e-commerce knowledge graph. The entity attribute information includes the product entity's superordinate concept, the product entity's inherent attribute tags, and the product entity's typical attribute values. Based on the entity attribute information corresponding to each of the product entities, a corresponding product entity attribute node is generated for each product entity, wherein the product entity attribute node corresponding to each product entity includes a product entity concept node and / or a product entity attribute node; An attribute feature edge is generated between each product entity node and its corresponding product entity attribute nodes, wherein the attribute feature edge is used to represent the subordinate relationship between the product entity node and its corresponding product entity attribute nodes; By utilizing the attribute feature edges between each product entity node and its corresponding product entity attribute nodes, each product entity node is connected to the corresponding product entity attribute nodes to form a multimodal enhanced scene graph of product data.
5. The automatic quality inspection method for multimodal product data on e-commerce platforms according to claim 1, characterized in that, Obtain a pre-trained semantic extraction model, and use the semantic extraction model to extract semantic features from the multimodal enhanced scene graph of the product data to obtain multimodal semantic features of the product data, including: Initial semantic feature extraction is performed on the multimodal augmented scene graph of the product data to extract multiple product entity nodes from the multimodal augmented scene graph of the product data; Adjacency node search is performed on each of the product entity nodes to obtain the product entity attribute nodes connected to each of the product entity nodes, and the connection relationship between each product entity node and the corresponding product entity nodes is extracted to form the corresponding initial semantic features of each product entity node. Using multiple weight matrices in the semantic extraction model, the initial semantic features of each product entity node are linearly transformed to generate product node feature vectors and product node adjacency feature vectors corresponding to each weight matrix. For each product entity node, in each set of weight matrices, the attention scores of each neighboring node corresponding to the product entity node are calculated, and all attention scores obtained for each product entity node are normalized. The adjacent feature vectors of each product entity node are then weighted and summed to obtain the adjacent aggregated feature vector corresponding to each product entity node. In each set of weight matrices, the adjacency aggregation feature vector corresponding to each product entity node is residually connected with the product node feature vector corresponding to each product entity node to obtain the product data semantic feature vector corresponding to each product entity node. The semantic feature vectors of the product data generated for each product entity node are integrated from each set of weight matrices, and the vector average is calculated accordingly to obtain the average semantic feature vector of the product data for each product entity node. The average semantic feature vectors of the product data corresponding to each of the product entity nodes are integrated to form an average semantic feature vector matrix of product data. The average semantic feature vector matrix of product data is used as the multimodal semantic features of product data, wherein the multimodal semantic features of product data include textual semantic features and image semantic features of product data.
6. The automatic quality inspection method for multimodal product data on e-commerce platforms according to claim 5, characterized in that, The multimodal semantic features of the product data are subjected to semantic alignment and semantic consistency processing to obtain the text-image semantic matching features of the product data, including: Cross-attention calculation is performed on the product data text modal semantic features and the product data image modal semantic features to obtain the product data image modal related semantic features corresponding to the product data text modal semantic features and the product data text modal related semantic features corresponding to the product data image modal semantic features; The semantic features of the product data text modality are concatenated with the corresponding semantic features of the product data image modality to obtain the product data text modality enhanced semantic features. The semantic features of the product data image modality are concatenated with the corresponding semantic features of the product data text modality to obtain the product data image modality enhanced semantic features. Obtain preset text attention weights, perform global average pooling on the product data text modality enhanced semantic features and the product data text modality enhanced semantic features, and perform weighted fusion on the product data text modality enhanced semantic features after global average pooling based on the text attention weights to complete semantic alignment processing and obtain product data multimodal collaborative semantic features; A preset semantic similarity perception model is obtained, and the multimodal collaborative semantic features of the product data are used as input to the semantic similarity perception model to generate image-text semantic matching features of the product data through the semantic similarity perception model.
7. The automatic quality inspection method for multimodal product data on e-commerce platforms according to claim 1, characterized in that, A semantic graph of quality inspection rules for product data is established, and the semantic graph of quality inspection rules for product data is used to perform violation inspection on the multimodal enhanced scenario graph of product data to generate corresponding semantic features of product data violations, including: By obtaining basic quality inspection rules and commodity data supervision regulations from e-commerce platforms, semantic recognition is performed on the basic quality inspection rules and commodity data supervision regulations to define multiple commodity data quality inspection violation patterns, and the corresponding pattern feature relationships are identified for each commodity data quality inspection violation pattern. By using each of the aforementioned product data quality inspection violation patterns as quality inspection rule nodes and identifying the corresponding pattern feature relationships of each product data quality inspection violation pattern as quality inspection rule edges, a semantic graph of product data quality inspection rules is constructed. Each product entity node and its corresponding product entity attribute node are extracted from the multimodal enhanced scene graph of the product data. The product data quality inspection rule semantic graph is used to perform violation inspection on each product entity attribute node corresponding to each product entity node. When the product entity attribute node corresponding to the product entity node conforms to the product data quality inspection violation pattern in the product data quality inspection rule semantic graph, the product entity node is considered to have violated the rules. Traverse all the product entity nodes, perform violation inspection on each product entity node to filter out each product entity node that has a violation, and add corresponding product data quality inspection violation pattern tags to each product entity node that has a violation according to the semantic graph of the product data quality inspection rules to form a violation product entity node. The various non-compliant product entity nodes are integrated to form a non-compliant product entity node matrix, which is then used as the semantic feature of non-compliant product data.
8. The automatic quality inspection method for multimodal product data on e-commerce platforms according to claim 1, characterized in that, A pre-trained violation prediction model is obtained, and the multimodal feature data of the e-commerce platform products, the image-text semantic matching features of the product data, and the violation semantic features of the product data are fused into violation features of the e-commerce platform product data. The violation prediction model is then used to perform quality inspection and judgment on the violation features of the e-commerce platform product data to obtain the quality inspection results of the e-commerce platform product data, including: The e-commerce platform's multimodal feature data, the product data's image-text semantic matching features, and the product data's violation semantic features are subjected to dimensional standardization processing to obtain standard e-commerce platform product multimodal feature data, standard product data's image-text semantic matching features, and standard product data's violation semantic features; Obtain preset feature importance weights, and then perform a weighted summation on the standard e-commerce platform product multimodal feature data, standard product data image-text semantic matching features, and standard product data violation semantic features according to the feature importance weights to obtain the e-commerce platform product data violation features; A pre-trained violation prediction model is obtained, and the violation features of the e-commerce platform's product data are used as input to the violation prediction model so that the violation prediction model can output the product data violation confidence level. A preset violation confidence threshold is obtained, and the violation confidence threshold is used to perform quality inspection on the violation confidence of the product data to obtain a quality inspection result. The quality inspection result is used as the quality inspection result of the product data on the e-commerce platform. If the violation confidence of the product data is higher than the violation confidence threshold, the quality inspection result is output as "confirmed violation". If the violation confidence of the product data is not higher than the violation confidence threshold, the quality inspection result is output as "confirmed compliance".
9. An automated quality inspection system for multimodal product data on an e-commerce platform, characterized in that, The method for automatic quality inspection of multimodal product data on e-commerce platforms as described in any one of claims 1 to 8 includes: The product data acquisition unit is used to acquire the original product data of the e-commerce platform, and to perform data preprocessing and multimodal feature extraction on the original product data of the e-commerce platform to obtain multimodal feature data of the e-commerce platform products. The multimodal feature data of the e-commerce platform products includes text feature data of the e-commerce platform products, image feature data of the e-commerce platform products, and image tampering feature data of the e-commerce platform products. The scene graph generation unit is used to acquire a preset e-commerce knowledge graph, construct a multimodal entity scene graph of product data based on the multimodal feature data of the e-commerce platform, and perform product entity recognition and entity attribute information extraction on the multimodal entity scene graph of product data and the e-commerce knowledge graph. The identified product entities and the entity attribute information are fused in the multimodal entity scene graph of product data to obtain a multimodal enhanced scene graph of product data. The image-text semantic matching unit is used to obtain a pre-trained semantic extraction model, use the semantic extraction model to extract semantic features from the multimodal enhanced scene graph of the product data to obtain multimodal semantic features of the product data, and perform semantic alignment and semantic consistency processing on the multimodal semantic features of the product data to obtain image-text semantic matching features of the product data. The data violation inspection unit is used to establish a semantic graph of product data quality inspection rules, and to use the semantic graph of product data quality inspection rules to perform violation inspection on the multimodal enhanced scene graph of product data, so as to generate corresponding semantic features of product data violations. The feature quality inspection and judgment unit is used to acquire a pre-trained violation prediction model, integrate the multimodal feature data of the e-commerce platform products, the image-text semantic matching features of the product data, and the violation semantic features of the product data into the violation features of the e-commerce platform products, and use the violation prediction model to perform quality inspection and judgment on the violation features of the e-commerce platform products to obtain the quality inspection results of the e-commerce platform products.
10. An electronic device, characterized in that, The device includes a memory, a processor, and a transceiver that are sequentially and communicatively connected. The memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the automatic quality inspection method for multimodal product data of e-commerce platforms as described in any one of claims 1 to 8.